HPE news. Tech insights. World-class innovations. We take you straight to the source — interviewing tech's foremost thought leaders and change-makers that are propelling businesses and industries forward.
PRAVEEN JAIN
30 years ago. the port speed used to be 10 megabits per second, megabits. now the speeds are 1.6 terabits per second for a single port. So it means the bandwidth demand is just too huge.
Here, any congestion in the network, GPU slow down or GPU cannot any congestion perform the tasks what they intend to do. So you've invested millions or billions of dollars in your GPU infrastructure, and if network is not delivering the performance, you will be wasting lot of high GPU cycles
SAM JARRELL
Wow. Well, I mean, at this rate, the network has gone from sort of a bicycle speed to warp drive, but basically the GPUs are saying, "Keep up or get out of the data path." it sounds like there's a lot of money at stake.
MICHAEL BIRD
Yeah, that's a, that's a great analogy.different bits of technology has been saying keep up at various points in time, But 10 megabits to 1.6 terabits, my goodness me.
Now, Sam, this week we're gonna be diving back into the hidden networks which prop up modern technology to see how much has changed in just six months since we last explored it, and it is a lot in the last six months.
I’m Michael Bird
SAM JARRELL
I'm Sam Jarrell
MICHAEL BIRD
And welcome to Technology Now from HPE.
MICHAEL BIRD
Now, for years, we've focused on GPUs and models when it came to conversations about AI. However, in the past 12 months or so, that conversation has expanded to include networking, because behind every AI breakthrough is an absolutely vast amount of infrastructure
SAM JARRELL
That's right. And of course, this infrastructure is evolving just as fast as the rest of technology because
with GPUs costing tens, if not hundreds of thousands of dollars each, you really can't afford to have them just sitting idle
MICHAEL BIRD
Yeah, e-exactly. And so the sort of underlying infrastructure that needs to support these GPUs needs high bandwidth, lower latency, really effective cooling, and it needs to be able to respond and mitigate any issues which could arise within the system. Basically, there's a lot going on
SAM JARRELL
Yeah, but solving problems takes time
MICHAEL BIRD
Yeah. And if there's a way to reduce how long it takes to deal with problems, say, oh, I don't know, something like a self-healing network infrastructure
SAM JARRELL
That sounds like it would be a great idea
MICHAEL BIRD
Yes. I'm so glad you said that, because today we are revisiting the topic of networking within data centers because just like all other aspects of technology, this is also advancing at breakneck speed. So, to find out more, I met with Praveen Jain, SVP GM for the data center networking business within HPE, and he started off by giving me a quick reminder of the difference between AI for networks and networking for AI.
PRAVEEN JAIN
So, let me answer this in two different ways. First of all, we have been running networks for decades, traditional networks, where we are connecting the compute, storage, and, whatnot, right? The traditional networking. Now, with the invent of AI, can I simplify that operation? Instead of every day you are monitoring what's happening, and whenever something fails, you immediately jump in, try to figure out what the root cause of the problem is.
It takes days and months to figure out all of this, right? So first factor is, which I call AI for networking. I use AI technology to simplify my traditional networking. In other words, I want this networking to drive by itself. You deploy the boxes, and done. It figures out where the problem is automatically, predicts where the problem could happen in future, and automatically corrects.
So, so far I covered AI for networking, where I'm using AI technology to simplify my network. Now let's talk about networking for AI. So, you know GPUs, everybody's talking about GPUs.
In a box, if you have eight GPUs, fine, that's one server box with eight GPUs, but you cannot do meaningful work or large scale work with these eight GPUs.
So you have to connect these GPUs to another set of GPUs, and guess what? To connect need network. And I call that as network for AI. So in other words, you need high performance gear or high performance networking. Just to give you an instance, we just released a product which is over hundred plus terabits per second.
Okay? It has sixty-four ports of one point six terabits per second each. That's the kind of capacity you need to connect these GPUs, and if you cannot connect these GPUs, you cannot do the meaningful workload what you are looking at. And everybody's talking about, "Hey, I can do the large scale models, I can do prediction here and there," but ultimately you need the network.
So that's the reason networking is so critical
MICHAEL BIRD
Can you talk through, some of the key differences of, a network built specifically for AI? You know, if you were building a network today, what that would look like versus a more traditional network
PRAVEEN JAIN
Very good. So I started my networking career, 30 years ago. At that point of time, the port speed used to be 10 megabits per second, megabits. And I just described now the speeds are 1.6 terabits per second for a single port. So it means the bandwidth demand is .
On top of that, I think the traditional workloads were more resilient to failures here and there.
Let's say one packet drops here and there, you have TCP/IP which will recover.
Application would not even notice. Here, in the network, GPUs slow down or GPU cannot any congestion perform the tasks what they intend to do. So you've invested millions or billions of dollars in your GPU infrastructure, and if network is not delivering the performance, you will be wasting lot of high GPU cycles. but also in addition, let's say we talked about the capacity, we talked about the congestion avoidance and making sure you utilize the network, not blocking the GPUs. Power requirements have skyrocketed, as you know. Everybody's talking about that. But last but not the least, cooling. Like my switch which I released, it's 100% liquid cooled switch, so it means I don't have any fan in my box. Compared to the traditional networking, it was always fan-based box. So those are some of the key differences.
MICHAEL BIRD
so quite different looking from an infrastructure perspective. In terms of, the topology or how you architect it, is that changing as well?
PRAVEEN JAIN
somewhat. obviously we used to do the factory topology as we used to call them. but now rail optimized designs are some of the designs people are following, to make sure that intra GPU communication, and all of that is done optimally.
MICHAEL BIRD
and so we're talking here about, I guess AI data centers. What about, traditional data centers for organizations? is the way that those networks are architected, Are they changing as well because of, the rise of AI, the change in technology in the industry?
PRAVEEN JAIN
So from the traditional data centre perspective, as
I was mentioning, topologies and the networks are not changing necessarily. But on the other hand, if I'm using AI to simplify day-to-day operations, you do not want the switches and routers and networking which was built in decades ago, and they don't know how to make use of the AI.
MICHAEL BIRD
And, from my understanding, AI workloads have different bandwidth usages in terms of, you know, on a traditional network, you do lots of caching and traffic can sometimes flow more in one direction than the other. and maybe also you'll have different times where that traffic is flowing.
and my understanding is that's different for AI workloads. and again, that has impacted the way that networks are architected.
PRAVEEN JAIN
Very good question. So earlier, and it's a little bit more technical, but I still want to mention it, lot of times we did the, load balancing. So let's say in the network you have multiple paths, to go to a destination. Let's say from your home you are going to some place, you have multiple paths to go there.
Depending on how much traffic it is and things like that, maybe the maps will reroute you to the appropriate place. Same thing within the network. And what we used to do in the traditional network was, and without going into too many details, we used to call it a five-tuple based load balancing. So you hash based on some key things and just go. In that, you were not necessarily looking at how much the link is busy or how much is it utilized. So we at HPE have innovated a lot of technology for this AI networks specifically, where we want to make use of this network in a best, possible way in two terms. One, utilize the network to the fullest.
Don't leave it underutilized. But also, let's say there is a heavy congestion happening on the link. You know that and you can see that in the network. Why don't you just reroute the traffic to non, congested link? So all of those technologies are, again, shipping to the customers, and customers are deploying those technologies today to effectively utilize their GPU cycles
Cycles
MICHAEL BIRD
So. this has mainly been a conversation about, networks for AI.
But I want to ask you about AI for networks. so are we at the stage now where AI can just run our network, run our data centers?
PRAVEEN JAIN
Yeah, actually, so good part is, what we describe here as self-driving network is not a vision. We are delivering it today. We are not fully there to the final level of the stage five self-driving network, but we are so far ahead in this journey, nobody else on the planet
MICHAEL BIRD
And can you, can you just talk through, like when, when we say self-driving network, what do we mean?
PRAVEEN JAIN
So, think about this as a self-driving car. Very e-easy, analogy, right?
Today, if you go to a regular car versus a self-driving car, you can see the difference. There's no driver. It knows how to turn, how to react, all of that, right? same thing in the network, where in one hand you have old networking boxes where you have to do manual on, operations on them, you have to configure them. When any issue happens, you have to jump into root cause. Compare that with HPE's self-driving network, which automatically detects issues. It also brings human in the loop before corrects, correcting it so that you can trust the system. So all of that. But also I want to point out one important distinction. Let me ask you a question. Pick a car. What car do you drive maybe?
MICHAEL BIRD
I drove a, I drove a people carrier Right.
PRAVEEN JAIN
All right, so let's take that car. Do you think I can take that car without changing the hardware? Can I make it self-driving?
MICHAEL BIRD
Not, I mean, if you push it down a hill, it'll be self-driving for a bit.
PRAVEEN JAIN
I said it So that's what our competitors are doing. They are taking their traditional products, they're pushing it down the hill and saying, "I got the self-driving." No, you cannot do that. So it means you need to have a purpose-built car, and in this case, purpose-built hardware in networking to do the self-driving operation
MICHAEL BIRD
So, what does the data center of the future look like?
PRAVEEN JAIN
Again, I think it's again our self-driving vision where I really want to get to a stage where people trust the system, and of course, they can still be notified and they can still be in the loop, but at a point where you trust it enough that you don't have to wake up for a 2:00 AM call saying that something is not working in the network.
It automatically detects, fixes it, and just sends you a notification, "Hey, this is what I fixed and this is the outcome." By the way, this, Network should not say that I fixed something. It should say, "Look, these applications were having problem. By fixing it, now I show you that the applications are no longer having a problem
MICHAEL BIRD
I guess you want a network where it's not just, being a firefighter, but being proactive. and, getting more value out of what you're spending money on.
PRAVEEN JAIN
Very well said. Exactly right.
MICHAEL BIRD
Praveen, thank you so much for joining us on Technology Now. It's been an absolutely fascinating conversation.
PRAVEEN JAIN
Thank you for having me here. I really enjoyed this conversation. Thank you
SAM JARRELL
one of the themes running through, several recent episodes of our show is that AI depends on layers and layers of infrastructure that most users never see. I think for a while there's been a lot of focus on GPUs and chips and things like that. But increasingly, networking really is becoming sort of the star of the show.
And, continuing with you guys' car analogies, you were making me think of, okay, you have all these fancy GPUs, you have, the compute power that's like a race car, and then you're taking it to, the children's pickup line at a school because the network is not fast enough. it's interesting
MICHAEL BIRD
Yeah, it's a really interesting point, the headlines in the news is, all about GPUs when it comes to AI.
the big conversation is, GPUs crossing borders or, which company's got the most GPUs, et cetera, et cetera. But, if you haven't got anything to plug them into, then it's completely pointless, isn't it? as you said, it's like having a sports car,
but putting square wheels on it
SAM JARRELL
the innovation can't just be happening with AI or the GPUs themselves. It has to be happening across the technology stack or it's all a bit useless. but one of the things that was pretty clear from the conversation is that the expensive GPUs can't afford to wait and, Praveen was talking a lot about how organizations are investing millions or even billions into the infrastructure and if they're stuck waiting on network, congestion or something else or a problem, then it's not even just like a technical issue anymore.
It becomes straight up a business-wide issue. the network isn't just supporting AI itself, but it's also protecting the value of these investments.
MICHAEL BIRD
Do you know what? It reminds me of, some of the issues we're having here in the UK with our grid. we've got all of these infrastructure projects where people are building, solar farms or wind turbines, but the big issue that they're having is they can't connect it to our grid. and there's, grid connections are taking decades, sometimes to be put in because there's just, there's such a demand.
It feels like it's the same thing. It's like, you can build all this incredible infrastructure, you can have these amazing GPUs, but if if the industry isn't innovating alongside it, then completely pointless.
you're just spinning turbines. You're just spinning wheels.
SAM JARRELL
that's pretty topical. And I mean, we talked a lot about energy consumption and, AI and data centers and things in the past couple of episodes too. It just does go back to the idea that none of this is happening in a silo. I think we're in kind of an interesting time period where there's a lot of innovation happening all at once, and it has to happen, or none of these things or the promises of these things will never be delivered.
but I think for most people, or at least most like media outlets and things, the focus just continues to be AI models, AI models, AI models. But there's so much more happening all around us right now
I also appreciated that once again I felt like Praveen reinforced this idea of like the human in the loop, and that, the self-driving network still does include people. ... But just the role of people increasingly focuses on what we're good at, which is judgment and oversight rather than manually hunting down individual problems.
MICHAEL BIRD
Do the stuff that humans are good at. Let the computers do the thing they're good at, yeah. Now, Sam, Praveen mentionedearlier that he's been in the industry for 30 years, and he talked about how much the bandwidth requirements have changed in that time.
So the final thing I wanted to know was, other than, port speed, what change has shocked him the most since he started in the industry?
MICHAEL BIRD
Now, Praveen mentioned earlier that he has been in the industry for 30 years and he talked about how much the bandwidth requirements have changed in that time so the final thing I wanted to know was, other than the port speed, what change has shocked him most since he started in the industry?
PRAVEEN JAIN
I think since I'm in the AI era, I would say AI was the biggest change which shocked me the most. And part of the reason we have been always talking about AI in some form or the other, even this concept of self-driving networks and all people have talked about here and there. But with the invent of latest LLMs and the AI, now it's possible.
So that's the change which shocked me the most
SAM JARRELL
Okay that brings us to the end of Technology Now for this week.
Thank you to our guest, Praveen Jain
And of course, to our listeners.
Thank you so much for joining us.
MICHAEL BIRD
If you’ve enjoyed this episode, please do let us know – rate and review us wherever you listen to episodes and if you want to get in contact with us, send us an email to technology now AT hpe.com
Sam, subject line suggestion?
SAM JARRELL
Race car in the school pickup line?
MICHAEL BIRD
Race car in the school pickup line. as a Brit, not that I know what a school pickup line is.
and don’t forget to subscribe so you can listen first every week.
Technology Now is hosted by Sam Jarrell and myself, Michael Bird
This episode was produced by Harry Lampert and Eva Higginbotham with production support from Alysha Kempson-Taylor, Nik Damarell, Beckie Bird, Zoe Revis, Alissa Mitry, and Jenessa Ayache. Our theme music was composed by Greg Hooper.
SAM JARRELL
Our social editorial team is Rebecca Wissinger, Judy-Anne Goldman and Jacqueline Green and our social media designers are Alejandra Garcia, and Ambar Maldonado.
MICHAEL BIRD
Technology Now is a Fresh Air Production for Hewlett Packard Enterprise.
(and) we’ll see you next week. Cheers!
SAM JARRELL
Bye y’all