
Loading summary
Andrew Feldman
This is the largest chip built in the history of the computer industry. It's 58 times larger than a GPU. And for AI, bigger chips process information more quickly, and therefore you get answers in less time. For AI work, big chips are undoubtedly the best way to go. There's no mode in inference. It takes you eight keystrokes to move from a GPU to us in the cloud. We solved a problem that nobody in the history of compute had solved, and we delivered it in 2020, and nobody cared. Nobody cared. Nobody bought any and nobody cared. Everybody said we were crazy, it would never work. So then we built the next one.
Matt Turk
Hi, I'm Matt Turg. Welcome to the Matt Podcast. My guest today is Andrew Feldman, co founder and CEO of Cerebras, the company that built the largest chip in the history of computing and just pulled off the biggest semiconductor IPO of all time. Andrew has been everywhere talking about the headlines, the $20 billion plus OpenAI deal, the IPO. But this conversation is a bit different. What is a wafer? And built up step by step. Why GPUs struggle with fast inference, the three shortages. Nobody talks about the decade in the desert when nobody wanted his chip. And why Andrew believes that Cuda is no longer a moat for Nvidia. If you want to actually understand the current chip landscape and how AI inference works at the silicon level, this episode is for you. Please enjoy this fantastic conversation with Andrew Feldman.
Interviewer
I found a fun place to start would be to talk about speed. So has speed become the dominant conversation for AI today?
Andrew Feldman
Yeah. What happened, I think was for a long time, AI was sort of a novelty, right? It was like a parlor trick. It was cool, but not useful. And what happened somewhere around the middle of 2025 was the AI got smart enough such that people began to use it. And we remember, we make AI with training, but we use it with inference. And suddenly people wanted to use it. And the minute you want to use it, the minute it's productive. Speed matters fast. Tokens are more productive. And so the conversation moved from everything else to how do we make our inference faster? How do we deliver tokens more quickly? Because those are more productive tokens. We get more done in less time, and therefore it's more valuable.
Interviewer
Right. And what does speed mean? Is that a question of token speed? Is that completion of the task? What's the right metric?
Andrew Feldman
So the right metric is tokens per second per user. That's how fast you get the first token, all the way through the last token in your. In your response. And it's true for queries from chat, but it's also true for agentic floats. Right. If they're sort of multi cycle turns, waiting is amplified. And so what you want is blisteringly fast responses so that the AI feels like it's in real time. You can engage with this a thing.
Interviewer
So it's the Brahmin moment for AI.
Andrew Feldman
I think, I think that's right. And I think that's a very good analogy. I think if you think of something like Netflix. When the Internet was slow, Netflix delivered DVDs and envelopes. You would get a DVD envelope. And when the Internet became fast, they didn't get more efficient at delivering DVDs and envelopes. They became a movie studio. Right. The speed enabled them to become something completely different. And that's what speed does in general and in particular for, for AI, it opens up a whole new domain. It allows you to use the AI differently. You will stay longer, you will come more often and you work on harder problems.
Interviewer
Yeah. So it's literally a question of the ux. Right. That's just. Nobody wants to wait a few seconds.
Andrew Feldman
That's right. I mean how big is the market for slow search? How big is the market for dial up is zero. How big is. How long will you wait for a website to resolve? Will you wait eight seconds? Nobody waits. And so it's the exact same with AI.
Interviewer
Yep. So no more people waiting with their laptops open while the agent.
Andrew Feldman
That's right. Well, it's running and running and running. I think that is not what what people want.
Interviewer
Okay, great, Wonderful. So I'd love to talk about the landscape of the chip industry right now to help people visualize and part where you guys are. So there used to be basically this concept of like one chip to do it all. And obviously this is evolving dramatically. This, you guys, this grok. There are GPUs, there are people have record of Trainium, people have heard of tpu. So help us compare and contrast. Like who does what for what?
Andrew Feldman
Well, I think there used to be one chip to do it all. It was called the cpu. Right. And there emerged a CO processor to do discrete graphics. And as the AI workload became interesting, we focused chips more and more on that particular workload. And today several Companies make traditional GPUs. So Nvidia and AMD it make very standard GPUs. There are a group of companies who the hyperscalers make some of their own parts. So the tpu.
Interviewer
So TPU is Google, TPU is Google
Andrew Feldman
and Trainium is aws. And then there were group lights, Cerberus. And we were among the pioneers to build a part from scratch optimized for AI and nothing else. And that we weren't inside of a hyperscaler. We weren't optimized for a hyperscalers problem or for one labs problem. We were building a chip for a collection of AI problems and all our thinking was around how to accelerate AI. And that's sort of the landscape to this day.
Interviewer
There's even more specialized ones.
Andrew Feldman
Right.
Interviewer
Like Edge developed specific transformers. Those are called ASICs as well. Can you maybe define what that term means?
Andrew Feldman
An ASIC is an application specific integrated circuit and it's a word that now has a wide range of meaning. It means you have made a series of choices away from general towards a narrower class of problem solving. And that you've made some decisions in your architecture that make it much better at some things and much much worse at others. Right. And that's choices that are made across the spectrum. So TPU has made some choices like that it can't do graphics, it's very good matrix multip, not very good at a collection of other things that we do in, in mathematics. Same for the gpu. We, we've all made different choices Right now in production there are nd, there are a collection, there's the TPU for Google Trainiums just coming up is a Maya part is apart from Microsoft Cerebras. And there was one other GROK that got acquired by Nvidia. Yeah.
Interviewer
And to the general chip versus pixelized chip. It was actually interesting because you guys and your CTO celebrated when Grok was acquired. What was that? Was it a recognition by Nvidia that your vision was right all along?
Andrew Feldman
Yeah, I think one of the things ideas, sort of most durable moats was the perception that the GPU could do everything and it was all you needed for AI. And the acquisition of Groq for $20 billion on the structure and the speed with which they chose to do it made clear to everyone that that wasn't true. That the GP architecture couldn't do, could not do fast entrance and that this market was large and growing quickly and we were the fastest at it and the largest and you know, our sales were more than 10 times the Grox and they paid $20 billion for the number two player. So that was a good day. That was a great day.
Interviewer
And the recent announcement of jalapeno between OpenAI, which is a very large customer of yours and Broadcom, where does it fit in that picture is that one of those highly specialized ASICS types.
Andrew Feldman
Right.
Interviewer
Grok and it's for inference.
Andrew Feldman
That's right. Jalapeno is a part that has been a long time coming. It was announced with Broadcom. Before OpenAI did the deal with us November we did a huge deal. This is the largest deals in Silicon Valley history, north of $20 billion. I think they have a yawning need for silicon and one of the things that Open Air has been, I think the best in the industry at is looking at an exponential curve, adoption, rate of growth, not being afraid of what it says. Right. Others have had to go out and strike really bad deals to get capacity. Where OpenAI saw this coming, they struck big deals for memory, for compute with us, with others. They've really been sort of visionary in understanding what it means to extrapolate from an exponential curve. And that curve is AI usage. It is growing so unbelievably fast. The whole industry is chasing demand. Right. Often it's the other way. Often people are building it, hoping it will come. And in our case all of us are chasing what people already want to do, let alone what they might do in the future.
Interviewer
As I put opinion madamy and I may, because that's such an interesting topic to finish on. GPAs versus inference versus ASICS. Is that the permanent situation from your perspective that we're going to be in this forever? Multi silicon kind of environment?
Andrew Feldman
Yeah, multi silicon environments are healthy ecosystem. Right. I don't think anybody would say that the x86 environment was a healthy place. There were for 20 years there were two players and if you notice when a new workload came around, cell phone workload, it was very similar but required low power and battery and both lost. Right. Intel got zero share, MD got zero share and a company nobody previously heard of became the largest seller of compute in the world. And that's R. And that healthy ecosystems have lots of different ways to solve problems.
Interviewer
Where does China fall in all of this? So there's no Huawei chips for deep sea can this emergence of a full stack Chinese AI factory, for lack of a better term. Is that divide happening or is that overstated?
Andrew Feldman
I don't know that that divide is real. I think they are an industrial adversary. I think they have made really interesting investments. They've made investments in power, the grid, which is a real weakness in the US Here in France you have nuclear which sort of turns out to be pretty cheap power and pretty clean. But in China they made huge investments in their grid. They are behind in Chips, but their approach was at the next level as open source models where they're producing some extraordinary models. Not as good as GPT or Anthropic or Google's Gemini, but very good. And I think they don't have the same chip strikes that the US does, but they've got leadership in other domains, particularly in power, which is what we need for data centers.
Interviewer
Okay, what does China fall in survey strategy? We don't sell questions for geopolitical reasons, for business reasons, for regulatory reasons, for
Andrew Feldman
regulatory reasons and for some geopolitical reasons.
Interviewer
Okay, is there a role for local AI and local chips? Nvidia had some announcement around just building chips for Windows computers. Is it for inference? Is that something that you guys think should be part of the multi silicon ecosystem?
Andrew Feldman
Yeah, I think if you look at the way the ecosystem for apps and clouds emerged, everything you can do, you should do on your phone or your laptop. But the, the, the ability to get real processing power to a phone or to a laptop is constrained because they're generally working off a battery. And so you want to do as much as you can as, as close to the data as you can. But the truth is, is in most situations for real compute, you have to go up to the cloud, to the data center. And that's exactly the way it's going to be with AI. We're going to do a lot of
Interviewer
work
Andrew Feldman
on the cell phones, on a laptop, but for the big work you're going to go to the data center. And that's where our focus is. Our focus is data center compute free.
Interviewer
Great. So as you mentioned, how powerful the demand was and of production. So just to ask the inevitable question around market equilibrium bubble or not. So like the most recent thing we all saw is that there was a little bit of a flash crash earlier in June where 1.5 trillion in market value was lost and Nvidia had lost 300 billion that day. Now that was corrected within, you know, 48 hours. So the market seems very jittery around that, that general concept. Maybe just putting words in your as well just based on what you said, like you're, you're reasonably on the other side of that okay. Concern.
Andrew Feldman
I, I, I, I think that, that if you spend a lot of time staring at the, at the public markets and watching their ups and downs every day, you're, you're watching the wrong thing. I, I think that, right. The wisdom received from the great thinkers of public market investing is that in the short term the public markets are voting mechanisms. Who's most Popular, but in the long term they're weighing mechanisms. Who's created the most value, the most weight. So what we can do as a public company, as soon as we cannot look at the day to day fluctuations and we can focus on just building extraordinary technology, winning new customers, making our existing customers happy and sharing they buy more, being better every week. And so I think the day to day fluctuations are out of our hands. And that's about the voting.
Interviewer
Yeah, maybe just to play a devil's advocate on the demand, it seems like a lot of the demand comes from the big labs which themselves are financed by venture capital, private equity, hedge funds, whatever you call it, sovereign investors. Is there any concern that demand for chips comes from the labs which are maybe artificially financed and that if you add a few circle of deals here and there, don't you sense any fragility? I.
Andrew Feldman
That's not exactly our experience. I mean obviously we have enormous demand from opening, but we have huge dozens of other customers who are trying to place very big orders off. And historically bubbles were when supply got out ahead of demand, when in the 90s we built out telco infrastructure, we built out fiber years before it was going to be used and it took six or eight years and it all got used. But it was a sort of if you build it, they will come mentality. Whereas what's different about AI right now is we're all trying to catch up. We're trying to build data centers faster, we're trying to increase our demand, our supply chains for demand that's already here today. And only a very small portion of the world are using AI anywhere close to its potential. And we're already sort of overwhelmed with compute. We're overwhelmed with the demand for memory, which is a real weakness in the GPUs. It's not a problem. We face the ecosystem, there are constraints left and right. And that doesn't feel like a bubble.
Interviewer
Do you want to actually double click on that? Because that's so interesting. Right. It seems like the bottleneck keeps moving because. So what's the current problem of a memory shortage that people may have heard of?
Andrew Feldman
So there are three major bottlenecks right now. The first is all GPUs and all ASICs except US use a type of memory called DRAM and a particular flavor of that memory called hbs. And that memory is made by three companies in the world, Hydex, Samsung and Microsoft. Three companies, and they're sold out. That's the number one problem.
Interviewer
We don't use it and it's Sold out, as in like the only way to. As old things, hardware, the only way to increase supply would be to build
Andrew Feldman
new factories, build more factories. And so you have this huge problem for GPUs and even now for CPUs in which you can't get this memory because of our architectural choices, we don't use it. The second constraint was a process at tsmc. Process was called coas. And this used silicon as a motherboard. And this is what Nvidia and AMD and others use as a motherboard. So they put their chips and the memory chips on it and it's more efficient than a traditional motherboard. Ahead, that process is sold out, is highly constrained. We don't use it. Third limitation in the space is 3 nanometer factory space at TSMC. So this is the most aggressive factory, so the most advanced technology, our chips are at 5 nanometers. So we don't use the 3 ganometer technology.
Interviewer
And so that's so fascinating. So to ask the layman's question, so what does that mean? Streamulator factory. That's a factory that's solely dedicated.
Andrew Feldman
That's right. They build factories and the factory etches transistors that are a certain distance apart. And the smaller the distance apart, the more transistors per square millimeter and the more you can do per square millimeter. And so the history of the computer industry making chips has been we've gotten better and better at putting transistors closer and closer. And so we used to be at 16 nanometers apart, and then we went to 10 and to 7 and to 5 and 3, and they're working on 2.1.8. And so right now the bulk of the GPU world is all three. And there's tremendous congestion at factory.
Interviewer
And that's literally different tooling at the
Andrew Feldman
factory, tooling at the factory. And we use the 5 nanometer factory. And this is our ship here, just the camera. This is the largest chip built in the history of the computer industry is 58 times larger than a GPU. It has 2 and a half or 3,000 times more memory bandwidth. And for AI, bigger chips process information more quickly and therefore you get answers in less time. And obviously you don't want a bigger chip for a laptop or a cell phone or for lots of other things. But for AI work, big chips are undoubtedly the best way to go.
Interviewer
Okay, that's unique. All right, so let's. I want to do a bit of a deep dive on the chip itself in a second, but to close on this. So three bottlenecks you also mentioned CPUs. A couple of two minutes conversation. And there seems to be a theme around the emergence of like a CPU shortage as well. Is that so? One, is that true?
Andrew Feldman
Two, what causes it so agentic? AI is a world in which AI doesn't just provide answers, it initiates action. So that action might be go to a website, it might be learn, gather some data from a website, bring it back, take another action. Those actions are done by CPUs. And so as AI gets better and better at doing things, at making instructions, calls for things to get done, we're using more and more CPUs, and that is driving up the consumption of CPUs and therefore the demand for CPUs. And so this sort of huge push for more CPUs is being driven by AI on machines like ours and GPUs doing agentic work and asking the CPUs to take an action, to go to a website to order Brito to find a piece of information to pull it from storage to. All that work is being done by
Interviewer
the CPUs, including in a system like yours.
Andrew Feldman
Including in a system like ours.
Interviewer
Oh, it's interesting. So what does it mean? So you have your chip and CPUs on the side, and those would be a system.
Andrew Feldman
It means that AI, when we do AI in Agentic, the AI processor is like the brains and the CPUs are like the body. They're taking action. They're doing things in the digital world under the direction of the AI, which is running on the accelerator, on the streamer system. And so as we do more and more AI work and more and more agentic work, we're making more and more calls to CPUs. And therefore the demand for CPU is just through the roof.
Interviewer
And CPUs experience the same memory, same
Andrew Feldman
number shortage, they see exact same number. That's right.
Interviewer
Okay, yeah, let's go back to the product in a second. But let's talk about you and your journey a little bit for people to have contact. So, you know, obviously you've pulled off an incredible IPO with perfect timing. And I know, I know it was a little bit.
Andrew Feldman
The way you have perfect timing is to have horrible timing for 10 years. So you have perfect doubt and very
Interviewer
much to this point. So, like, you guys started building this, you know, bigger chip, focused on inference over a decade ago.
Andrew Feldman
2016.
Interviewer
2016. Which makes perfect sense in retrospect, given how long it takes to build a technology like this. But what was the vision? 2016? I guess it was like four years after ImageNet, deep learning was a thing.
Andrew Feldman
It was all vision, it was all evolutions. And that's why we, we don't believe the right thing to do is to embed the latest and coolest model into your circuitry. That's a mistake. We're the fastest in the world of transformers and our architecture was set before transformers existed. We're the fastest in the world of diffusion before an architecture was set before diffusion. What you want to do is get the underpinnings of those so that when the market moves, you can be good at that as well.
Interviewer
Otherwise, because of short life to the ASICS question earlier. So that's what others do. They build the architecture of the model into this.
Andrew Feldman
Some have, some have. And historically that's been a structural mistake.
Interviewer
It's been a structural mistake. So you're doing the broadly Oracle platform on which any kind of model you
Andrew Feldman
want to think about what is the underlying calculation? The underlying calculation at all. This work is sparse linear algebra. Yeah. And if you can accelerate that, whatever the bottle makers invent, you can make faster. And that was our approach.
Interviewer
Yeah. So you started in 2016, you had a prior company that you sold to AMD. What were some of the lessons you on there that you took into Cerberus?
Andrew Feldman
I think the lessons are large and many. I think what, what are the things of, you know, I guess experience is, is another name for having made mistakes and learning from them. Right. I, I think, you know, we, we have as a team built dozens of chips over the past 25 years, and the returns to experience in chip making is enormous. And we built a different type of computer at, at C Micro, a type of computer optimized for low power and optimized for a workload that was very different than AI for something like web browser. But the fundamental underpinnings, the questions you ask as a computer architect are always the same. What can I do to make this work faster? And is there enough of it to make it worthwhile? These are the two questions we ask. Should we build a part for it? Should we. What could we do to build a chip optimized for AI? And will there be enough AI so that you can build business around it? Those are the questions we asked in 2016. And you know, the, the flip side of that was, Arnie, wouldn't it be a surprise if the GPU, which had been optimized for graphics for 20 years, had been pushing pixels to a monitor, was suddenly good at a new world? Wouldn't that be serendipitous? And we came to believe that it wasn't the right architecture for it, it was just better than the cpu.
Interviewer
Right.
Andrew Feldman
And that we could build an architecture that would be faster, that would use less power and could drive down the cost of it.
Interviewer
And that was the journey. So service was very early, but then
Andrew Feldman
we spent time in the desert,
Matt Turk
then
Andrew Feldman
we wandered in the desert.
Interviewer
Yeah, maybe walk us a little bit through those years for the area founders, especially NDPIC founders. So listening to this, so what was. So first of all, what was the issue? Was it market timing? Was it the technology was not working? And then how did you go about it as a team? You'll be on board and your investors and raising more rounds as you presumably didn't have the proof points that you wanted to need.
Andrew Feldman
So for us at the beginning, we were honest with our VCs and told them we were going to attack a really hard problem. We weren't going to build something that was a little bit better than a gpu. And that our idea, our strategy was that you will never be a great company like Nvidia by doing something a little bit better, that they're going to buy everything for less, they're going to have price and pressure, they're going to be able to bundle. And the right strategy would be to do something incredibly hard in engineering that was way better. 10, 15, 20, 30, 50 times better. But to do that, the ordinary and obvious paths are all closed. Everybody else has taken the bar. And so what we observed was that speed in inference was going to be a function of memory, and that there were two types of memory. There's this drab of HBM and they can store a lot, but they're sort of. There's another type of memory called srum. It is unbelievably fast, but per unit area can't store very much. And for graphics, everybody had always used dram, that used hpf, and that the AI workload was fundamentally different. In graphics, you move data to the GPU and then you work on it for a long time and then you send the result. So the time spent in total of movement plus work is dominated by work in inference. In AI, it's the exact opposite. You move a huge amount of data, all the weights from memory to compute, and you do one calculation to generate the next word and then you have to do it again. So all the time is dominated by the movement of data. So that's why GPUs have so much trouble being fast. So we observed that if we chose a strategy using sram we could be faster, but then we have to overcome the trade off with SRAM that it can't store very much. That led us to the solution. If we could build a part vastly larger than any parting history, the size of a dinner plate, we could stuff it to the gills with SRAM and thereby get the benefit of sram that it was fast and overcome the weakness that it can't store very much. And that led us to a strategy called wafer scale. This chip is made from a single wafer. It comes out of tsmc.
Interviewer
Do you want to define what a wafer is?
Andrew Feldman
A wafer. All chips are cut from a wafer. A wafer is a circular piece of silicon that's 300 millimeters across. And the process of chip making stamps out like your mother does with a cookie cutter stance out chips. The biggest chip that had ever been built before us was 800 square millimeters. 840 to be exact. And this is 46,000. So we had to invent all sorts of new technology to build a chip this big. And once you build it, you have to invent ways to power it, to cool it to there are no vendors waiting for you because you look like nothing else ever made. And that took years. And it was a deep tech problem. It had never been solved before. And we had an 18 month period where we were spending 8 million a month and we couldn't build them. So for your deep tech founders, 8 million a month.
Interviewer
Now why so much? What was the cost? Just to understand how these businesses, because
Andrew Feldman
what everybody thought was hard, we solved quickly. And what nobody else knew about because they'd never actually done it, turned out to be really hard. Imagine, I tell people, imagine the first group that was going to climb Everest and they're at base camp and they're having tea with a group that just failed. And the group that just failed says halfway up there's this part. It's unbelievably hard. We couldn't do it. Your team climbs up, makes it all the way to the top, comes back, they're having tea again. And the team that made it leans over. The team that hadn't made it said that part in the middle. That wasn't the hard part because nobody had gotten past it. Nobody had gotten past certain things. So they didn't even know what to be afraid of. We now know. And it was something, a step called packaging. And that's how you affix a wafer to motherboard, how you deliver power to it and how you cool it. And nobody had done it before. And over that 18 month period, we became the best in the world at from approximately zero. And we did that by sailing again and again and using good engineering methodology and doing a failure analysis. Every single failure. So we fail differently again and again and again, again. And we told our board, you know, we met with our board over six weeks and yeah, this is the strategy. Nobody's ever done this before. This is what we're going to do. And we had to invent new materials, we had to invent new techniques. We ended up building things that, that everybody else had partners who could do. But In August of 2019, WinHouse, we had solved it. And my co founders and I, the first time it worked, it was sitting in a tiny little lab and we couldn't believe it. We were the first few things that we need to make one work. And we just stared at. Watching a server run is about as exciting as watching paint dry. There we were, the five of us, just staring at this machine, not believing it might work. We'd made it work.
Interviewer
Was that, was it a bigger moment that actually ringing the bell or just
Andrew Feldman
completely different events, like emotionally, emotionally, it, it was a completely different thing. It was that we had made our idea work. And I think for deep tech founders in particular, there's always this little thing in the back of your mind that says, maybe it's shit, right?
Interviewer
Maybe we're actually crazy. That's right.
Andrew Feldman
Maybe it's gonna fail and maybe it's not gonna work and maybe I don't have time. Maybe we're going to run out of money. Maybe, maybe, maybe. And the flip side of that is the joy that this was my co founder's ideas. I mean, these are their ideas manifest in the world. And that is an extraordinary thing. It's when someone's ideas take physical shape and then the next step is when you watch other people's work sit on top of your idea. Then you know you love making infrastructure. And so when we rang the bell and we went public on May 14 this year in the largest semiconductor IPO in history, and we did something unusual, we invited all our engineers who'd been with us since the start, everybody who'd been with us more than nine years and their families, and we all rang the bell together, we all got up on stage, that was sort of a moment of a different pride that we had done this together and that we had managed not to die. Sometimes with a startup, you know, people aren't honest, they don't tell you that. But we, we'd Avoided some. We'd made plenty of mistakes but we've been avoided the fatal ones and we, we'd made it through to a, a level of success that gave us the opportunity to, to pursue a new level of success. Right. That's what an IPO is is. It's not the end, it's you've gotten to a plateau that you can chase a new level of success see in the public market. And, and that felt pretty good. I would make.
Interviewer
Amazing, amazing. Thanks for sharing. So just go a little deeper on some of the product and technical stuff. So one of the trade offs of building a bigger chip, one that may come to mind, is failure mode. If you have all of little chips on a bigger big way for Prismware you can isolate the problems. If you have a big one, then everything would fail at the same time.
Andrew Feldman
So you have to think very carefully about failure mode. And we invented a technique that had about a million identical tiles and if one fails we can shut it down and we can use a redundant ones and they keep going. And so if you're going to go, you have to think about in the very architecture of the computer how you're going to manage failure. The JP's have a huge failure rate, so I'm sure you guys have spoken about this. Infant mortality is enormous and they fail all the time. There's good data. Facebook put out a paper on the number of failures they get in a big cluster. Now we can do some other things because we have all this compute in one spot, we can invest more to cool it. So we pioneered water cooling in AI systems and we run these much colder and GPMs. And the failure mode in electronics is temperature. And so by running them standard we are more reliable. And so that was an advantage. But again we had to advance the technique to cool off that ship this big. And so I think we have a system mentality, right? If you're going to do something big, you've made trade offs and you have to think about an architecture that allows for redundancy and repair. You have to think about the pros and cons, every aspect of the architecture and that's how you do something different and new.
Interviewer
And again, just to drive the point home, to make sure this I guess clear takeaway for people from this conversation. So explain it to me like I'm maybe not 5, but 15. Why is this faster than a GPU? Like what fundamentally makes it faster and what country is super fast with the gpa?
Andrew Feldman
How tokens are generated in inference is why it's faster. In an inference to generate a word. And our answers are a whole stream of words. They can be code, they can be pixels, but call them words talk. The weights of the model are moved from memory to compute. A calculation occurs and that generates the
Interviewer
word is a pre fill versus decode,
Andrew Feldman
that is both steps.
Interviewer
Okay, and do you want to maybe explain what pre fill and decode are?
Andrew Feldman
Okay. There are two steps in the computation necessary to do inference. When you type into chatgpt, explain to me the history of this village prior to World War II and it came to you. Two things have happened. The first thing is your prompt has been processed as step one and step two, your answer has been generated. The way your answer is generated is called decode. Decode is sequential processing. Your prompt, which we call presill, can be paralyzed, so you can process many of them simultaneously. But the speed with which you get an answer is a function of the decode and it is step by step by step in sequence. And that can't be changed. And how you do that step is you move weights, which are the intelligence from the model to compute to generate a word. You do a calculation and you get a word and then that word is used to generate the next word where weights are moved from computer. So the process is one of moving weights from memory to compute. So, well, how big are the weights? In a little model, like a 70 billion parameter model, the weights are about the size of 100 HD movies. So to generate a single word, you move 100 HD movies from memory to compute and then you have to move them again for the next word. And you want to do this a thousand times to get a good answer, a thousand word answer. This is where HBM is slow. This exact step is where HBM is slow. And that exact step, it is where by having all the SRAM here, we're blisteringly fast. And so the speed of moving weights to compute is about two and a half thousand times faster here than on a woven gp. And so that's the essence of what we've been able to do here and why it's so much faster. Yeah, fascinating.
Interviewer
Do you think that models need or will evolve as well in that fast AI world or is that just a problem?
Andrew Feldman
No, I don't exactly. Because remember, two things are happening. One, we want the models to be smarter, or two, one of the ways models are getting smarter is with rl. And RL uses inference inside a training. And so the faster you can do the inference insider training, the faster you're training.
Interviewer
Interesting is that A market for you guys in that's a market for us as well. So you're not just inference.
Andrew Feldman
We do RL and we do traditional trading too. Not for the largest models for the largest lab, but for the next tier
Interviewer
for pre training. For pure pre training.
Andrew Feldman
Pre training, fine tuning our full set.
Interviewer
So you can do pre training, I mean post training with RL, some pre training for the other labs. GPUs are still better than you for what job?
Andrew Feldman
So GPUs in training have some challenges that have been solved by a very narrow selection of the community. GPU is a very small chip and the calculations that we need to do in trading are very large. And one of the most complicated parts of training is the breaking up the calculations and spreading them apart on multiple GPUs. That's called distributed compute. That has historically been the domain of the supercompute world and is very difficult. Not just because cracking a problem and having lots of others work on it is hard. But they have to constantly share information and that sharing information is why they needed to buy Mellanox is they needed to control a fabric over which all is sharing in order to get an answer would happen that breaking up a big matrix multiply a big calculation is called running tensor model parallel. You are breaking up the tensor and spreading it apart. And the best labs in the world are good at that, but nobody else is. When we run training we don't have to run that way. We run what's called model data parallel. Data parallel is very simple. It allows teams who are good and very good to quickly test ideas in training. And so we are easier to use and faster because we allow them to use a technique which is much simpler.
Interviewer
Let's talk about the agentic world, maybe starting with reasoning. I think I saw a blog post where you guys have given me stated that reasoning was not always the right solution for all problems. Like how do, how do you think about this? Where does that fit in world?
Andrew Feldman
We're using it as a technique a little bit like when you were in eighth grade and you wrote different drafts of the paper. Single shot in front of you get an. You write a query, you get an answer. The easiest way to think about reasoning is you're going to do several drafts. It's going to break the problem up, it's going to solve in parts, it's going to bring them together, it's going to review the results, it's going to improve the results and then it's going to give you an answer. And so that is going to take more compute. And if your compute is slow, that's going to be more than an irritant. That could be crippling speed. And so as the best models, all of them, whether they're domestic or Chinese, whether you're OpenAI or anthropic or Gemini, whether you're any of the Chinese models, they all move to a reasoning approach. But that meant more compute was being used during inference. That was a huge advantage for us because we were fast and it made the GPU slowness stack up and it made our speed have an even bigger advantage. And so this was a huge boon for us. We think this is going to stay as a fundamental essence of the way these models are run right now.
Interviewer
And you also wrote about verification and whether what was the bottlenecking agents was whether the reasoning was good enough, whether they were smart enough, or whether there was a verification problem.
Andrew Feldman
Sure. I think the verification problem is a little bit like the guardrail problem. And what you'd like to do is after you written two or three drafts of your answer, you would like to be sure that it wasn't wrong. And that's your verification step. And you can do that maybe with a different model, you can do that by asking your existing model a similar question in a different way.
Interviewer
Right.
Andrew Feldman
All of these are ways you can pressure test your result. Guardrails can work the same way. You want to review, either with a model or with another technique through a scoring mechanism, that this question isn't out of line, that this answer isn't about how to make biological weapons or calling upon information that that you are direct to the FBI. All of that takes compute time. Whether you're trying to improve through reasoning or whether you're running guardrails. By being faster, you can get results in less time.
Interviewer
So what do you think the world is evolving for agents? Is it a bunch of smaller model running faster, doing more verification versus a large model?
Andrew Feldman
I think those work together. I think your big model produces an answer. Then you want to double. You just want to check your data. I mean, in the journalism industry, people would write papers that have data checkers. Right. Somebody would go and make sure. Back in the world, journalists check data via data checker. Right. That was a job. Each claim was checked independently. That's a different model. The main model wrote the piece, and then a little model checks some of the answers. And I think that's a very good way to go about it.
Interviewer
Where does a multimodality for the larger models fall in your world?
Andrew Feldman
We just announced Sort of that we were fastest in the world on one of Google's multimodal models. I think the truth is that there's very little text that doesn't have charts and graphs. Right. You must to understand text, be able to understand illustrations, graphs, charts. And so that's sort of the first and easiest part. And then you ought to be able to create book and then you want to understand images. And I think the new models are very, very good at that. Obviously what follows that is video because a video is just a collection of images. But that takes an enormous amount of compute right now. And that's one of the reasons it's been sort of set aside by the leading labs. So unbelievably computation intensity.
Interviewer
Great. Let's talk about the business a little bit because you guys are obviously a chip maker provider as we discussed, you're also a cloud provider, data center provider. What are the different parts we make
Andrew Feldman
compute and that computer is optimized for AI. It is the fastest at AI in the world. If you have a data center, we will sell you hardware for deployment in your data center. If you don't have a data center and you'd like to rent it by the month or the year, we have data centers so you can rent our, our equipment through our data centers and through our cloud. And so that allows us to get AI natives as well as large enterprises and governments.
Interviewer
And what's the.
Andrew Feldman
It was 50, 50 last year and I, I think this last quarter it was maybe, maybe 75, 25 in savor of hardware sales. I think this year it might be 50, 50. And as our OpenAI deal continues to unfold it will probably be 3070 with 30 on premise deployments of hardware at a SAP and D Cloud.
Interviewer
Great. Let's talk about that operating ideal since it's such like a major historical milestone record making. So it's providing up to 750megawatts which is, which is interesting by the way as a metric because you're a chip provider. But this is power. So is that shorthand for.
Andrew Feldman
It's a shorthand. It turns out right now and we didn't talk about this because there's sort of in the adjacent supply chain we went through the shortage of memory, we went through the shortage of a process called CO ops and 3 nanometer capacity. The other limitation in our industry right now is data center availability and that is a limiting factor for everybody. And that's why Anthropic did a huge sort of very expensive deal with Elon Musk for Data centered capacity. Our deal with them with, with OpenAI was because data center capacity is eliminating constraints measured the way data centers are measured in, in megawatts. The the deal is 760megawatts, 250megawatts in 26 on a multi year lease, an additional 250megawatts in 27 on multi year leaks and additional in 28 multi year leaks.
Interviewer
And you're doing that data center for them or you're providing the chips that go into the data center?
Andrew Feldman
Data center for them. We're delivering a full cloud solution so they connect to us via an API basically.
Interviewer
You mentioned 20, 26, 26 immediately. That's terrible.
Andrew Feldman
I am looking for data centers. My next meeting is in fact with a data center provider. We're doing a lot in Europe right now, a lot in, in the Nordics.
Interviewer
Is that because it's closer to power sources?
Andrew Feldman
Yes, it's because there is low cost power, clean, low cost power and low cost cooling. And does it matter?
Interviewer
Just like in the cloud business where the data center is located in terms of Spain.
Andrew Feldman
Yeah. There is an additional latency called transport latency and that's the speed of light through fiber to get from Helsinki to New York. And you have to account for that if your customers in New York and your data center is at Helsinki. Yeah. I mean it's usually about 2/3 the speed of light in case how long it takes you would like data centers on the same cognitive platforms.
Interviewer
That's the OpenAI deal. There was an exciting deal with AWS as well where you. It's a, it's a code chip solution.
Andrew Feldman
That's right. It's the disaggregated solution you mentioned before where their palladium part is doing the pre fill and is doing the parallelizable step. So that's trainium, that's training is doing the pre fill step and our chip is doing the decode and so you get a lot more bullish freely fast tokens. It's a good deal for us. It uses their data centers. So these are deployments in the AWS data center.
Interviewer
Yeah. It's fascinating, right? Like talking about this industry is how it seems that flexibility is so important. There's just like delivering a solution like you need chips, you get chips, you need data centers. Everybody's buying from different suppliers to reduce dependency.
Andrew Feldman
I think that's one of the reasons why we, we went from being a traditional chip consistent provider to also offering data centers. Is that what our customers want are fast tokens and anything we can do to make the delivery of fast tokens easier. For some of them that's in their data center. For some of them it's with an API. Just point your traffic to us and we'll point the fire hose of a fast token back.
Interviewer
Yeah. And just to get a sense for where you start and where you stop is you or do not or at least currently provide the cloud version. So if I want to run Kimi like I know you have like incredible stats for QE and Gemma in terms of speed, feel free to mention them. But you don't run those as a service or do you?
Andrew Feldman
We do.
Interviewer
You do. Okay, so you have a service also competing with the base tense and fireworks.
Andrew Feldman
Yeah, I think on we have an on demand service where you can come to our site and book a month. I think you can even buy buckets of tokens for Kimi or GLM or some of these models. Many of our customers come there, get excited about it and then move to a dedicated offering where they take hundreds of machines for a year or two or three or four once they've proven out the benefit for them in their work. Often they do a B tests. Not surprising. People like faster.
Interviewer
So just more like a testing.
Andrew Feldman
It's a full environment. You can go and use it. It's at cerebral AI play around.
Interviewer
Fascinating. But that could become like yet another big clock. Okay, so you have chips, you have data centers and you have a cloud business running inference on so it's going to go fasting thinking about moats. So you know famously Nvidia as cuda as well discussed mode. What's your equivalent of.
Andrew Feldman
Well, I don't think cuda is a modus Korea. We should talk about that. I think two years ago every state of the art model was trained in a CUDA flow and right now Gemini is trained without cuda Anthropic Claude is trained without CUDA OpenAI is trained with Kuda. So in a one or two year period they lost 70% share of training models or the state of the art Because Gemini is trained on TPUs it's
Interviewer
their
Andrew Feldman
anthropic is trained on cranium. Got at your and so I think the story of the moat is still present where the data show the mode is clearly shrinking. There's no mode in inference. It takes you eight keystrokes to move from a GPU to us in the cloud and it strokes oh eight feet. Keystrokes.
Interviewer
Eight keystrokes.
Andrew Feldman
That's it. To move your traffic from GPU API to us. And so obviously Pluto Was sort of enormously important in the creation of our industry and didn't allow the graphics processing here to be more general than graphics processing. But since 2324 Quick Live, I think its ability to solve the durable moat shrunk substantially.
Interviewer
And you have a whole ecosystem strategy as you think about your mode, to the extent that any mode can be, for example, in this industry, you build a whole ecosystem. Is that.
Andrew Feldman
Yeah, yeah, we built an ecosystem. I think our moat comes from the fact that by virtue of our architecture, we are doing things no one else can do. And it's not that they can spend more money or they can't pay more for this wookie pews. If you want Trask, you can't. It doesn't work. And so that's where we're building sort of our strength and that's how we're delivering value to customers.
Interviewer
How do you think about supply chain? We mentioned supply chain constraints for others, but what are your supply chain constraints? Are you old tsmc?
Andrew Feldman
Your tsmc. We have very close collaboration. They were investors of us. They'd been exceptional partners. I'll tell you an unusual story. In 2017, we showed up and we were about 30 guys total and we showed up in August. Horrible trend in Taiwan. I don't go to Taiwan in August. Brutal. Not that it's so nice here today. It's only 90 years you to where
Interviewer
you're embarrassed to me.
Andrew Feldman
Yeah, but all the humidity and. Yeah, you got put on a suit and we met with the leadership of tsmc. We said we would like a little pipsqueak company. We believe we can solve a problem that nobody solved in history. And here's how we would modify the way you make chips to make this possible. They thought about it and they said we agree, let's do it. In the meeting.
Interviewer
In the meeting.
Andrew Feldman
In the meeting. It wasn't go away for a month in the meeting was that they had
Interviewer
a prepared mind or they were just exceptionally fast on their feet.
Andrew Feldman
First, the salesperson had gathered the decision makers. Second, we are. Our proposal was sort of really good at allowing them to use what they were good at. And it didn't require them to change a huge amount, but it did require them to make real changes. And I think they saw this as sufficiently bold that they would learn as they did it. And they also knew that AI was better on big chips. And so the combination of fair mind, willingness to take some risk, bold thinking from a very large company. Fascinating. It is fascinating. I mean that's how big companies win, right? And how rare is that was extraordinary.
Interviewer
And what happened next, like how long does it take between a decision in a meeting like that, which senses exceptionally fast to two years. Two years.
Andrew Feldman
Chip making is a long hard process and most of the time your first chip isn't a winner. And there are lots of startups now, some of them with really smart guys like their first chip will not be order the tpu. Google had some of the best guys in the industry. First chip wasn't to win it or the second, nor the third. Fourth was really good. Now they're on their eighth and it's really good chip. Hey, the Anapuna team at aws first chip wasn't great. Second, third chip really good. It takes time. And so we built a chip, we delivered it. And this gets to an earlier question you asked. We solved a problem that nobody in the history of compute had solved and we delivered it in 2020 and nobody cared. No pitch cap, nobody bought any and nobody cared. You were like, oh God, nobody. Everybody said we were crazy, it would never work. It now it works and I know we want something. So then we built the next one. I mean the first one we probably sold quite a year fit and nobody
Interviewer
wanted it because the market was not ready or because the product was not good enough.
Andrew Feldman
Nobody wanted it because AI was a hobby. It was right. And who cares if your hobby's really fast? You care about fast when it's in production. You care about fast when you use it every day. And so we built another one and that one we sold three or five hundred. We built the third one and we sold tens of thousands.
Interviewer
Amazing. Yeah.
Andrew Feldman
Isn't that interesting?
Interviewer
And so going back to supply chain, do you need to think about on shoring diversification?
Andrew Feldman
So it's very hard to diversify away from tsm. Chips are so hard. And you actually when you design a chip, part of the design is for the rules of that factory, right? So you can't take your design from TSMC and go to somebody else because a huge amount of the work has been to be sure your design is within their rules. And so England I think only with one or two exceptions in history, each chip generation goes to one fat. So that's we're going to be with TSMC for our next generation as well. I think we have a supply chain that is built in many parts, but we bring the chips back from TSMC to the US repackage in the US and we assemble in the us we do our manufacturing in the US and then we ship from the US I Think when you're growing as fast as we are. Right. There are a full range of garden variety supply chain challenges. A vendor screws up a batch, it gets stuck in customs. The number of ways that things can go wrong. The supply chain is unbelievable. But we manage these every day and we're increasing our manufacturing throughput exponentially. And so it's really that part of the business. And call very well.
Interviewer
Incredible. So maybe to zoom out as a last question, what's your best guess about where all of this is going? Obviously who knows in AI in the next few years but in the next few year or two.
Andrew Feldman
Well, we know some things. We, we know that the model we use it you're using Today Fable or GPT5 6 will be the worst model you ever use. And whatever you think is cool about it right now is going to be boring and backwards in six plus. And that is so exciting. And I think I watched the way our young engineers use it and it's very different than the way I'm using it. And it's sort of a fun time where you can learn from your young team members. They're using AI very differently. I think the business of dashboarding and the business the AI doesn't the damage is doing the SaaS is I think unrepairable. For monies you could ask your AI build me a tool like Salesforce. 30 seconds later you have a working tool that is just unbelievable. And all the things that were difficult because they cut across your inside organizational silos. Right. One of the things that's really hard if you want to know for your top performing people when was the last time they got a stock option refresh and how much holding power is left so how much unvested stock they have left at today's price it was like five systems. You're in work days, you're in your stock, you're in your. Your car. You're in car. None of them can. And that's what a CEO wants.
Interviewer
Yeah.
Andrew Feldman
How much holding powers for my top guys.
Interviewer
Yeah.
Andrew Feldman
And I used to have little tools. I wrote for this and I've got a little app that I had and spark rights worms.
Interviewer
You go, well, what a story. It's just incredible to hear all of this from you. What a journey and what an exciting future. So thank you very much. I learned a lot and this was terrific. Thank you. Andrew.
Andrew Feldman
Thank you for having me on your show. I really appreciate it.
Matt Turk
Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review, or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
Podcast: The MAD Podcast with Matt Turck
Host: Matt Turck
Guest: Andrew Feldman, CEO of Cerebras
Date: July 23, 2026
This episode features a deep-dive conversation with Andrew Feldman, co-founder and CEO of Cerebras, the company behind the largest chip ever built—central to OpenAI’s infrastructure and the recent largest-ever semiconductor IPO. While Feldman has recently made headlines for Cerebras' $20 billion+ OpenAI deal and IPO, this discussion with host Matt Turck goes further, unpacking the technical, business, and personal sides of building specialized silicon for AI, industry bottlenecks, the shifting chip landscape, and the tenacity required to push deep tech innovation.
AI Inflection Point (2025): Feldman explains that as AI matured from novelty to productivity, "the minute you want to use it... speed matters." Faster tokens mean greater productivity and value.
“For AI work, big chips are undoubtedly the best way to go.” — Andrew Feldman (00:00)
Token Speed as Metric:
The key metric is "tokens per second per user"—determining the quality of UX and productivity. Real-time feedback is crucial, especially for agentic AI.
“The right metric is tokens per second per user. That’s how fast you get the first token, all the way through the last token...” (02:39)
Analogy to Netflix:
Just as broadband transformed Netflix from mailing DVDs to streaming studio content, fast inference is redefining what’s possible with AI.
“When the Internet became fast... they became a movie studio. Right. The speed enabled them to become something completely different.” (03:20)
"That was a good day. That was a great day." — Andrew Feldman (08:23)
OpenAI’s Approach:
OpenAI is “the best in the industry at looking at an exponential curve,” securing compute and memory capacity ahead of time.
“They struck big deals for memory, for compute with us, with others. They’ve really been sort of visionary in understanding what it means to extrapolate...” (09:34)
A Healthy Ecosystem:
Feldman argues the future will be multi-silicon—more vendors, more application-specific approaches, avoiding past concentration (e.g., x86 duopoly missing the mobile revolution).
Not a Bubble:
The demand for AI compute is not “bubble” behavior: “We’re all trying to catch up... overwhelmed with the demand for memory, which is a real weakness in the GPUs. It’s not a problem we face.” (17:27)
Three Major Bottlenecks:
“We solved a problem that nobody in the history of compute had solved, and we delivered it in 2020, and nobody cared.” (00:00/66:18)
2016 Origins:
Cerebras began with the insight that real AI acceleration required “unbelievably fast” memory (SRAM), packaged at wafer scale, not just squeezing more performance from commodity GPU architectures.
“You will never be a great company like Nvidia by doing something a little bit better... The right strategy would be to do something incredibly hard in engineering that was way better.” (30:26)
Wafer-Scale Innovation:
Building a chip out of a single wafer—“the size of a dinner plate”—was unprecedented. Overcame enormous challenges in power, cooling, and redundancy.
Failure and Triumph:
Cerebras spent 18 months and $8M/month crossing technical hurdles no one had anticipated, particularly around packaging the chip and making it reliable.
“What everybody thought was hard, we solved quickly. And what nobody else knew about... turned out to be really hard.” (34:35)
Emotional Milestones:
The joy wasn’t from ringing the IPO bell but from seeing the first server run successfully—a deeply personal win for the technical founders.
“There’s always this little thing in the back of your mind that says, maybe it’s shit, right? Maybe we’re actually crazy... And the flip side... is the joy that these are my cofounders ideas manifest in the world.” (37:12)
“The speed of moving weights to compute is about two and a half thousand times faster here than on a woven GPU. And so that’s the essence...” (44:13)
“If your compute is slow, that’s going to be more than an irritant. That could be crippling.” (48:30)
“We just announced that we were fastest in the world on one of Google's multimodal models.” (52:45)
“Their trainium part is doing the pre fill step and our chip is doing the decode... It uses their data centers.” (58:11)
“There’s no moat in inference. It takes you eight keystrokes to move from a GPU to us in the cloud.” (62:32)
“Whatever you think is cool about it right now is going to be boring and backwards in six plus. And that is so exciting.” (70:08)
Perfect Timing Myth:
“The way you have perfect timing is to have horrible timing for 10 years.” — Andrew Feldman (25:55)
Validation by Acquisition:
“The GPU architecture couldn't do, could not do fast inference and that this market was large and growing quickly and we were the fastest at it and the largest... So that was a good day. That was a great day.” (08:23)
Lonely in the Desert:
Feldman recounts 18 months of burning $8M/month with no product or sales, uncertain whether the dream would ever pan out.
“There’s always this little thing in the back of your mind that says, maybe it’s shit, right? Maybe we're actually crazy...” (37:12)
CUDA's Decline:
“There’s no moat in inference. It takes you eight keystrokes to move from a GPU to us in the cloud.” (62:32)
Industry’s Future:
“[Public markets] in the short term are voting mechanisms... but in the long term they're weighing mechanisms. Who's created the most value, the most weight.” (15:51)
On AI's SaaS Disruption:
“The damage AI is doing to SaaS is, I think, unrepairable.” (70:08)
This episode stands out for the clarity with which Feldman demystifies both the technical and market dynamics of the AI chip wars. The journey from deep tech “desert” to industry-defining innovation is candidly told, with takeaways not only for engineers and investors but for any listener eager to understand where AI infrastructure—and the entire software landscape—is heading.