
Loading summary
A
I do think there will be a general superintelligence, but I think that it won't be one lab that has built it, but it'll be kind of the plurality, like the collection of all intelligences will be a general superintelligence.
B
Welcome back to the MAD podcast. Today my guest is Misha Laskin. Misha is a former DeepMind researcher who left Google after shipping Gemini 1.5, convinced that LLMs had crossed a critical threshold and started Reflection AI, a startup that emerged recently out of stealth with $130 million in funding and a goal to build superintelligence autonomous systems. We talked about superintelligence a lot in this episode, what it means, how all building blocks to create it already exist today, and how coding and enterprise AI are the best paths to build it.
A
I think we have everything, the sort of series of breakthroughs that need to happen. I think many of them have happened. We're really focused on coding to start for a number of reasons. This is a super intelligence complete problem. Like if you really solve this Oracle organization, it's just for coding. You've basically built all the capabilities you need to have superintelligence.
B
We also chatted about Asimov, their brand new product announced today.
A
As we publish this podcast, we're launching our first product and the first milestone on the path to superintelligence. The best in class code research agent built for organizations. It has this new concept that I don't think has been introduced to date.
B
Finally, we covered what it's like building a super intelligent startup at a time when the big AI labs and max7 companies are poaching top AI talent left and right.
A
I think the entire makeup of the research team was earning a lot of money. At big labs we tend to attract people who have that kind of internal drive in them of wanting to be part of that next story.
B
Please enjoy this deeply insightful conversation with Misha. Hey Misha, welcome to the MAD podcast. Thanks for being here in person today.
A
Yeah, thanks Matt. Thanks so much for having me.
B
You are the co founder and CEO of Reflection AI, which is a super exciting startup that up until recently was pretty much in stealth. You sort of came out of stealth a few ago and you have some exciting product announcements that we're actually going to discuss today. Before we do that, tell us about the company itself.
A
We're a research and product company and AI that was founded by myself, my co founder Yanis. We are both longtime researchers at DeepMind and generally the makeup of the team is a lot of folks who did Large language model training and reinforcement learning, building reinforcement learning systems within companies like DeepMind, OpenAI and Anthropic. We came together to build this company that co designs, research and product to build superintelligence. I think that superintelligence can sound like a very abstract thing. And I think what makes it special doing this at a startup or a smaller place is that you can be very concrete in what that means to you. Because I think that generally when people think about superintelligence, they might think of something that becomes really good at mathematics, you know, like Olympiads or code Olympiads or competitive coding, things like this. But these things to us don't really. I mean they're impressive, but maybe not. So it's unclear how they're useful. Like what is the superintelligence doing that is useful for kind of an end user. But if you co design kind of a product and the research alongside it, you can be much more opinionated about where it is that you want to go. And after talking to a lot of organizations and enterprise customers after we started the company, we realized that probably the form factor of an organizational superintelligence, the thing that's really going to help organizations get a lot of stuff done, is probably going to be something like an oracle. Like an oracle that understands the entire organization extremely deeply, that can answer questions at the level of, let's say the senior most person on that team. Right. It kind of has that like a principal level engineer or like the senior most sales leader has like full context on the company and it'll probably be super intelligent also in the sense that it's doing all those things at once. Right. As opposed to having like these distinct, different functions. And so that is what we think an organizational superintelligence is going to look like. That it's going to be a system that really deeply comprehends the organization.
B
It's a great term, organizational superintelligence as opposed to the super abstract concept of superintelligence. Just to say for a second on the concept itself, superintelligence. Does the industry actually agree on what that means versus AGI? So, you know, we're recording this. Obviously the big news is our friend Zuckerberg showering top researchers with extraordinary amounts of money and poaching them from various organizations to create a super intelligence lab. And then this, you know, safe superintelligence as a startup. Do people agree on what that means or is that just more like something that directionally feels very impressive, but not everybody exactly know what that means.
A
You know, I would actually say that superintelligence in these contexts of the large lab context is actually just being used synonymously with what AGI used to be used for. There's a scene in Indiana Jones and the Raiders of the Lost Ark, the first scene where he goes and tries to steal a sculpture, and he kind of takes a sculpture and replaced it with the sandbag and hopes that nothing changes. Yeah. And I kind of think one of those.
B
That scene does not end well. If I remember correctly.
A
It doesn't end well. But, you know, that's.
B
But the analogy stops here.
A
Yeah, the analogy stops here unless we really figure out safety. But I think that what happened is that by many accounts, many people would argue that AGI hasn't been achieved, but now some people might argue that AGI has. And so it's unclear that this is a binary event that hasn't been achieved yet, even though that was the case before. And so I think the goalposts were just moved like, oh. What I meant by AGI was actually super intelligence. So I think that is to say that it's as vaguely defined as AGI was before. And I don't think anyone actually that there was any agreement on what we actually meant by AGI. Right.
B
I mean, it does sound cool, though. 2025 superintelligence is 2024 AGI. Okay. All right.
A
Yeah.
B
So coming back, I really like that term of organizational. Superintelligence does sound more tractable in your analogy of what a senior person would know, especially, I assume, a senior person that actually does the work. Right. Because typically the most senior people in an organization are actually disconnected from the reality of what happens in the troops.
A
Yeah. I think it's more. Yeah. The person who. On every team, there's a go to seasoned, experienced person who's very much in the weeds and understands everything. Getting systems first that are like that, and then they can hold a lot more context in their heads and sort of be even better than those critical kind of members of the team. I think that's kind of like an oracle for an organization. That's what a superintelligence looks like. Because once you deeply understand the problems that you need to solve in any discipline in enterprise, acting upon them to actually go and solve it is the easy part. And so we're really focused on coding to start for a number of reasons. We kind of think you solve this problem, it sort of solves the more general problem.
B
Yeah. What are those reasons?
A
There's one reason is that the way today we think of coding is just for Software engineers. But when you build a coding model, you've just trained a language model that can interact with a piece of software through code. And so the way these language models are going to interact with any piece of software, not just software engineering software like Salesforce and other CRMs and creative tools and so forth, the majority of those interactions are going to be through function calls, APIs. So through code. So today we think of coding model as something that is just for software engineers. But that's not true. If you build that system, it can actually interface with any piece of software. It's sort of what you've built is kind of the hands and legs of a digital AI, right? In the same way that a humanoid robotics robot can do anything with the same hands and legs that humans do.
B
So would coding be the brain and then everything else becomes the legs and the hands that are. Yeah, it's kind of connected to the brain. Or is that the wrong.
A
Yes, it's kind of like you can kind of think of the model is the brain and then what people call a scaffolding agent scaffolding. That's sort of the affordances. Right? The things that you can actually do. All these affordances for software, at least, you know, for digital intelligence, are going to be primarily through code. That is, the other option is teaching a model how to drag a mouse around, and it's called computer control. Some of that will happen, but I just don't think that that's going to be the majority in which a language model interacts with software. So if you solve coding, you've just solved how does a language model, how should it interact with software?
B
And also it's a more tractable problem. Is that fair? Because code is more structured and code is closer to grammar, lends itself well to the type of work LLMs do. Or is that.
A
No, I think that's very much correct. It's more that code is almost in. It's kind of intuitive to LLMs because they were trained on the Internet initially, and so they kind of have almost a muscle memory for code. It's kind of like humans have been evolved to have spatial geospatial muscle memory. And so we learn very quickly how to use our hands and so forth as toddlers. Code is actually very unintuitive to us. Right. That's why the GUI was invented, because the first way to interface with computers was through code. Very unintuitive for humans. So you invent the GUI for language models. It's the opposite. They've never seen any GUI data Really on the Internet, there's no mouse movement data out there. So the thing that is native to them is the exact opposite of what's native to us. So if you build things that are really good at coding, you're sort of just amplifying the existing sort of base knowledge of the LLM. It's just intuitive to a language model. So that's why coding. But then within coding you can kind of think about what it would take to create a super intelligent coding agent. You basically need a system that can generate code really well. And you need a system that comprehends large code bases and all the knowledge around them. Like everything that's in all the docs and the stuff that's kind of tribal knowledge that's in engineers heads and where I'd say every coding tool today is focused on the code generation piece in an autonomous, semi autonomous in your ide, in your terminal, but it's all around code generation. And as a result, because companies today I don't think are really focused in earnest on the understanding piece, what we're building towards I'd kind of call as almost today Maybe we have AI engineers that are at the level of an L4 engineer, maybe next year L5, L6 going to L9. I think the path there is clear and what we're going to get to if we don't solve the comprehension piece is basically L9 engineers with amnesia. So an L9 engineer, if they came into your company but had amnesia, they.
B
Wouldn'T be really good news and bad news, right?
A
It's kind of if you tell me exactly what to do, I'll go do it really well. But otherwise I know nothing about your code base, I know nothing about your organization and I'm not going to gather any of that either. The thing to solve now I think is building what we kind of think of as the context engine. Like what's the thing that's kind of brain that your coding agents can query in order to have all the context that they need to go do meaningful work. Not just junior engineering work, but the stuff that engineers spend really most of their time on, like critical infrastructure bugs, dealing with legacy code, things that if you actually hover over a shoulder of an engineer at any large organization, or even a startup with a really sizable code base, and you look at what they do, you'll find that a minority of their time is actually spent coding. 70% of their time they're digging basically through information, trying to understand stuff, asking other engineers questions. And you kind of want to build a Superintelligence that has that same DNA. What it's spending most of its time on is collecting information. And then obviously once it has it, it knows how to act on it.
B
How is that different or similar to basically doing a combination of coding and rag And RAG to bring in the context of the enterprise?
A
Yeah.
B
Is that the general direction?
A
It's kind of. Maybe it's almost like if RAG worked. So the purpose of RAG is to get all the information that your coding agent needs in order to do some work. But it's a very primitive form of doing this and by and large fails. It's kind of.
B
I mean, why is it primitive? Because it's just search oriented. It's just a query or it's more.
A
So there are a few things about it, but the first thing is that it's kind of what we call sparse because it's the stuff that it grabs from a large code base has a lot of false negatives and false positives. So it's kind of. And it typically only does it once. Right. It'll grab it and then that's all you have. And most likely for any meaningful query, it will not have given you the information that you need to actually go do the task. So RAG agents actually are pretty weak. What's happening now is that there's kind of a new kind of retrieval that I think people are calling agentic search, which is more what Claude code does. And it's kind of an agent that uses the same kind of command line that a human does and goes and kind of looks for files and uses the same commands that an engineer does to go look for files and search for keywords and grab stuff. And it does this agentically. So as opposed to it just being a one step process where it just pulls a bunch of embeddings in, it will use a file search tool and then it'll look at the stuff that it did and think about it and then use another tool and so forth.
B
And will remember it. Is there a concept of memory built in?
A
It does store the stuff in a context. But the way I would think about it is that imagine you are dropped into a large dark jungle and all you have is like a tiny flashlight. That's basically what agentic search is today, where you're going and you're kind of exploring and you have this tiny flashlight and you have to then remember everything in your head. And obviously that's a first form of comprehension, but it's also a pretty weak form of comprehension. It doesn't really. It doesn't scale to large jungles. Like if you have like a little tiny jungle in your backyard, then you might be able to navigate it.
B
A bonsai jungle.
A
A bonsai jungle, yeah. So that's kind of how I liken the existing state of the art agentic search. Whereas the kinds of systems that you need to build are ones that, well, basically expand, right. The kind of lens, like the aperture of your beam, so you're seeing more, are smarter about where they go and look for stuff, how they remember it. These are basically fundamental research problems like memory. Long context reasoning is what people call it. But I would say it's just expanding the aperture of your beam, figuring out how to actually store it. And also, what sources of information should you be pulling on other than just your code? Because information for organizations lives in their chats and their project management tools, and a lot of it is in their team's just brains, right. So oftentimes when that senior engineer leaves who understands some legacy piece of code, all of a sudden that becomes known as a haunted graveyard because no one else understands it. And so how do you build products that capture that knowledge that's in people's heads? That's why I kind of think that it's so hard to think about superintelligence in the abstract, because half of it is a product problem. The comprehension problem is sort of, I would say, deep. And what people used to say, like when a problem was AGI complete, it just meant that the problem has so much depth to it that if you solve it, you get AGI. I would say this is a super intelligence complete problem. Like if you really solve this oracle for organizations, just for coding, you've basically built all the capabilities you need to.
B
Have a superintelligence because you can generalize from there. It's an idea that I certainly the rounds, by the way, that coding is a path to superintelligence. I think a variation of that idea has been that it's intelligence building intelligence, doing it automatically, which I think feels like a different avenue. But what you're saying here, or maybe not, but I think. But what you're saying here is more that it's such a complex problem that if you solve that problem, then you're there. I mean, just to play back what you just said.
A
Yeah. I think that the notion of in the sense in which coding is ASI complete is that you can use an intelligent coding system to build another intelligent coding system. If people like to do research in the abstract, that sounds awesome because it's kind of Like I don't have to think about the actual problem that's being solved. If I just build an intelligence that builds an intelligence, it will figure it out somehow. I don't think that's how it works. I think coding intelligence that builds better coding intelligence will make algorithms more efficient, basically. And maybe it's so intelligent that it can even start making all the product decisions for you in terms of what questions you should be asking users and what features you should be building in. And to me it actually seems like a much larger, more practical problem that there's almost no, I think research without co designing product with it is sort of a meaningless pursuit now that that was only meaningful when the ingredients for how to build artificial general intelligence or ASI were not known. I think now they're known.
B
So you think that we have everything and maybe list what everything is in terms of components?
A
Yeah, I think we have everything. Even though the capabilities will continue being tightened. But the sort of series of breakthroughs that need to happen, many of them have happened. There are still some, but I would say that the ones that come to mind and there I'd say are four of them. The first ones are just making deep neural networks work. So that was just the Imagenet moment in 2012, when you can have a deep neural network that's classifying images at human and then superhuman level. I think that was the first breakthrough. Second breakthrough was reinforcement learning. Even though I think it was not clear maybe then, but now it's clear again. So reinforcement learning basically told us, how do we take a system and if we have a reward, make it super intelligent. Superintelligence has been built a few times now. AlphaGo was super intelligent. AlphaStar was. If you put more computer into it, it would have gone super intelligent. So we've known how to build narrow superintelligence for a while. So deep neural networks, reinforcement learning.
B
So 2012, AlphaGo was 2016 or 1916.
A
Yeah, 2016.
B
Okay.
A
Then scaling up transformers, the GPT series.
B
Of works, 2017, 18, kind of 18.
A
But then really I think started shining in like the early 2000 and twenties with GPT3. And that was both an architectural innovation and a data innovation. It was an architecture that could consume data on the Internet. So it's kind of both hand in hand. Transformers are not just architecture, it was also data. And then I would say the last ones are RLHF, which was, okay, we know how to align models in the basic way. We know how to do basic reinforcement learning on top of language models. And then Reinforcement learning coming back again with the reasoning models. So if I was to summarize, I would say neural networks, deep neural networks, transformers and Internet scale data, and reinforcement learning in kind of its various incarnations, these are the ingredients that are required to build a superintelligence.
B
And as a return of reinforcement learning at the end that you described, does that obviate the need for RLHF? Was that a temporary solution or are those parallel?
A
I think they're parallel capabilities because they do different things. RLHF was more the design, there was more to align a language model, like a pre trained language model. If you play with one of these base models before they're aligned, they are really useless. They are. They feel like stochastic parrots. They don't follow instructions. They're extremely kind of high entropy. And the fact that you could just align them, tweak them to align them with something that is human consumable, I think was a pretty big breakthrough and that was RLHF. Whereas RL with reasoning is more RL to drive kind of intelligence capabilities, which is make it really good at coding, make it really good at math, make it really good at whatever target domain that you have rewards for. And you use these in tandem where RL in kind of the reasoning phase, it sort of really expands on a capability and, and then RLHF aligns it to be kind of human consumable, but they're the same thing. In fact, I would say it's just the same thing. The machinery is the same.
B
So we have all the components for AGI. Sai, your goal is to create this in an organizational enterprise context. And I guess let's get into the big news of this week, which is that you're launching your first product. Tell us everything about it.
A
Yeah, we're launching our first product. Kind of the first milestone on the path to superintelligence. It's called Asimov, like the science fiction writer who had some thoughts on the subject. And this Asimov is the best in class code research agent built for organizations. So it's different than a coding agent, which is a thing that will go and write code for you. But that today we feel are pretty context poor. So you have to. If the task is well scoped, coding agents today will do it pretty well for you. If you are spending most of your time trying to understand some hairy problem in your engineering, why some infrastructure bug is happening, or really doing what engineers spend most their time doing, which is this kind of work. There's no tool today that really unblocks them in that way. And so there's, I'd say 70% of software engineering is actually, it's a code research job rather than a code monkey job. And we're building a product that addresses that problem. So the majority of time that engineers spend unpacking problems and trying to understand why something's happening and oftentimes it's bottlenecked because they have to ask someone a question and it takes a few hours for that person to get back because they're busy. We've built a product that helps overcome these kinds of problems.
B
So just to play it back like deep research for code.
A
It's very similar in the concept to a deep research for code, a thing that will go explore your code base and other sources of knowledge. It might take a bit longer than the traditional kind of snappy ask kind of products, but it'll come back with much better answers. Things that kind of make it different are that, well, one engineering knowledge does not just live in a code base. It lives in all sorts of other surface areas and other software products like project management tools, chats, documentation like things like this. And Asimov pulls from those things. So it's not just your code base, it's aggregating all these sorts of information. It has this new concept that I don't think has been introduced to date that we call team wide memories. So today products have individual memories which are remembering the personal preferences of the developer, but they don't have a team wide organizational understanding. Like suppose you have a senior software engineer that understands some microservice A there you are, not that engineer and you now you need to interact with their microservice and you don't have the context they do. So now with Asimov, it kind of organically captures that information as it happens in chats, but also engineers can just teach it directly.
B
That's super cool. And then that addresses the problem. You mentioned that I think at some point earlier, which is like if somebody leaves, then the institutional memory goes with them. So you have permanent organization one memory.
A
Exactly. It's kind of building a permanent organizational system of record for your engineering knowledge to start. A side note on that is something that's been interesting, is that some of the most excited users have been these senior, you know, senior staff kind of level engineers who are fielding people's questions all the time. And so what we've seen is that we, you know, we'll go to an organization and the first three weeks there'll be four kind of staff, senior staff level engineers who are just populating its knowledge to kind of. It's kind of the first time I've actually ever seen engineers excited about documentation effectively. That's kind of a unique thing. And then the final thing is really around agent design that enables you to kind of what I was calling earlier, increase the aperture of your beam so that it's kind of looking at much larger code bases than agents were able to before. This is still in that sense a work in progress and that long context reasoning is just a big fundamental problem and I think there's a lot of work to be done there. But I think this new agent design is a step in that direction directionally.
B
Why and how are they able to do that?
A
This is not unique to us. I think when I look at how state of the art agents are being designed, they're being designed very much in kind of as these multi agent systems. Agent design is effectively it's a big reasoning agent and I think that's pretty standard. But it dispatches small long context reasoning agents to go and kind of go search for different parts of relevant chunks of information in the code. And so I think that it's this kind of decoupling of a big reasoning agent, like maybe a smaller context with a bunch of these little retriever scout agents. It's a different design than what's happened to date though I'm sure that other companies will converge on it as well. You really want to design your agents for the problems that you're trying to solve. So if you want like a really snappy agent that's going to answer things immediately, then this is probably not the best design for it. Right. You might want something that does rag which is very fast or like a very basic search agent. Kind of like what Claude code or cursor might do, which is something that just uses the terminal and like the file system there and use the same commands. As an engineer, that's a lot snappier than sending a lot of these retrievers out. So I think what we'll start seeing is this is product co evolving with the problems that you're solving and there'll be different agent designs for different problems.
B
And is a sort of like snappy one shot agent nessily a bad thing if you direct them at small problems and then you try to put them together or does the sort of individual little agent needs to be smarter?
A
No, I think that it's a really interesting question and this is kind of one of the things that's interesting at research now as opposed to before Is that before it used to be just around models, but now a lot of the research is in your agent design. And so what you ask is kind of an open question. You could, there could be a hybrid system that kind of routes some queries to one agent, to one agentic system and then routes other queries to another agentic system. It could be that you figured out some elegant, simpler kind of multi agent system that can do both things. It's kind of an open research question. It's very exciting and in some sense it parallels a lot of the unspoken research that was happening at DeepMind and probably in OpenAI as well during the pre language models. For example, projects like OpenAI's Dota 5 or DeepMind's AlphaStar project, which trained these expert level agents to play pretty complex video games like starcraft and Dota. A big question there that I don't think that many people appreciated is how do you design the environment for your neural network to actually dispatch actions? So the most simple thing you can think of is well, it learns to use a keyboard and mouse like a human does. That turned out not to work. So when you read the AlphaStar paper, you see that they actually figured out this particular way of factoring out the actions in order to make neural networks draw StarCraft. And what that actually meant is that that was agent design. That's sort of. So now agent design is back in terms of designing scaffolding. Back then it used to be called environment design, but it's the same thing. And a lot of the project was in these big projects was not even on training the models. It was figuring out how the agents should actually interface with the environment that you're training it in and where do you get your data from, what's your data, how do you collect it?
B
So on that point, so you had mentioned the three things like multi agent design. I think the first point was what sources of data it accessed. So is there, I think to what you just said, is there some difference in how those agents access data versus rag, which is pretty much like straight up search, I guess as a side question or related question, do those data sources need to have a special protocol to lend themselves to agents or like MCP style kind of infra so that the super intelligent agents that go around and sort of grab information everywhere can interact with them?
A
It's a, it's a really good question. So there are kind of a couple, couple interesting things to unpack there. The first thing I guess I would say is that kind of going back to some Depending on the problem you want to solve, like different types of search. These are all these all fall into basically search. And maybe I would say that if you want really fast, but it doesn't really matter how accurate it is or you know, just needs to be some kind of of ballpark accurate, but really fast. So kind of the weakest form of search. RAG is great. Then there's this more kind of agentic search we spoke about that the agent uses tools similar to the ones available to humans on a computer. And finally, which is kind of in between the spectrum of fast and you know, so it's slower than rag, but it gives you better answers. And then even slower is what I'd call neural retrieval, which is you have a really long context model and you ask that long context model to retrieve stuff for you. You feed it everything. You can maybe use multiple of them if everything doesn't fit in one, and then you ask that model to look at what you put in its context and retrieve the relevant stuff to you. That's kind of called neural retrieval. And that is going to take the longest. It's not guaranteed to be the best, but you can train it, you know, to be really good. And so that's kind of the spectrum of search capabilities that I see them today. The difference between something like this and then kind of MCP as it pertains to interacting with different sources of knowledge is that MCP is kind of like, it's stateless. It's just a way for you to interact with another piece of software. But what you need to do here is you actually need to collate an index of knowledge, right? You need to take data from software, store it somewhere and make it searchable. For an agent that ends up being a bit of, I would say a blind spot to a lot. You know, there's a question of why hasn't this been done. One of the reasons is that just from a like forget intelligence, just from a kind of business model perspective, it falls into a bit of a blind spot for existing coding tools which were meant as to be served as SaaS offerings for broad kind of consumer base. But an enterprise, this is tip. It will not want this leaving into SaaS. And so you have to kind of rebuild your entire business to basically deploy the stuff on that enterprise's resources. And so most companies have been basically staying away from this problem of indexing and integrations at this level of depth because it requires them to basically change their entire go to market and business model. But it's also why it's really easy to switch around between various different coding providers today because they don't really integrate deeply. And so, you know, you can try Cursor today, you can try Claud Code tomorrow you can switch to Windsurf. And there's really, because they're pretty thin skins, I guess on just the language model API from a developer perspective, it's pretty easy to switch around and try different things.
B
Interesting. So is a consequence of that that Asimov needs to be able to work in a sort of air gapped context. Virtual private clouds on Prem.
A
You know, we do have a SaaS offering because ultimately if you flip something into On Prem, like it has to start off as SaaS anyways. But the primary kind of benefit to enterprises is we're not going fully on Prem today, but we are doing vpc and that ends up being sufficient for a lot of big enterprises out there who already have their cloud infrastructure on AWS or Azure or gcp. This notion of it being deployable as VPC is extremely important to an organization. This is definitely. It's really a deal breaker.
B
You can't even start reinforcement learning part in Asimov. How does it manifest? I mean you guys are super world class RL specialists. So is that the whole idea? Like it keeps sort of learning and getting sharper with every interaction?
A
Let's say before language models were useful, you kind of had to be in this world where you build the best language model and then you figure out the product. That's kind of the world we were in. And Anthropic spent a few years building the language model and then took off with quad 3. OpenAI GPT2 is not really productizable. GPT3 not really either. And it wasn't really until 3.5 and 4 that they were able to productize it effectively. We're in a different world today where language models are pretty good. And so our strategy has been we build this kind of multi agent system. Some parts are we're training models for. We kind of see blind spots from third party models and other places a third party model stays for today. Over time we're going to kind of abstract all of it, but we're being a bit more strategic about which parts of the system you need to go after as a startup because that's kind of, you know, what you have as a benefit as a startup is that you can be a lot more focused. And the problem at hand, obviously the downside is that you have to be a lot more strategic about the bets that you're taking. You don't have the resources to go and train everything all at once. So you kind of have to take it one step at a time. And so in the long term, this is going to be a system that just learns end to end reinforcement learning. And in the short term, we're applying reinforcement learning to fix problems that we're seeing as kind of blind spots in the existing sets of models when we're deploying them.
B
I love pragmatic undertone to everything that you. You're saying, which is really interesting and dare I say, somewhat refreshing. I look at this different ways of building wonderful things in AI, but the pragmatic tone is really interesting and not that widespread. So that's the product that you just launched. Maybe take us back a little bit. So I alluded to your background as being world class reinforcement learning. How did that all come about? What was your journey to starting all of this?
A
As a kid I got pretty obsessed with physics and wanted to be a theoretical physicist.
B
As regular kids do.
A
As regular kids do. Well, you know, if you're a Russian Jewish kid, dropped in the middle of nowhere America, which is what happened to happen to me.
B
So let's go into that. So you're born in Russia.
A
Yeah.
B
Then immigrated to Israel as a kid.
A
Born in Russia, immigrated to Israel as a kid and then immigrated to the States for the second half of my childhood.
B
And is that a thing that like top people in AI do? Because like I can think of other people that were born in Russia and immigrated to Israel, then came to actually Canada.
A
Yeah, Canada is a big one.
B
Referring to Alia, I guess now at ssi, I think.
A
Well, there was, I mean the Soviet Union was an extremely, I would say kind of technically academic culture and a lot of basically the young scientists of, you know, when the Soviet Union fell apart, left to Israel, Germany, America, Canada. And when we arrived in the United States, it was actually.
B
How long were you in Israel for? Just as a, like a. I was.
A
There for eight years.
B
Eight years?
A
Yeah. So it was, you know, from one to nine was there and then, you know, was just born in St. Petersburg. So have no memories obviously of living there, Just visiting and then, yeah, arrived in the States.
B
So where was nowhere America that you mentioned, what city was it in?
A
Rural Washington State. So in Washington State. Washington State. Many people don't know this, but it has a hairline where the west side of the state is a lush forest and then the east side of the state is a desert. And you can see exactly where that starts. It's not a gradual transition it's just like trees, you know, a dense forest just starts somewhere and then it transitions to immediate desert. I lived on the desert side. There's a national lab there. My parents are chemists, and they got jobs in this national lab. Right. The culture of that town was. It was one of the sites during the Manhattan Project. It was called the Hanford site. It's where the plutonium was enriched. So it was the sister side to Los Alamos.
B
This is a whole vibe.
A
It is a vibe. You know, everything is themed around that event. The bowling alley is the atomic bowling alley. The brewery is the atomic brewery. The streets are like uranium, mercury, plutonium. You know, like the park there by the river, it's on the Columbia river, is called Leslie Groves park, who is like the ruthless general in charge of that project. So it's an intense town. Actually, the most intense thing is that the high school mascot, the town's called Richland, and the mascot, it's the Richland bombers. So the B52 bombers. And there are mushroom clouds on the basketball court. And that's it. It's an intense town.
B
So hence. Okay, all makes sense now. That would be the escape.
A
Yeah, I guess I was. Well, I didn't even think about the town's history as physics, even though that definitely is the case. But it was more that I was learning to speak a new language, had a lot of time on my hands, and my parents had their lecture books from. They bought this, the Feynman kind of series of lectures. And that was around. And I just spent some time reading it and just got into it. So that was the path to physics. Ended up going through and doing a PhD in theoretical physics and actually defected to artificial intelligence. I realized that first. What I saw. I saw Alphago come out. And the short of it is I realized that I had picked an interesting science, but not the science of our time. That was kind of it, you know, the. All the things I was learning in physics, all these kind of very interesting, great things were done basically 100 years ago, 60 to 100 years ago. And when you're studying something as a student, you don't really think too much about the timelines. You're like, oh, this is the cool. This is so cool. I want to do this. But then, you know, 60 years later, the field has crystallized in many ways, and it's not as dynamic, where the frontier is just moving so fast as AI. And so when I saw Alphago come out, to me, it just seemed like, oh, this is actually the science of our Time, and I needed to do that.
B
And as a quick segue on that note, it was super interesting to see the Nobel Prizes a few months ago, or maybe that was last year. At this point, everything converging towards AI. Do you think AI is eating all those other scientific fields?
A
I think it's augmenting. And in some sense, I think that part of what happened was that AI's impact in the world had clearly become hard to ignore. But there's no Nobel Prize for computer science, which is a field that. So, yeah, the Turing Awards have gone to AI for a number of years now. And so I think that the Nobel committee, I'd imagine, felt like it needed to somehow shoehorn this. And so obviously, there were very impactful AI breakthroughs with AlphaFold that resulted in a prize. But what's interesting is that the physics Nobel Prize was given to something that has not really had that much impact in physics. But I still buy it because it's kind of. There's a physics smell to the breakthroughs that led to these systems called Boltzmann machines and Hopfield networks that Geoff Hinton and Hopfield got the prize for. They're very physicsy. They look like the same exact objects that physicists study.
B
That's super. Just to play it back, like part of what you were saying is the Nobel Prize going to AI is less a function of AI sort of eating everything, but almost like a political thing at the Nobel Academy or whatever it is, that they should have had a prize for computer science and they don't have it. And now it looks kind of silly because AI is the fastest moving field in the world, therefore they're sort of retrofitting it.
A
Exactly. And I wouldn't even say that. I don't mean it in a negative way, I think. But, yeah, I don't think that as a physicist, when I look at, again, these objects, Boltzmann machines, Hopfield networks, very fundamental objects in the development of AI, even though they're not really even used today. But some concepts from them are. Those things haven't permeated physics in any way whatsoever. So I thought that was. But they just look like objects that a physicist would study, and they kind of mathematically look very, very similar.
B
Physics, E. Yeah. Okay. All right. So you evolved from physics to AI, and then what was the next step?
A
Had an interim where I started a kind of small startup that went through Y Combinator, and it was basically doing machine learning kind of prediction for inventory management. Really felt like I had to get on the frontier of AI research. I felt that that was going to be where a lot of scientific impact accumulates. And so I ended up joining UC Berkeley as a postdoc where I worked in this lab called the Peter Abbeel Lab, which is one of these great labs for reinforcement learning and what was called unsupervised learning research, which is basically large language models are probably the biggest. Large language models and diffusion models are the biggest kind of outputs of unsupervised learning. I didn't realize that at the time, but that was kind of a, in some sense like a. I wouldn't say like miracle year, but it was a special year in that lab given the people who are in it. A large portion of that lab ended up going to starting impactful companies or you know, being scientists, like doing very impactful work in large labs to give you a sense. First people I worked with was Arvind Srinivas, who now runs Perplexity, and Dennis Yaras, his co founder. We worked together on some papers there. Jonathan Ho, who is one of the inventors of diffusion models, was there and invented diffusion, like the big paper that made them break through there. And this guy named Jay Jain, who was on that paper and they started Ideogram and genmo, which are two startups in the kind of video gen and image gen space. Deepak Pathak, who is the founder of a company called Skilld, which is one of the kind of premier robotics companies, was there.
B
What year was this? What rough time period?
A
2020.
B
2020, yeah.
A
I mean, now look back at it, it was a pretty incredible group of people like Aditya Grover, who started Inception, which is a company that does diffusion models for coding. Yeah, it was just like another. I was just thinking, I saw this guy's tweet, his name is Kevin Liu the other day, who was in the lab and he was an undergrad then and most recently was leading a lot of the, or doing a lot of the kind of, of work for the mini models at OpenAI. So just kind of an incredible group of people at that time. Not obvious at all then. I mean, the research people were doing was very interesting, but would never have predicted that so many kind of companies would have come out of there.
B
And the next step after that was DeepMind.
A
And then I went to DeepMind. At the time I was really interested in Toronto. Yeah, Toronto and New York. So I joined the group of this researcher named Vladmini who was largely credited with starting the field of Deep rl. He was the first author of the Deep Q Networks paper, which was the Paper that got neural networks to play Atari. And his first set of papers actually largely defined deep reinforcement learning as a field. Then were basically DeepMind's claim to fame for a very long time. And so I joined his group to study the problem of what we called. I mean, the team we built together is called the general agents team. And so the whole point was to do research that enabled for us to figure out how do we build general agents. I think that it was much more opaque, I think then than now. And the big problem we were trying to solve is what people called and still do, but called unsupervised reinforcement learning, which is really how do you train reinforcement learning systems that are capable of assigning their own rewards if you don't have rewards without supervision, in the same way that you can give things some rewards. But kids and animals, when you look at them, they learn a lot in an unsupervised way. They interact with their environments. No one is telling them to. And so we were thinking about how do we teach reinforcement learning systems in this way? I actually think the subject is coming back in vogue in the age of language models with the question of how do you do pre training scale reinforcement learning? Like, how do you generate a lot of synthetic data if you don't have explicit rewards? I think it's actually a really interesting question now again, but that was kind of the general agenda of what I joined to study. And the short of what ended up happening was that language models started working. And once they started working, that really change, I think my entire perspective on what problems matter and what didn't. Because a lot of the problems that we thought were fundamental problems were solved in this kind of brute force way for us. And so you kind of had to recalibrate on what are the problems you want to solve. And so I joined a small project at the time that was tens of people. And that project became Gemini 1 and 1.5 and then obviously 2 and so forth. And I joined with my co founder, Yannis was leading the reinforcement learning team, the RLHF team. I joined his team and led a lot of the work for training reward models for Gemini and implementing the algorithms and so forth. And it was a very exciting time when I'd say a group of 10 to 20 people were basically those are all the people doing the RLHF work there.
B
And when was the decision to leave and start a company and what was the thinking?
A
Well, we shipped Gemini 1 and 1.5 and we realized that language models cross this threshold of utility where they're no longer research objects, they're going to be very useful. This was early 2024, realized that the ingredients were kind of in place to build a superintelligence. We felt that everything was there. There was one more piece to solve of going from RLHF to making reinforcement learning work. And that basically happened over the last year with reasoning models. So we felt that that would happen. Then the question was superintelligence for what? That was basically, you know, we felt that you can't answer this question in the abstract by kind of being a researcher that's really far away from product and customers. You kind of had to. You really had to go in and define what that means from a product vision and what problem you're trying to solve. Perspective that it's like we're not interested in building a super intelligence that will be super intelligence in mathematical Olympiads. And the difference between this era of reinforcement learning and the previous era of pre training is that when you did pre training, you made the models kind of generally better at everything. Reinforcement learning is much more jagged. It makes them good at what you want them to be good at. So just because you made them good at competitive code, that kind of improves the general coding capabilities, but that does not mean that you'll have not even a super intelligence, but even just a useful intelligence for software engineering code. And I think an example of that is I think Anthropic has done a really good job of building models that are meant for users of their products rather than benchmarks. When I look at academic benchmarks, the CLAUDE models are consistently worse than whatever else is out there by, you know, like oftentimes not even close. Like they're consistently worse, and yet from a user perspective, they're consistently better. Something has to explain that. And I think the explanation is that when you train large language models with reinforcement learning, they become jagged in the sense that they become good at what you want them to be good at. And there are some generalization capabilities, but they're much weaker than people think, which.
B
Is a little counter the narrative that you hear a lot, which is that generalization is always going to win. And I guess I don't know if that's true to it or bastardization thereof. But like the rich, sudden, bitter lesson. And so what you're saying is sort of not the opposite, but like that the solution is a combination of generalization and specialization. Is that fair?
A
Well, I think that it's like the bitter lesson is actually doesn't say anything about generalization. The Bitter lesson says that the systems that we should be thinking of and building are ones that scale well with search and compute. That's kind of it. And so what he's saying is if models are limited today, this was actually basically the lesson was for researchers, but I think it's for product builders as well. If you're building your product with the assumption that these models are going to stay at their current intelligence level and you make a bunch of hacks around your product to overcome those things, then in the next iteration of models, a lot of the hacks that you put in place will probably be bitter lessened. And that was scientific. Researchers were seeing these kind of limitations of models and plugging them in with kind of temporary hacks that would make them better at certain benchmarks. So it's kind of the same lesson translates. There's. But the lesson is more to build systems that kind of are good at soaking up compute and scale well with search.
B
Yeah, we'll put that in the show. Notes. I think that's probably one of the most often misquoted blog posts in the history of AI.
A
Well, I think what's happening with when we think of generalization, what's really happening is that if your training distribution is everything, then your test distribution just falls in your training distribution and you have generalization. Maybe one point of view that I don't think that that many people share it, but I do think there will be a general superintelligence. But I think that it won't be one lab that has built it, but it'll be kind of the plurality. Like the collection of all intelligences will be a general superintelligence. Because if you think of pre training as kind of the soil or the substrate from which you can now grow superintelligence in various categories, medical superintelligence, organizational superintelligence, superintelligence for scientists and math, these are all different types of superintelligence. Some of them you can merge into one model. But I think that there'll be, you know, if you think of different superintelligent plants growing from the substrate, yes, like a frontier lab will be able to capture some of them. But I think that there will be new frontier labs built that build out sort of, you know, kind of grow other plants and that the collection of this garden is going to be a general superintelligence rather than one company going in and growing all the plants. I think from a research and compute perspective, that's possible. But from a product perspective, if you want to build organizational superintelligence you have to go and integrate with all these customers and you have to have solution engineers that support them and you have to have salespeople that support them. And in this kind of rosy picture of a researcher who just kind of trains models and hopes to model as a super intelligence, I just don't think that the world will play out that way because it's meaningless if it's not coupled to a product and its deployment.
B
Double clicking on the product today, what is the reality of something like an autonomous coding agent? How good is the state of the art right now versus what hopefully it will be in the future? Are we in the teens in terms of SWE benchmark? Are we higher than that for certain tasks? Where does it all land currently on the benchmarks?
A
This is kind of a quick side comment that I think that sui bench is either. I mean, I would call it. I mean it's going to get, it's getting to saturation. And it's funny that even though you have like These numbers, like 70% or something like that on sweebench, those coding models are good, but they're not solving 70% of engineering tasks. So there's a sort of benchmark to real world problem misalignment, which is always going to be the case when your benchmark is not the actual thing that customers are using it for. But in terms of progress, I've been, if there's one thing, I had fairly aggressive timelines on progress in my mind. And I would say things have moved faster even than I would have expected. I think that we've gone from autocomplete engines to things that are kind of semi autonomous to now things that for junior tasks they can just do them autonomously. It's pretty incredible. There is, in some sense we are probably at a L4 kind of junior engineer level of autonomy, which is pretty incredible.
B
And autonomy means 100% success. No need for code review.
A
You still need to do code review in the way that you would do with an L4 engineer. But a level of reliability that, you know, the code review will, you know, there'll be some nits that you pick off like you do with a normal engineer, but that the thing they gave you is useful enough for you to review it in the first place, as opposed to this is just garbage. And why am I spending time reviewing this code? So I think for junior code, like, you know, little UI changes and small things kind of here and there, of which there are a lot. So it's very useful. We're probably at L4 as a field, I think that over the next couple of years that will just. The code generation ability will continue improving and we'll have things that are quite intelligent. I don't know where exactly I'd place them because I think they'll be really good at generating code when you give them all the specs, all the requirements, exactly what needs to be built. But the whole job of a staff level and above engineer is figuring out the requirements in the first place. So that's kind of 70% of their time is spent on that of designing stuff, figuring out what the requirements are, figuring out planning in advance. Like if I build this, will it conflict with these things? Soliciting information from other teammates. That's really what a staff engineer and above does. And then the implementation part is usually the sort of straightforward part. Okay. You spend like 20% of your time actually writing code and implementing things. And so the way things are progressing, I think we'll have very capable agents at the implementation level once you give them something to do that's really concrete and specific. But if you don't solve this kind of contextual gathering or, you know, build out the contextual core, I don't think they'll be at, you know, at that staff level as a whole.
B
But the L9 with memory, not the L9 with amnesia that we were talking about earlier, feels like a tractable, near term kind of problem to solve from your perspective. Like we're well on our way there.
A
Yes. So I think if you, I think it's a very tractable a. It's a very hard, very tractable problem. And that the combination of this, you know, L9 with amnesia and the L9's context core will together, you know, that will become the principal level engineer in an AI engineer. And so I actually think that that's not too far away. That's, I would say in, you know, a couple of years away.
B
So when you say rewriting back to the beginning of the conversation that you, starting with a coding agent, then is the idea that a lot of those principles that you described can be horizontalized across the company. So the institutional memory, which I find a fascinating concept that you can plug in, that could be the coding institutional memory. But next it becomes the marketing the product, the HR institutional memory.
A
Yeah, that's right. It's kind of at that point, from a build out perspective, it's very similar. You're now on prem or in the VPC of an enterprise customer. You have a centralized kind of knowledge around their code and you've already centralized knowledge around other tools for them as well that are immediately adjacent to the next thing. If you're centralizing knowledge from jira, that's both engineer and product management adjacent. Right. So I think at that point it just becomes an bringing other tool or other kind of integrations in based on where you're seeing pull from the enterprise and then enabling the ability to act on the user's behalf when they want to. So instead of just being able to ask questions, enabling the actual agent to go and do stuff for them. So I think that, I mean, that's going to come sooner than later, right? It's in the coding space. I think once you have that contextual core, you can integrate it into existing coding products. Right? You sort of like, if we think about the other coding products as being these autonomous software engineers with aminesia, you can fill that gap for them. Obviously you can build out one of your own as well. But I think it ends up being kind of a notion of customer choice. Right? Like you want the overall solution to be best for the customer and you're as a company focused on what you believe is kind of fundamental building block of enabling superintelligence.
B
So maybe zooming out to close. I'm curious on a few thoughts about the reality of building an AI startup today. I guess from a talent perspective, to start with, as we were saying a few minutes earlier, in this weird moment where ridiculous amounts of money are being offered to talent to move from one company to the other, how does one recruit and keep talent in this environment when you're very impressive, obviously, but still a small startup?
A
I think the entire makeup of the research team was earning a lot of money at big labs, basically. Obviously Giannis and myself included. The thing I guess to remember is that a lot of people get into this field because they're scientists at heart or they're kind of builders at heart. And so there is the financial element, like you certainly need to be. You need to pay enough for, you know, be generous enough, where that's not really like on top of minds, but people really care about discovering. Right. The next frontier and the next breakthrough. Right. The most exciting time to be in an AI lab is when it's before it's obviously the frontier lab. Like, I think that most exciting time at DeepMind was building Deep Q Networks and building AlphaGo. Those were the times, I think, where sort of like DeepMind, I'd say Golden Days. And similarly with OpenAI, was building out the GPT series of models, like at the 1 2, 3 stage. And I think for Anthropic it was like really in the 1 and 2 stage, Claude 1 and 2 stage, where obviously now they're kind of reaping the benefits of that breakthrough work that was done there. We tend to attract people who have that kind of, I would say, internal drive in them of wanting to be part of that next story because they're already at a big lab or they could join a big lab and they'll always be able to. I mean, always. I mean, we'll see when ASI comes around. But it's like that's not really that scarce of an opportunity. When you actually look at which startups are out there that have the clarity and the team and potential of starting a new frontier lab, they're not that many. So there are actually a lot more spots at the big labs than there are at startups that have a shot at this. So I think that people end up in a sense kind of self selecting like we win over candidates over OpenAI and anthropic meta DeepMind regularly. Obviously they get a lot more equity in this company as a percentage of its ownership. And if they kind of do back with a napkin math of if they had joined Anthropic at this stage and got that percentage ownership, what it would be worth now? It would be, I mean, absolutely generational. So that tends to be, I think if you don't have a good kernel of the initial team is not strong enough, then it becomes very hard. But if you have a very strong initial team and people see that potential for breakthroughs, then you've become like in a sense a scarce option because there are not that many places where you can do this. And you know, later this year we'll be shipping things that I don't think anyone ever thought a startup could do. Like, I think that we're going to be shipping some things on the research side that I think everyone thinks you need to be a giant lab with 100,000 GPUs to do. And I think it'll be quite interesting.
B
And surprising on the product front, again to the discussion about product versus research. You know, you guys are super deep, PhD type, world class AI researchers, but in a context precisely where you want to build product. Was it part of the core team to bring in people that would bring product or how did you think about it?
A
Yeah, we've hired out, I guess the company is, we kind of think about half product, half research. And so we've hired out a research team, we've hired out a Product team. And then there's, I'd say the majority of the makeup of the company is probably 2/3 of it is people who have research backgrounds at some of the big labs. Of those people, a bunch of them are kind of in this role that is between research and product. That for example, the design of the agent, design, research. That's a vary between research and product or evaluations. What are you evaluating your models to be good at? That typically is something that's just on research for us it's kind of a cross functional, end to end thing. The data that you're generating, the synthetic data that you're generating to train your models, that also cuts across all those things. So in a sense I think it attracts, well, maybe people similar to Yanis and myself that came into this and just wanted to be. We just did not want to maximize another academic benchmark. We just wanted to solve real problems and have of real evaluations. And so for those people, this ends up being a really good place I think for people who would much rather kind of sit in a known entity and really focus on some specific piece of work of like training the model because right. They're so big that when you enter them you're kind of given, you know, okay, like this is the sliver that you own. But that's really interesting nonetheless because you kind of become a craftsperson. So I think that's actually a very important skill to pick up. But if that's where people are in their lives, then yeah, I think the big lab is definitely a better option for them.
B
And you've raised a bunch already. I think like 125, 130 million, whatever the number in the press is. Does capital matter as much for a company like yours? Thinking about moats in AI today, there certainly is a well publicized capital raise for certain type of companies. Does your generation of startup and your type of startup require as much?
A
Capital matters a lot? I think the difference is that you can't operate at 100x less capital than a frontier lab, but you can operate at say 10x like an order of magnitude less capital when you're really focused. So I think that capital matters a lot and it really is determined by when are you ready to scale up your GPU count. That's basically it, right? I mean there's obviously headcount and data and so forth, but the primary cost for any of these companies is their GPU expenditures. So that's sort of. You kind of raise capital commensurate to when you're ready to scale to the next stage. And so I think capital matters a lot, but you can be a lot more efficient than traditionally Frontier Labs have been.
B
All right, well, we covered a bunch from the past to superintelligence to Asimov, to your background, to what you're building and what you're building next. I'm excited for this announcement. Asimov and then what you alluded to that's coming on the research side in the next few months. Thank you so much for being here. This was terrific. Thank you.
A
Yeah, thank you, Matt. This was fun.
B
Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
Episode: Ex-DeepMind Researcher Misha Laskin on Enterprise Super-Intelligence | Reflection AI
Date: July 17, 2025
Guest: Misha Laskin (Co-founder & CEO, Reflection AI, former DeepMind researcher)
Host: Matt Turck
This episode features an in-depth conversation with Misha Laskin, an accomplished AI researcher who left Google DeepMind after helping ship Gemini 1.5 to found Reflection AI. Laskin shares why coding is the most tractable path toward building "enterprise superintelligence," introduces the team's new product Asimov, and discusses the fundamental research, product co-design, and talent dynamics shaping Reflection AI. The exchange covers the history of AI technological breakthroughs, current bottlenecks in coding agents, and reflects on the practical challenges — both technical and organizational — of bringing superintelligent autonomous systems from labs into enterprises.
[21:14] — [25:11]
[26:46], [34:50], [62:46]
| Timestamp | Topic | |------------|------------------------------------------------------------------------------------------------| | 00:00 | Opening thoughts on superintelligence vs. AGI | | 06:23 | Defining organizational superintelligence | | 08:48 | Why coding is intuitive for LLMs; coding models as brain/hands/legs metaphor | | 10:50 | Limitation: “L9 engineer with amnesia” — current AI agent shortcomings | | 12:20 | The limitations of RAG and move toward agentic/neural retrieval | | 17:39 | Key breakthroughs and AI “ingredients” needed for superintelligence | | 21:14 | Launch and deep dive into Asimov: the code research agent for teams | | 25:11 | Multi-agent design in Asimov — product and research co-evolve | | 32:38 | Security and deployment: need for VPC/on-prem options in enterprises | | 34:50 | How Reflection applies RL pragmatically today | | 41:07 | Nobel Prize discussion—AI’s cross-pollination with sciences | | 47:28 | Laskin’s path from DeepMind to starting Reflection AI | | 53:32 | Autonomy of coding agents today, limitations & progress toward L9-level (senior) agents | | 57:01 | Generalizing institutional memory — coding to other company domains | | 59:32 | Recruiting top talent against Big Tech lab poaching | | 62:46 | Product/research team makeup and the value of startup focus | | 65:01 | Venture capital needs and capital efficiency vs. frontier labs | | 66:07 | Closing remarks |
The discussion balances the technical—sometimes almost academic—language of cutting-edge AI research with frank, pragmatic insights about product development, team building, and the messy reality of deploying AI in organizations. Both Matt and Misha show humor, humility, and candidness throughout, lending an approachable tone to otherwise advanced topics.
For anyone looking to understand the current state, challenges, and high-stakes ambitions at the intersection of AI research and enterprise productization, this is a must-listen conversation—with lessons for researchers, builders, and business leaders alike.