
Hosted by Nathan Benaich (Air Street Capital) · EN

Poolside trained Laguna S 2.1 on 4,096 NVIDIA H200s and published the open weights sixty days later. The 118-billion-parameter mixture-of-experts model runs multi-hour autonomous coding sessions, handles up to a million tokens of context, and fits quantized on a single desktop-class machine.Nathan Benaich of Air Street Capital walks through the Model Factory that made that schedule possible, the system co-founders Jason Warner, GitHub’s former CTO, and Eiso Kant built to take a research idea to a production run in under a week. The episode covers the 409,000 training environments behind the model, the 50-minute session in which it built a browser rendering engine with no human intervention and no ability to see its own output, why the licenses attached to open weights differ so sharply between Poolside, Tencent and Moonshot, and what it means for a company to run a capable coding model inside its own security boundary rather than renting one through an API.Read the full piece: press.airstreet.com Poolside’s launch post: poolside.ai/blog/introducing-laguna-s-2-1 Evaluation trajectories: trajectories.poolside.ai RAAIS 2025 - Eiso Kant, Inside poolside’s path to AGI: press.airstreet.com/p/eiso-kant-poolside-ai-raais-2025From Air Street Press. Subscribe at press.airstreet.com, read the State of AI Report at stateof.ai, and leave a rating - it helps the show.

Description:Black Forest Labs has launched FLUX 3, a multimodal foundation model that learns jointly from images, video and audio within a single architecture, and mimic robotics has released FLUX-mimic, a video-action model built on that backbone and being tested on real assembly work in Audi's Production Lab. Nathan Benaich of Air Street Capital, an investor in Black Forest Labs, reads his Air Street Press essay on why a model trained to predict how scenes evolve turns out to be a usable robot controller. Covers the FLUX 3 preference results against Runway Gen-4.5, Grok Imagine Video, Kling v3 Pro, Seedance 2.0 and Gemini Omni Flash; how a lightweight action decoder reads the video prediction path without ever generating video; the frozen-backbone ablation against π0.5; the 101-millisecond system reaction time on a single RTX 5090; and what Air Street's own robotics deal flow says about where the constraint really sits.Chapters (estimated at ~150 wpm, slide proportionally against final audio):0:00 A robot arm in Audi's Production Lab0:40 What Black Forest Labs and mimic released1:20 Early access and open weights1:50 The preference-test results2:40 What a model must represent to predict video3:30 Reading actions off the video path4:20 The frozen-backbone ablation5:00 101 milliseconds, and the Audi tasks5:50 What we see in robotics deal flow6:30 The road to physical intelligenceLinks: FLUX 3 · FLUX-mimic (BFL) · FLUX-mimic (mimic) · Odyssey Series B · BFL Series B ·

Raia Hadsell, VP of Research at Google DeepMind, makes the case that intelligence is more than language: the same recipe that learns the patterns of text can learn the patterns of any complex system. She walks through DiffusionGemma and text diffusion, the Genie world models, and DeepMind's robotics stack, where world models now generate training data you cannot tell from the real thing. Recorded at RAAIS 2026.Timestamps0:00 Introduction (Nathan Benaich)0:35 From philosophy to DeepMind: the frontiers of intelligence2:37 The twenty-year lesson: one recipe for complex systems4:34 DiffusionGemma and the Gemma 4 open models5:32 How text diffusion works7:34 Speed, self-correction, and the sudoku test10:51 World models: better agents need better worlds12:56 Genie 1 to Genie 315:01 Genie 3 demos: typing a world into being18:20 World models for education19:38 Grounding Genie in Street View20:41 Robotics: a brain and a spine23:14 Gemini Robotics-ER 1.6 and Boston Dynamics' Spot24:42 The vision-language-action model25:45 The data bottleneck and closing the loop27:24 Beyond language: the domains still to crack

At RAAIS 2026, Nathan Benaich sits down with Hadrien Canter, co-founder and CEO of Alta Ares, the air-defense company building AI-guided interceptors to shoot down cheap drones and cruise missiles. They get into why Europe has lost air superiority for the first time in modern history, what the data loop looks like when you run it on a freezing front line instead of a laptop, why "quantity is the quality" in the industrialization race, and how the talent pool is shifting toward European defense. Recorded live at RAAIS 2026 in London.Chapters00:00 — Introducing Hadrien Canter and Alta Ares00:47 — 2022 in Ukraine, and how Europe lost air superiority03:23 — The Series A and the Airbus partnership05:13 — The data loop, edge AI, and the three phases of a mission08:22 — What no simulation can reproduce10:09 — Hiring for the mission; defense as the precondition for peace12:36 — Open research questions and the human in the loop14:22 — How the adversary uses AI: evasive Shaheds, drone mesh, China17:27 — Two interceptors, and why quantity is the quality20:50 — Iron Dome math, budgets, and peace-time vs war-time23:41 — Q&A: keeping pace with a fast-changing front26:31 — Q&A: talent and the shift toward European defense29:30 — Q&A: re-arming without permanent war32:56 — Freedom doesn't come for free

World models are the bet that AI should learn the world by watching it and acting in it, not just by reading about it. At RAAIS 2026, Odyssey co-founder and CTO Jeff Hawke makes the case: a world model is a neural simulator - an interactive stream of pixels that runs in real time, models physics, and answers back.He walks through Odyssey's four research fronts - streaming interactive pixels (Odyssey-2), joint audio and video (Starchild-1), shared multiplayer worlds (Agora-1, demoed live as a fully generated game of GoldenEye), and PROWL, which sends a reinforcement-learning agent to find and fix a world model's own failures - and argues the field is at its GPT-2 moment: promising, but pre-ChatGPT, with the GPT-3-style commercial unlock still ahead.Recorded at the 10th Research and Applied AI Summit (RAAIS), London, June 2026.Timestamps00:00 Intro: Nathan on Odyssey and world models01:05 Jeff Hawke: from self-driving to world models01:40 The bet — a missing form of intelligence02:40 Why world models suddenly matter (the late-2025 flip)03:16 What a world model actually is (and isn't)04:45 The neural simulator06:34 Two principles: end-to-end learning and generality07:19 The "GPT-3 of world models" and four research themes08:46 Odyssey-2: streaming, interactive pixels10:33 Starchild-1: generating audio and video together13:03 Agora-1: multiplayer world models13:57 Live demo: the room plays GoldenEye16:20 PROWL: improving the model by breaking it18:39 Where Odyssey goes next19:55 Still the GPT-2 era21:30 Q&A: physics limits, safety, compute cost, merging with LLMs

Ted Moskovitz leads the Science of Scaling team at Anthropic, the group that works out how to turn compute into smarter models. In this RAAIS 2026 fireside with Air Street Capital's Nathan Benaich, he argues that frontier scaling has become an empirical science - a discipline for cutting uncertainty before spending the compute, not just buying more of it.They get into the honest measure of AI acceleration (it's the counterfactual, not the benchmark), why a bigger model can be cheaper than splitting a task across small ones, whether a model can have research taste, and why safety and capability turn out to be the same axis. Plus the highest-leverage AI work to do in 2026, and why Anthropic's London office no longer feels like a satellite.Recorded live at RAAIS 2026 in London.Timestamp:00:00 - Meet Ted Moskovitz and the Science of Scaling team00:45 - What "the science of scaling" actually means01:18 - Why scaling is a science, not an art02:55 - Big labs vs the new "neo labs"04:47 - How a research finding reaches the product06:44 - What neuroscience carries over to AI (and what doesn't)09:12 - "When AI builds itself" and the real measure of acceleration10:33 - Trust, bypass mode, and the latest model jumps11:42 - One big model vs many small ones13:13 - Can a model have research taste?15:39 - How safety research makes products better17:36 - Emergent misalignment and the alignment race19:14 - The highest-leverage AI work in 202620:21 - Inside Anthropic's London office21:34 - Audience Q&A

Vivek Natarajan, Research Lead for AI, science and medicine at Google DeepMind, on porting the self-play and search recipe behind AlphaGo into scientific and clinical reasoning. He walks through the AI co-scientist, which generates and debates hypotheses (one matched a decade of lab work in two days), and AMIE, a diagnostic dialogue system trained in simulation. Recorded at RAAIS 2026.Chapters:0:00 Welcome and introducing the AI co-scientist1:41 Origins: Med-PaLM and the leap to hypothesis generation5:10 System 1 versus System 2 thinking6:34 Borrowing from AlphaGo: self-play and search8:02 Generate, debate, evolve, and tournaments11:47 Testing in real labs: Imperial College and antimicrobial resistance13:29 Ten years in two days: Penadés reacts15:44 More breakthroughs: leukemia, liver fibrosis and vorinostat18:44 Plant immunity and protein design20:09 Democratizing medicine: from Med-PaLM benchmarks21:28 AMIE and the value of experience23:12 Diagnosis, empathy and augmenting doctors25:19 Real patients: the Beth Israel feasibility study27:21 The co-clinician and the new triad of care28:31 Audience Q&A

At RAAIS 2026, Google DeepMind's Roberta Raileanu lays out a recipe for superhuman scientific discovery: AI systems that make groundbreaking discoveries across domains faster than people can. She walks through three ingredients - reinforcement learning to discover solutions where progress can be measured, open-ended divergent search to find new problems rather than climb known ones, and meta-learning to speed up discovery on problems no one has posed yet. The through-line: we can search for anything we can measure, but we still cannot measure what makes a discovery good. The bottleneck isn't the search. It's the signal.Chapters:00:00 - Introduction00:51 - Defining superhuman scientific discovery01:40 - The state of play: real progress, real plateau06:58 - Ingredient one: discovery as reinforcement learning (Move 37, MLGym)12:51 - Ingredient two: open-ended search and why greatness cannot be planned18:30 - Rainbow Teaming: quality-diversity in practice21:07 - Ingredient three: meta-learning the process of discovery (DiscoBench)25:06 - The recipe, and the missing signal

Angelos Perivolaropoulos, a research engineer at ElevenLabs, on turning GPU scarcity into an inference-engineering problem: how to serve far more users on the same hardware, from batching to frontier architecture changes. Recorded at RAAIS 2026.00:00 Introduction: ElevenLabs and the GPU squeeze00:38 The question: how to scale when you can't add capacity01:11 About Angelos: Scribe, speech-to-text and text-to-speech01:56 GPU scarcity meets exponential demand02:44 What a token actually costs: compute vs memory bandwidth03:38 Prefill, decode and the KV cache05:53 Batching and continuous batching (1 → 15 users/GPU)08:37 FP8 quantization and quantize-aware training (→ 20)11:29 Speculative decoding and multi-token prediction (→ 28)15:13 Compressing the KV cache: TurboQuant and distillation (→ 70)17:27 Frontier architectures: MLA, linear attention, state-space (→ 140)20:39 Trade-offs: nothing is free22:03 Q&A: papers vs production, token subsidies, TTS evals

The fifth State of AI Compute Index, in collaboration with Zeta Alpha. After a soft 2025, open AI research citations rebounded in 2026 - and NVIDIA still appears in ~91% of them. But the bigger story has moved off the page: Hopper is now the live installed base, Blackwell is mostly still pipeline, and frontier labs have started buying compute by the gigawatt. Nathan walks through what changed, what didn't, and why "GPU count" is becoming the wrong question.Read the full piece and explore the live charts: https://www.stateof.ai/computeChapters(00:00) What's new in v5 - citations, infrastructure, and gigawatts(01:25) The breather was short: 2025 was a pause, not a rollover(03:15) NVIDIA at ~91%, and the challengers - AMD, Huawei, Apple, TPU(05:05) Inside NVIDIA: the handover from A100 to Hopper to Blackwell(06:55) Startup silicon fragments - Groq, Cerebras, and the NVIDIA deal(08:20) Hopper is the installed base: 460k deployed GPUs(09:50) Blackwell is mostly pipeline: 80% still announced(11:00) The demand side, measured in gigawatts(12:15) Looking ahead, and why a GPU order isn't a clusterLinks:Full index and charts: https://www.stateof.ai/computeState of AI Report: https://www.stateof.aiAir Street Press: https://press.airstreet.comIf you found this useful, rate State of AI with Nathan Benaich five stars and share it with someone building in AI infrastructure - it genuinely helps.