
Hosted by Ravid Shwartz-Ziv & Allen Roush · EN
Two AI Researchers - Ravid Shwartz Ziv, and Allen Roush, discuss the latest trends, news, and research within Generative AI, LLMs, GPUs, and Cloud Systems.

Daphne Koller wrote the book that many of us learned probabilistic graphical models from, founded Coursera, and now runs insitro, which is trying to make drug discovery a machine-learning problem.We start with the bitter lesson. She agrees with most of it and then says where it stops working: biology doesn't have enough data, structure is how people understand anything, and making a drug is a question about an intervention that hasn't happened yet, not a pattern in data you already have.Most of the episode is about why drug discovery is hard. Ninety percent of drugs that reach the clinic fail, and mostly not because the molecule was bad. The molecule usually does what it was designed to do. It just turns out the thing it was designed to do had nothing to do with the disease. Only 22% of diseases have any approved drug at all, and she calls that an upper bound on what we understand, not a lower bound.She also gets into what agents are and aren't good for in a wet lab, why cells don't grow faster no matter how many GPUs you point at them, what it would take to have real foundation models for biology, and why almost all of biology is still out of distribution.Plus GLP-1s and what human data keeps teaching us, whether AI can make the kind of leap that turned a bacterial immune system into CRISPR, and what she'd build if she were starting Coursera today.Key TopicsThe impact of scaling and data in machine learningThe importance of structure and causality in AIChallenges in drug discovery and biological understandingThe role of foundation models in biologyEthical considerations in AI and biomedical researchChapters00:00 Introduction to Machine Learning and Drug Discovery02:00 The Bitter Lesson and Its Implications06:48 Challenges in Drug Design and Discovery11:48 Ethical Considerations in Human Research17:20 The Drug Discovery Pipeline Explained29:30 Integrating AI in Experimental Design35:38 The Role of Human Judgment in Drug Design37:14 Future of Drug Design: Efficiency vs. Automation39:37 Challenges in AI and Data Availability for Biology41:08 Foundation Models: Potential and Limitations43:39 Causality in Biological Data: Importance and Challenges45:18 Creativity vs. Understanding in Drug Design48:17 Balancing Investments in Data, Algorithms, and Experiments50:07 The Value of Simulations in Drug Discovery52:03 Mathematical Frameworks in Biology: Utility and Limitations54:14 The Future of Drug Discovery: Optimism and Innovations56:28 The Impact of Coursera on Education01:00:33 The Role of Universities in Lifelong Learning01:04:06 Connecting Dots: The Fun of Variety in Work01:05:46 Optimism for the Future of Drug DiscoveryMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

In this episode, Joseph Suarez from PufferAI explains why he thinks RL never had an algorithm problem, but it had a code problem. Every part of the standard RL stack was running about a thousand times slower than it should have been, and once that got fixed, problems that used to take months started getting solved in seconds on one GPU. We talk about what makes a simulator good for RL, why most of their sims run on CPU, what he wants to do with scientific simulation, and why he open sources all of it instead of writing papers. Key topicsTypes of RL and their applicationsChallenges in scaling reinforcement learningThe role of simulators and hardware in RLRL in gaming: from chess to complex games like NetHack and RuneScapeFuture directions: scientific simulation and biological modelingChapters00:00 - Introduction to RL and Puff AI01:50 - Different settings for RL: Games, Robots, Finance04:10 - RL in LM and other domains07:00 - Challenges and solutions in RL scaling09:55 - Building fast, efficient simulators15:10 - RL for scientific research and simulation19:57 - RL in complex games: NetHack, RuneScape, Dwarf Fortress29:55 - Future of RL: Scientific discovery and beyondResourcesPuff AI - Official Site - https://puffer.aiNetHack - https://www.nethack.org/RuneScape - https://www.runescape.com/Dwarf Fortress - http://www.bay12games.com/dwarves/OpenAI Gym - https://github.com/openai/gymMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

Florian Brand builds evals at Prime Intellect. The premise of the conversation is that writing a benchmark is the easy part now. Keeping the model from cheating it is the job, and it takes longer than the benchmark itself.We get into why he thinks you can't evaluate a model apart from the CLI it runs in, what happens to statistics when a single run costs five figures, and whether the feeling that a model just works can ever become a number.He also has a few stories about agents finding their way around the scoring that are worth hearing cold.Timeline00:13 Intro01:00 What evals are for04:05 Agentic benchmarks07:10 Kimi K2 and model diversity08:23 Long-horizon coding tasks10:29 Building a benchmark12:15 MirrorCode14:27 Rubrics and LLM judges16:30 The cost of expert labelers17:49 Long runs and variance19:44 Evaluating the harness24:29 Chinese labs building CLIs30:00 More reward hacking37:45 Tau-bench and economic tasks39:43 Benchmaxxing and GLM 5.245:15 Statistics and cost47:56 Frontier convergence52:04 Misuse in open and closed models55:35 Self-improvementMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.AboutThe Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

Most labs build language models by scraping the web and filtering afterward. Pierre-Carl Langlais runs it the other way around. At Pleias, the French-German lab he co-founded, the models are built from data he can actually account for, which in practice means open and public-domain sources plus a lot of synthetic data the lab generates itself. It sounds like a self-imposed handicap. It mostly isn't. One of their models is a 600 million parameter system that runs live inside the Paris subway's monitoring pipeline.We cover the SYNTH pretraining dataset and why he thinks "ethical data" has to mean more than copyright-free. He explains why barely 2% of their Common Corpus appears in typical web crawls, and why that gap is really a preservation problem. From there, he gets blunt about benchmark maxing and whether GLM really earns its Opus-class reputation. He also argues that the quiet move by closed labs to hide reasoning traces is mostly about claiming ownership of model outputs. He's skeptical of sovereign AI, and not shy about how Mistral drifted from frontier research toward French corporate consulting. We finish on NVIDIA's persona datasets and the odd idea of training on the conditions that produced a text rather than the text itself.Timeline(00:02) Welcome and introductions(00:49) Why synthetic data matters, and the SYNTH set(04:15) Three reasons to control your training data(07:18) What "ethical data" actually means(11:08) How Common Corpus got built, from Wikipedia to PDFs(16:35) Agentic harnesses and synthetic data(20:03) Evaluating data when you train on reasoning traces(25:27) General versus specialized pretraining(27:08) Benchmark maxing and the GLM question(31:51) Getting diversity in, and the NVIDIA personas(35:02) Hidden reasoning traces and the fight over model IP(38:17) Mid-training and the "It's All Training" thesis(41:47) Can small models actually compete(45:01) Cybersecurity and Europe's strategic gap(47:08) Do you need a big model to orchestrate the small ones(52:08) Sovereign AI and the limits of national champions(56:42) Scaling laws when you control the data(01:00:41) The NVIDIA persona datasets(01:04:52) What you actually do with synthetic personas(01:08:22) Closing thoughtsMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.AboutThe Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

Dhruv Batra spent years leading Embodied AI at Meta, training virtual robots to navigate photorealistic 3D scans of real buildings with pure reinforcement learning. Then he left to co-found Yutori and build agents for a very different environment: the web browser.In this episode, Dhruv explains why he sees these as the same problem. Web agents, in his framing, are robots that act in a browser (pixels in, actions out), and the web turns out to be just as messy an environment as the physical world.Along the way, we cover his definition of intelligence as "navigation in idea space," why robotics is lagging LLMs, the sim-to-real gap and why you can't fake friction coefficients, the teleoperation counterexample to the "it's a sensor problem" argument, and his provocative claim that under the current paradigm, we solved machine learning and didn't even realize it. He also makes the case for why the scaling hypothesis isn't falsifiable, why JEPA-style arguments deserve to be grappled with, how Yutori trains its Navigator models with RL on live websites, and what happens to the ad-supported web when agents, not eyeballs, do the browsing.Timeline00:01 — Intro00:54 — What embodied AI actually means06:47 — Intelligence as navigation in idea space13:26 — Habitat: training robots with pure RL, no maps20:04 — Why robotics is behind LLMs28:24 — Sim-to-real: what you can and can't fake33:34 — "We solved ML and nobody noticed"37:12 — Leaving Meta, founding Yutori43:21 — Web agents: screenshots in, actions out48:15 — Why the web won't rebuild itself for agents53:32 — Training Navigator: RL on live websites1:01:04 — Who pays for the web when agents browse?1:09:17 — What Yutori means, closing thoughtsMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.AboutThe Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

Ion Stoica has done what almost no academic ever does — repeatedly turned university research into billion-dollar companies. He co-founded Databricks (now valued at over $100 billion), Anyscale, Arena AI and Conviva, while his Berkeley lab produced the open source projects the entire AI industry runs on: Ray, vLLM, and SGLang.In this episode, we ask him how it's actually done. His answer is surprisingly unromantic: solve a problem people already care about, build an artifact good enough that they adopt it, and pay attention to the moment users start asking "who maintains this after the students graduate?" - that's when a project becomes a company. He's also insistent that the credit belongs to his students.From there, the conversation goes deep into what he's watching now: why the AI stack has become an order of magnitude more complex than the Hadoop/Spark era, why maximizing GPU utilization is "the name of the game" for any enterprise, and why coding agents will struggle with distributed systems long after they've mastered web apps. He shares a memorable reward-hacking story — a load balancer that maximized throughput by dropping requests — explains why the gap between open and closed models sits at about six months, and closes with his case for regulating AI by outcomes, not capabilities.Timeline00:00 — Introduction: welcoming Ion Stoica01:21 — The playbook: how research projects become companies05:22 — Will vLLM and SGLang stay open source?07:47 — The real bottleneck in the AI stack: complexity, not just hardware14:31 — Should algorithms follow infrastructure, or the other way around?16:13 — Can AI coding tools write distributed systems and GPU kernels?21:09 — Verifiers, harnesses, and the limits of outsourcing understanding25:41 — Reward hacking: the load balancer that dropped requests25:58 — How should enterprises consume GPUs? Utilization as the name of the game30:23 — GPU scarcity: will the compute crunch ever end?35:27 — Hyper-optimization and the risk of locking in today's architectures37:17 — Open vs. closed models: why every company wants to own the stack40:35 — The six-month gap, and the rising cost of training frontier models43:58 — Kimi, Qwen, and who's incentivized to keep open models alive45:39 — Regulation: outcomes, not capabilities47:41 — Self-regulation, concentration of power, and auditing open models48:32 — Wrap-upMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.AboutThe Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

Jean-Francois Puget is a Director and Distinguished Engineer at NVIDIA, where he leads the Kaggle Grandmasters team, and he's ranked third on Kaggle's all-time list. We caught him on the day NVIDIA announced Nemotron Ultra and its new agent skills repo. We talk about what skills actually are, why they beat MCP tools on context cost, and how NVIDIA built an evaluation pipeline to separate skills that help from skills that don't.From there we talk about the thing JFP cares about most: evaluation. He explains why most LLM benchmarks reward overfitting, how his team discovered O3 could pick the right files to fix SWE-bench issues without reading them, and why the only benchmarks he trusts are the ones where you commit before you see the score, which is exactly how Kaggle works. He predicts a "bloodbath" for the wave of competitors letting coding agents chase leaderboard scores with no notion of validation.We also get into what coding agents are actually good for ("a mix of a genius and a dumb person"), the multi-agent system at NVIDIA that built a working PyTorch clone that runs 10x slower than the real thing, his unfiltered take on frontier lab PR and the Mythos release, whether AI is a bubble, and the story of how his team won ARC-AGI with a 4-billion-parameter model at 20 cents a task, including jumping from third to first in the final hours of a seven-month competition.Timeline00:00 — Intro01:05 — NVIDIA's announcements: Nemotron Ultra and the agent skills repo07:21 — Skills vs MCP tools, and progressive disclosure10:24 — Agents that write their own skills: a new form of learning13:33 — When overfitting is fine (and when it isn't)15:47 — Why most LLM benchmarks reward overfitting17:06 — The SWE-bench contamination story: O3 picks files without reading them19:45 — How LLMs changed Kaggle, and the coming "bloodbath"25:40 — What makes a good data scientist: evaluation and one-bit experiments28:56 — Running Codex at scale: the top token consumers at NVIDIA29:37 — Did coding agents kill AutoML?30:16 — Genius and dumb at once: the limits of coding agents35:21 — Humans in the loop, sandboxing, and the teenage hacker who never wrote code37:42 — Mythos, frontier lab PR, and open source40:08 — Why NVIDIA builds open models, and where it's already frontier43:48 — World models, robots, and the coffee test49:20 — Why agents still can't play Dota50:24 — Is AI a bubble?53:14 — Winning ARC-AGI with a 4B model at 20 cents a task57:39 — Kaggle is a legal drugMusic:"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

In this episode, we talked with Dimitris Papailiopoulos, researcher at Microsoft Research's AI Frontiers lab and professor at the University of Wisconsin, about doing research in the age of agents. Dimitris told us about the Sunday morning that changed how he works: he handed Claude Code and Codex a question he'd been sitting on for years, went about his day, and came back to an answer. After a few days of dread about what's left for humans, he landed somewhere more optimistic, calling this the golden age of asking questions.We talked about his "smallest transformer that can add" leaderboard, a symbolic GSM8K solver built from if-else statements, and what happened when he put two Claude Code instances in the same file system and told them to do something cool (one pair invented a communication protocol, the other played Battleship). We also got into diversity and slop in agent-generated ideas, why agents get stubborn after a million tokens, harness overfitting on Terminal-Bench, continual learning and world models, whether agents need vision, and where information theory actually helps in AI and where it's a katana used to make coffee.Timeline00:00 Intro01:45 How agents changed the way Dimitris does research04:30 A Sunday morning with Claude Code, Codex, and GSM8K07:15 The dread, then the golden age of asking questions08:20 Taste and verification, and how we train students now09:53 Will models make human verification obsolete?11:30 The smallest transformer that can add 10-digit numbers13:40 Humans as initializers for gradient descent in idea space15:32 Allen on diversity, slop profiles, and high temperature research21:44 When Claudes meet: Battleship, invented protocols, and a grokking paper25:53 Single agent vs multi-agent under fixed compute30:28 Auto-research benchmarks and what agents actually accelerate35:14 Inside the symbolic GSM8K solver (with a live progress check)40:04 Idea overfitting and why agents refuse to change course44:00 Learning from failure traces and harness overfitting48:04 Continual learning, memory files, and world models51:30 Why don't labs personalize models on your own history?57:52 Agent-to-agent communication: is Jira the right tool?1:01:25 Multimodality: vision as a tool vs one unified model1:05:40 Information theory and AI, or making coffee with a katana1:11:23 Closing thoughts: ask bigger questionsMusic:"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

Phillip Isola, professor at MIT, joins us to talk about representation learning: what makes a representation good, why different models seem to converge on similar representations, and whether pre-training is really over.We discuss the platonic representation hypothesis and its limits, why clustering structure matters more than global geometry, and Phillip's new neural thickets paper arguing that post-training is easier than people think because pre-trained weights already sit near solutions to downstream tasks. Phillip also explains why he thinks LLMs are already world models, why he's betting on RNNs making a comeback, and why his most exciting current direction is artificial life: putting LLM agents in open environments with no fixed task and studying them like new organisms.Timeline:00:00 Intro song00:13 Intro01:05 What is representation learning and why it matters04:09 What makes a representation good: minimality and sufficiency10:03 How cross entropy and contrastive learning shape representations14:35 Dimensionality reduction and why dimension isn't the right complexity measure16:35 Compression and geometric clustering during training19:27 The platonic representation hypothesis and what actually converges22:53 Local neighborhoods vs global structure: the Aristotelian follow-up24:33 When convergence is strong: truth vs the space of possibility28:09 Is there true similarity in the world? The Bouba-Kiki effect30:56 World models vs autoregressive LLMs32:14 Diffusion LLMs as a special case of autoregressive models33:42 What architectures win in five years: the case for RNNs36:11 Grad student descent, or do we actually have principles?40:51 Feathers and wings: what to take from biology43:17 How close are we to brain-like models? Marr's three levels47:01 Are better models becoming less human-like?49:38 Is pre-training all you need? The neural thickets paper54:18 LoRA, low rank fine-tuning, and why post-training is easier than we thought56:01 RL environments and what our benchmarks actually test1:01:11 Artificial life: LLM agents as new organisms1:07:20 What's overlooked in AI research right now1:08:36 Why stay in academia, and doing science in the age of OpusMusic:"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

Most AI-for-science companies are selling shovels. Qichao Hu wants the gold.In this episode, we talk with Qichao, the founder and CEO of Molecular Universe, the AI-for-science platform that grew out of SES AI, a high-energy-density battery developer he's run for fourteen years. His core distinction is that companies from the AI world build tools, such as foundation models that predict properties, while companies from the science world care about the final product, such as the new battery or material that actually ships. Molecular Universe sits firmly on the science side, and the difference shows up everywhere from what they publish to what they refuse to.We get into the actual workflow of materials discovery and where AI compresses it. A single trial in a traditional lab can take a year with maybe a 40% success rate; the goal is to run a thousand candidates in parallel and turn that year into a week. Qichao walks through improving low-temperature fast-charging for EV batteries: from hypothesis generation through molecule-, material-, and device-level property prediction, down to autonomous labs that synthesize and test the top candidates without a human touching a pipette.The hardest problem, it turns out, isn't predicting molecular properties or measuring device performance, but it's the black box connecting the two. In batteries, that's the solid-electrolyte interface, which the field has been hand-waving about since the seventies. And the thing standing in the way of cracking it isn't a clever training trick but data: companies sitting on twenty years of records are finding it too messy, incomplete, and poorly labeled to train on, and are having to start collecting from scratch with new protocols and robots.Timeline00:13 — Intro and welcome;01:19 — Shovel vs. gold05:18 — Why the world's smartest scientist doesn't automatically give you a better battery07:25 — The discovery workflow09:37 — Exploration vs. exploitation11:54 — Safety and filtering: screening novel molecules against banned and toxic-substance lists17:55 — How hypotheses get generated, and where frontier LLMs help20:29 — From hypothesis to ~400 formulations: property prediction, ranking, and handing off to autonomous labs26:37 — "A foundation model for everything" — and the black box between molecular properties and device performance30:01 — World models and physics33:09 — The great unknown in batteries37:08 — Simulation vs. reality: calibrating massive simulated datasets with a sliver of experimental data41:47 — Lab robotics: how fast the hardware has caught up, and what a floor of autonomous labs looks like43:50 — The real bottlenecks50:21 — Pre-training from scratch vs. post-training LLMs, and why training tricks haven't reduced the need for good data52:42 — Evaluation55:42 — Publish the B+ model, keep the A model58:05 — Five years out1:00:37 — Closing thoughts and wrapMusic:"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.