
Hosted by Jordan Nanos, Doug O'Laughlin · EN
Everything semiconductors and AI
Covering the spectrum

OpenAI and Anthropic are adding close to $30 billion of ARR per month, and the driver is gross margin expansion, not new compute. Jeremie Eliahou Ontiveros (@JeremieEO) and Reyk Knuhtsen (@robotknower) join Jordan Nanos (@JordanNanos) to trace the $100 million per megawatt per year figure from real inference workloads on GB200 and GB300 up through SemiAnalysis simulation. The Google deal prices GB300 capacity near $14 an hour against a $3 average, a premium justified by a 90 day cancellation clause and the ability to turn on megawatts immediately.From there the conversation moves to SpaceX's 10 gigawatt ambition, the sites and supply chain needed to hit it, the permitting playbook, and why Microsoft becomes the largest offtaker. The bull and bear cases both get airtime, followed by a Hugging Face security scare to close. Subscribe for weekly coverage of tokenomics, datacenter energy, and AI compute economics.CHAPTERS:0:00 – Intro0:47 – The $100M Thesis3:59 – Pricing & Google's Deal10:38 – Training vs Inference13:21 – Sites & Supply Chain20:55 – Permitting Playbook25:40 – Microsoft's Role33:31 – Paying For It41:01 – Bull vs Bear43:29 – Hugging Face Security Scare & Wrap-UpReferenced:SpaceX 10GW in 2027 – Why It's Real, Will Drive $300B ARR for SpaceX, and Why Microsoft Will Be the Largest Offtaker: https://newsletter.semianalysis.com/p/spacex-10gw-in-2027-why-its-real

Jeff Dean and the Gemini leads are leaving Google. Jon Y (@asianometry) makes his FIRST appearance to unpack the exodus with Doug O'Laughlin (@fabknowledge) and Jordan Nanos (@JordanNanos). The crew debates the Demis CEO question, whether Google acquired its innovation or invented it, and how the Bell Labs comparison reads once disruption arrives. First engineers who spent thirty-five years in one job now leaving for passion projects.The episode opens by revisiting the 2024 Dwarkesh clip where Jon questioned whether GPT-5 would be any good. Two years later the verdict is mixed: GPT-5 was a dud, 5.2 was ass, but 5.6 is strong and GPT-6, codenamed Doug, is rumored to write well. Then the Chinese transceiver ban, where the West depends on a supply chain China already owns. 00:00 Intro01:02 Was GPT-5 Good?05:00 The Transceiver Ban08:57 Everyone Leaves Google12:56 Google's L Culture19:20 Do Legends Matter?23:44 The Protestant Church28:44 Elon's $1T Pull-In30:55 Roll Your Own Software38:17 Start With Memory

Doug is back this week!Timestamps:00:00 Market Update07:02 Comparing to Past Bubbles in Taiwan and Korea10:08 Memory Prices, LTAs, and Market Cycles18:11 Future Demand for AI and Model Usage37:40 Scaling Laws and Supply Constraints44:45 Financial and Capital Constraints in Tech Expansion50:57 Geopolitical Risks, Policy Impact, Long Term Outlook

AI and datacenter CapEx hits $11 trillion cumulatively from 2024 to 2029, and $7.1 trillion of that needs funding, roughly 75% debt financed. Dan Nishball (@dnishball), Zane Fong (linkedin.com/in/zanefongzq), and Kang Wen Cheang (linkedin.com/in/cheangkangwen) sit with Jordan Nanos (@JordanNanos) to break down the AI project Trinity: capital, offtake, and data centers. Today only one deal reliably clears the lending bar, a five-year offtake from an investment-grade hyperscaler. Everything else struggles to finance.The crew explains how NVIDIA backstops are structured, how lenders price GPU loans against them, and what NVIDIA actually does with GPUs it takes back. AI debt financing is on track to become the second-largest US asset-backed market behind the $13 trillion mortgage market. Subscribe for weekly analysis on semiconductors and AI infrastructure.References: https://newsletter.semianalysis.com/p/nvidia-gpu-debt-backstop-unleashesCHAPTERS:00:00 Intro|01:23 The $11T Funding Problem09:12 Startups Can't Get GPUs11:32 How the Backstop Works20:26 The Central Bank of AI21:35 The Bullseye26:23 CoreWeave Credit Spreads36:58 APAC Examples42:41 ClusterMax48:58 Training vs Inference

Coding drives over 70% of lab API revenue, and token austerity policies mostly miss the point. Crystal (@crystalthegg), Max Kan (@maxkan), and Joey Brookhart (@SaasquatchC) break down why blocking teams from Opus saves nothing, while power users at the 99th percentile burn $100k per employee per year. Jordan (@Jordannanos) and the team run break-even math on the Max plans and Anthropic's margins.00:00 Intro00:53 Token Budgeting: Maxing vs Austerity03:05 Coding Eats the Token Market04:49 The ROI Question06:42 Subscriptions vs API Pricing09:08 Break-Even Math on the Max Plans10:42 Anthropic's Profit Margins12:07 Consumer vs Enterprise Mix16:04 The Two-Horse Race18:45 Codex App vs CLI20:41 Who Comes in Third?23:14 Clawbacks and the SpaceX Playbook26:59 Meta's NeoCloud Backstop29:58 Token-as-a-Service Market Forecast32:03 Hyperscalers vs Inference Startups38:50 MSL and the RL Scaling Law41:58 How to Build a Five-Figure RL Task47:10 Vibe Checks48:13 The $400B Anthropic BetReferenced:TokenBudgeting: Our Conversations with Enterprises on Token Spend: https://newsletter.semianalysis.com/p/tokenbudgeting-our-conversationsMeta Compute: Everyone Wants To Be A Neocloud: https://newsletter.semianalysis.com/p/meta-compute-everyone-wants-to-beAnthropic 3Q26 Profit Over $1B: The Anthropic IPO Financials Sneak Peak: https://newsletter.semianalysis.com/p/anthropic-3q26-profit-over-1b-theThe Future of Meta Superintelligence: A 1 Year Progress Update: https://newsletter.semianalysis.com/p/the-future-of-meta-superintelligenceComplete Launch Kit

A year ago, the big three was OpenAI, Anthropic, and Google. Things have changed.Moonshot's Kimi K3 sits above Gemini on every composite benchmark, and it's open source in 10 days.New episode: what K3 reveals about frontier margins, model sizes, and who's actually still in the game. 00:00 Intro00:11 Is Kimi K3 the Third Best Model?04:04 Why Delay the Weights?05:30 2.8T Parameters and Serving Constraints06:48 Frontier Margins and the 3x Price Hike11:10 New Architecture, What Comes Next14:09 Will Open Source Catch Closed?19:51 Built for Chinese Accelerators22:57 The Harness Is the Product28:49 We're Still Early

SMIC's N+3 node shrank the M0 layer over 15% and cut SRAM area 10 to 20%, all without EUV. SemiAnalysis built a lab just for that; Andrew Wagner and Afzal Ahmad walk Jordan Nanos (@JordanNanos) through the STEEL teardown of Huawei's Kirin 9030, from package to transistor. They explain die shots, FIB and TEM cross sections, NPU discovery, cell height, and standard cell libraries. Then backside power, GAA, and where SMIC goes next. 00:00 Intro: The STEEL Teardown Lab01:02 What Is a Teardown?02:17 Who Uses Teardown Data03:22 SMIC N+3 and the Kirin 903005:02 Inside the Lab: Sourcing to Silicon09:10 Die Shots Explained12:35 The NPU Discovery15:42 Scaling Without EUV17:57 FIB, SEM, and TEM Cross Sections20:59 Cell Height and Transistor Shrink23:56 Standard Cell Libraries26:22 Export Bans and Huawei's Response27:40 What's Next: Backside Power and GAA32:43 Data Center GPUs and Logic Folding34:28 Closing ThoughtsRead More: https://newsletter.semianalysis.com/p/steel-smic-n3-teardown

Bloomberg said half of 2026 US data center capacity is delayed. The SemiAnalysis Data Center, Energy, and Industrials team pulled the underlying report and found a broken denominator. Amazon alone built 4GW in 2025 and is adding 5GW plus in 2026. CoreWeave adds a gigawatt, all under construction. Jeremie Eliahou Ontiveros (@JeremieEO), Reyk Knuhtsen (@robotknower), and Ellie Holbrook join Jordan Nanos (@JordanNanos) as they walk through why the number is wrong and what the real forecast shows. The team covers behind the meter power generation reaching 40GW by 2028, Oracle's New Mexico problem, and the gas turbine supply chain that is running toward peak. They break down the three types of data center delays, how OEMs are responding, and where solar, batteries, and nuclear fit. 00:00 Intro00:51 The "half of capacity is canceled" myth04:43 Why early stage projects get canceled08:06 Three types of data center delays08:40 Oracle's New Mexico problem13:47 Behind the meter: 40GW by 202818:50 How OEMs are responding20:24 Signed deal to powered GPUs23:44 Hyperscaler market share27:49 How SemiAnalysis tracks data centers31:34 What behind the meter means33:31 The gas turbine supply chain38:22 Peak turbine42:54 Solar, batteries, and nuclear46:13 Final thoughts48:20 Favorite projectsReferenced:Stop Saying Half of 2026 US Datacenter Capacity Is Canceled: https://newsletter.semianalysis.com/p/stop-saying-half-of-2026-us-datacenterUS Grid Constraints: Towards 40GW+ of Behind-The-Meter Datacenter by 2028?: https://newsletter.semianalysis.com/p/us-grid-constraints-towards-40gw

DeepSeek V4 claims a 100x KVcache reduction versus a standard MoE model, hitting 1M context length through compressed sparse attention and heavily compressed attention. Kimbo (@Kimbochen), Cam Quilici (@noslawextratost), Bryan Shan join Jordan Nanos (@JordanNanos) to break down what changed from V3, why the new MHC dimension tripped up NVIDIA on day zero, and how Mega MoE fuses compute and communication into a single kernel. The vLLM versus SGLang NDA access gap and the Huawei day zero optimization guide circulating on Twitter.The crew walks through the InferenceX article on going from day zero to day 43 support and what that grind actually looks like across different hardware. Subscribe for weekly semiconductor and AI infrastructure analysis from the SemiAnalysis team.Referenced:DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time - GB300 NVL72, Huawei, MI355X, B200: https://newsletter.semianalysis.com/p/deepseekv4-16t-day-0-to-day-43-performanceChapters:(00:00) DeepSeek V4 vs V3 changes(01:00) Sparse attention and KV cache reduction(03:04) Day zero runtime support challenges(05:34) What Mega MoE actually is(08:38) Downsides of fusing kernels(10:25) MegaKernel benchmark claims(12:59) AMD FP4 optimization gains(15:14) Compounding step by step improvements(17:59) vLLM versus SGLang competition(19:34) Open source vs vendor libraries

Unitree is going public, boasting 67% gross margins on its humanoid robots. Jordan Nanos (@JordanNanos), Reyk Knuhtsen (@robotknower), and Niko Ciminelli discuss how the Chinese company achieves this through aggressive pricing, rapid iteration, and a focus on "good enough" hardware for the research and hobbyist markets. This strategy allows Unitree to dominate, much like DJI and BYD did in their respective fields.The discussion explores the reality of humanoid robot deployment versus market hype. While industrial applications are in their "baby days," Unitree's approach leverages economies of scale to create a significant moat, challenging US competitors to match their production volume and cost efficiency. The team analyzes if the US can truly compete with China's manufacturing might in the emerging robotics sector.Join SemiAnalysis Weekly for expert insights into the semiconductor, AI infrastructure, and robotics markets. Subscribe for deep dives into AI supply chain, chip economics, and market analysis.Article: https://newsletter.semianalysis.com/p/chinas-unitree-will-dominate-globalTimestamps:00:00 — Intro, deployment reality vs. hype02:39 — Unitree Business: why go public, margins, pricing05:48 — Parallels to DJI and BYD, economies of scale as China's moat16:52 — When is Claude Code moment for robotics18:38 — Real use cases and future demand shocks26:39 — Shenzhen and the humanoid BOM35:56 — The bear case: who actually buys them?42:17 — Can the US compete?