
Hosted by Nova & Alloy · EN

Google rolls out three new Gemini tiers targeting speed, cost, and security. A judge signs off on Anthropic's $1.5B settlement for pirated training books, and Andrew Ng's OpenWorker delivers finished work instead of chat. Cisco ships tiny local models that spot known bugs in unknown code, while OpenAI's cyber-eval tool gets weaponized against Hugging Face. Plus XcodeBuildMCP 2.7.0, openagent 2.85.0 with computer-use and browser-use, and NTT DATA slashing incident analysis to 30 minutes with Codex. Show notes: https://tobyonfitnesstech.com/podcasts/episode-92/

A 27-billion-parameter open-weight model now fits in laptop RAM. Jack Dorsey launches Buzz, a chat, AI agents, and Git mashup. Anthropic's $1.5B pirated-books settlement clears final court approval. Semble's 98% leaner code search engine targets AI agents. XcodeBuildMCP 2.6.2 gives coding agents real Apple build access. Plus Gemini 3.6 Flash surfaces in listings, OpenAI and Hugging Face disclose a model-evaluation security incident, and Bristol Myers Squibb deploys Vera Rubin for drug discovery. Show notes: https://tobyonfitnesstech.com/podcasts/episode-91/

Cosmos 3 Edge from NVIDIA lands on Hugging Face, plus Unity's MCP v10.1.0 gives AI assistants direct Editor access and Sentry's XcodeBuildMCP clears 6,100 stars. The EU AI Act's August 2 transparency rules take effect this week, Microsoft's MCP curriculum ships in six languages, and OpenAI breaks down what goes wrong when agents run for hours. Also: Ternary on huggingface, whodb 0.121.0, codebase-memory-mcp, holaOS at 5,500 stars, and Upsonic 0.77.3. Show notes: https://tobyonfitnesstech.com/podcasts/episode-90/

Today's rundown: Moonshot ships Kimi K3 at 2.8 trillion parameters with steep pricing. Codebase Memory MCP reports a 99% drop in coding-agent token use. Prism-ML's Ternary-Bonsai-27B surges on the local-AI charts. Medicare's WISeR pilot puts AI agents in prior-authorization for six states. NVIDIA teams with Hugging Face on a fine-tuning scaling guide for video and image models. LongStraw and VideoChat3 push research boundaries. Plus the Agent Stack Release Readout covering Codex rust-v0.144.6, FastMCP 3.4.4, and Unity-MCP 10.1.0, and Microsoft's MCP for Beginners crossing 16,700 stars. Show notes: https://tobyonfitnesstech.com/podcasts/episode-89/

Today's AgentStack Daily covers OpenAI Codex rust-v0.144.5 and Claude Code CLI 2.1.205, plus Kimi K3 and Meta's Muse Spark 1.1 — both arriving on OpenRouter with million-token context windows. We look at a code-search tool claiming 98 percent token savings, clidey's whodb 0.121.0, Transformers 5.14.1 with two Inkling fixes, and a 2-bit 27B chat model trending online. NVIDIA Nemotron 3 Embed tops RTEB, FastMCP passes 26K stars, and OpenAI argues for a 'reverse federalism' AI safety approach. Show notes: https://tobyonfitnesstech.com/podcasts/episode-88/

GPT-5.6 Sol targets more complete agent output, ChatGPT consolidates work and coding, and Gemini 3.5 Flash operates screens and builds software. We also examine Bonsai 27B, offline Pixel AI, AMD’s 128GB desktops, JetBrains Copilot backend support, Anthropic’s risk-based model access, small-business results, Claude’s robotics boundary, government action in New York, Australia, and GOLD EAGLE, plus Google’s AI reconstruction of Pelé’s lost 1959 goal. Show notes: https://tobyonfitnesstech.com/podcasts/episode-87/

Today's AgentStack Daily covers three harness releases — OpenClaw v2026.7.1, OpenAI Codex rust-v0.144.4, and Claude Code CLI 2.1.202 — plus Kwaipilot joining OpenRouter. New agent research includes ABot-AgentOS for robot control, Amap's ABot-N1 for visual navigation, LightMem-Ego for wearable multimodal memory, and JobHop v2 for career trajectory reasoning. We also examine a multi-agent backdoor study, evidence-backed video QA from Salesforce, the MM-ToolSandBox visual grounding benchmark, Requential Coding's generalization bounds, and AdvancedMathBench for doctoral-level mathematical proofs. Show notes: https://tobyonfitnesstech.com/podcasts/episode-86/

OpenAI ships Codex rust-v0.144.3 and rust-v0.144.2; vLLM 0.25.0 promotes Model Runner V2 to default for dense models. Apple sues OpenAI over alleged trade secret theft by ex-employees. Plus Freya-TTS hits Turkish speech with a 183M flow-matching DiT, SAGEAgent cuts glioma diagnostic burden 55%, Agora moves from a router to an auction over reasoning steps, a two-agent system posts 0.402 on QANTA 2026, PAC-ACT trains Action Chunking Transformer policies with chunk-level RL, and Semantic Pareto-DQN addresses fraud collapse without resampling. Show notes: https://tobyonfitnesstech.com/podcasts/episode-85/

Today’s AgentStack Daily examines Codex 0.144, OpenAI’s GPT-5.6 Sol, Terra, and Luna lineup, and SpaceXAI’s Grok 4.5 release. It also covers GPT-Live’s simultaneous listening and speaking, Mistral’s 8B Robostral Navigate model, ChatGPT Work, Microsoft Flint, and new research on continuous-control memory, citation judging, coding evaluations, proactive agents, delegated web research, procedural code retrieval, and energy-market agent testing. Show notes: https://tobyonfitnesstech.com/podcasts/episode-84/

Today's AgentStack Daily: Hermes Agent v2026.7.7 ships, OpenAI Codex lands rust-v0.143.0, and Claude Code CLI releases 2.1.197. AionLabs ships Aion-3.0-Mini roleplay on OpenRouter. Kokoro runs high-fidelity TTS on low-power CPUs. Rowboat debuts on Show HN with 162 points as a Claude Desktop alternative. Security: GitHub AI agent prompt injection leaks private repositories. Plus early-failure probes for agent loops, Danus fact-graph memory, FreqDepthKV, DepthWeave-KV, RuBench 1.0, VAORA, and the Anthropic developer-relations API migration story. Show notes: https://tobyonfitnesstech.com/podcasts/episode-83/