
Hosted by Teresa Torres · EN

Guests Matthias Kleverud - Co-Founder, Momental Charlotte Kleverud - Co-Founder, Momental What we cover in this episode: What "GitHub for product management" means: finding merge conflicts in strategy, not code The product chain: signals → learnings → decisions → principles, and how AI maps it Three trees that model an organization: the product tree (OKRs to epics), the wisdom tree (decisions and their reasoning), and the people/time tree How a document processing agent uses OODA-loop thinking to extract and connect context across documents Why traditional chunking and RAG breaks down at scale and what Momental does instead The origin story: building a team of AI agents in 2024, only to discover agents hit the same alignment problems as humans Starting in 2022 with DaVinci 002 and learning that the market wasn't ready for AI-assisted product thinking How conflicts are detected, auto-resolved, or escalated to humans with merge options Why metadata—who said it, when, and in what context—is critical to preventing hallucinations The self-improving agent: collecting user feedback weekly and rewriting its own prompts Moving from chat-first to UI-first to proactive agents as an AI product design pattern Design partner strategy and what's next for Momental's public launch Resources & Links Momental - GitHub for product management Spotify - Where both founders started their PM careers Claude Code - AI coding tool discussed in the conversation Perk episode on Just Now Possible - Referenced episode about eliminating shadow work Chapters 00:00 Meet The Founders 01:14 GitHub For PMs Explained 03:19 Strategy Merge Conflicts 06:49 Product Chain Model 09:49 Capturing Context Fast 12:17 Context Graph And RAG Limits 16:52 Origin Story Since 2022 20:01 From Agent Team To Foundation 25:26 Three Trees Of Context 28:42 Two Agents Secret Sauce 31:37 How Document Processing Works 34:55 Agents Ask Better Questions 35:41 Human In The Loop Context 36:38 Data Models And Graphs 39:21 Beyond Documents And Vectors 42:25 Specialized Tools Win 44:50 Quarterly Planning At Scale 49:38 Discovery Versus Vibe Coding 51:00 Tree Building And Conflicts 53:49 UI Over Chat Interfaces 56:01 Proactive Agents In Practice 58:00 Quality Evals And Feedback 01:00:22 Launch Plans And Mission 01:03:19 Eliminating Shadow Work 01:03:56 Closing Thanks

Guests Yuri Vela Tulopov -- Co-Founder and CEO, ShowMe Quique Gomez -- Co-Founder and Lead Product Engineering, ShowMe What we cover in this episode: How ShowMe builds AI digital workers that function as inbound sales reps The origin story: spotting the conversion gap at a previous company and realizing AI could fill it Why the first MVP was a voice agent with product videos and a simple RAG knowledge base Adding a realistic avatar via HeyGen and how it changed user engagement through better affordances Decomposing a single sales conversation into multiple specialized sub-agents (greetings, qualifying, pitching) The three agent types: conversation agents, evaluator agents, and creator agents How deterministic workflows manage the lead-to-close journey across days Building toward a smart orchestrator agent that breaks out of rigid workflow paths Ingesting sales transcripts and training materials to teach agents company-specific sales skills Customer-driven evaluation loops that start at 100% review and taper to ~5% over time Creating automated tests from customer feedback to prevent prompt regression Confidence scoring and frustration detection for real-time human handoff decisions Treating the agent as a coworker: onboarding via Slack, weekly reporting, CRM integration Future plans: self-serve PLG motion, smart orchestration, and expanding to customer success Resources & Links ShowMe - AI digital sales reps for inbound teams HeyGen - AI avatar platform used for ShowMe's video calls Chapters 00:00 Meet the Founders: Juri & Kike Introduce Show Me 00:45 What Show Me Builds: AI Sales Reps as Digital Coworkers 02:17 Why Inbound-First Sales Agents (and Not Outbound Spam) 03:51 Origin Story: The Website Conversion Problem That Sparked the Idea 08:11 MVP Launch: Voice + Video Product Demos in Two Weeks 10:45 Bootstrapping the Knowledge Base: Videos, Docs, Scraping & RAG 11:54 Beyond Demos: Multi-Stage Buyer Journeys, Follow-Ups & Orchestration 14:34 Building Trust: Avatars, Video-Call UX, and AI “Affordances” 20:18 Whiteboard Architecture: Agents, Workflows, and the Orchestrator Layer 29:43 Where Agents Run: Creators, Evaluators, and Breaking Deterministic Flows 32:18 Sales Is High-Stakes: Personalization vs. Hallucinations & Revenue Risk 33:35 Conversation Agent Evolution: From Q&A Bot to Guided Sales Discovery 34:10 Why One Agent Becomes Many: Decomposing Stages for Latency & Memory 36:36 Orchestrator + Tooling: Routing Between Greeting, Qualify, Pitch, Next Steps 38:46 Teaching Sales Skills: Generic Prompting vs Company-Specific Playbooks 42:52 Ingesting Real Calls & Onboarding Like a Teammate (Transcripts, Training Docs) 45:05 Real-Time Voice + Avatar Demos: Latency Tricks and Video Clip Libraries 47:33 Creator vs Evaluator Agents: Data Cleaning, Custom Fields, Sentiment & Confidence 49:15 Human Handoff Guardrails: When Confidence Drops or Users Get Frustrated 50:15 Proving Quality in Production: POCs, A/B Rollouts, Dashboards, and CRM Logging 53:21 Evals at Scale: Customer Feedback Loops, Regression Tests, and the 5% Review Set 58:19 What’s Next: Smarter Orchestration, Self-Serve Setup, and More Digital Workers 01:02:07 Closing Thoughts: Customer Insight as the Moat in a Fast-Moving AI World

Guests Mark Barbir – CEO, Earmark Sanden Gocka – Co-Founder, Earmark What we cover in this episode: How Earmark differs from generic AI notetakers by producing finished work, not just summaries The pivot from Apple Vision Pro presentation coaching to a web-based meeting assistant Running multiple agents in parallel during live meetings Template-based agents: Engineering Translator, Make Me Look Smart, Acronym Explainer Personas that simulate absent team members (security architect, legal, accessibility) Why ephemeral mode (no data storage) became a selling point for enterprise Reducing AI costs from $70/meeting to under $1 through prompt caching Why GPT 4.1 still beats newer models for prose quality in their use case The limits of vector search for analysis questions across meetings Building agentic search with multiple retrieval tools (RAG, BM25, metadata queries, bespoke summaries) Designing for product managers as the extreme user to solve for everyone Their vision for an AI chief of staff that goes beyond automating deliverables Resources & Links Earmark — Productivity suite where the work completes itself ProductPlan — Roadmapping tool where both founders previously worked Granola — AI notetaker mentioned for comparison Assembly AI — Speech-to-text service used by Earmark OpenAI API — LLM provider with prompt caching support Cursor — AI code editor with build integration in Earmark V0 by Vercel — AI prototyping tool with build integration in Earmark Chapters 00:00 Introduction to Earmark Founders 00:28 Background and Experience 01:05 What Does Earmark Do? 01:23 AI and Productivity 03:09 Comparing Earmark to Competitors 03:41 Earmark's Unique Features 05:53 Templates and Personas 10:06 Technical Details and Development 17:12 Early Product Versions and Challenges 28:44 Understanding Prompt Caching 29:49 Managing Multiple Tools and Costs 30:59 Optimizing Transcript Summarization 35:11 Challenges with Context and Reasoning Models 38:10 Innovative Search and Retrieval Techniques 44:06 Creating Actionable Artifacts from Meetings 48:30 Ensuring Quality and Managing Hallucinations 58:20 Future Vision for AI Chief of Staff

Guests Jennifer Deal – SVP of Product Development, Healio Casey Utley – Senior UX Designer, Healio Matthew Skepner – VP of Technology, Healio What we cover in this episode: Why physicians need AI at the point of care—and how they actually use it (hint: it's preparation, not bedside) The surprising discovery that physicians wanted help with patient communication and empathy, not just clinical answers Building a working prototype in a weekend with Cursor after starting with Figma mockups How Healio's RAG system combines lexical search, vector search, and semantic search across multiple trusted sources Why "just use PubMed" isn't simple—five different ways to access the same data, each with trade-offs Designing citations that physicians trust: subscripts, hover states, and progressive disclosure Serving contextual ads while the LLM processes queries—a practical monetization approach HIPAA compliance and input guardrails for masking personal health information Eight LLM judges for evals: safety, medical accuracy, faithfulness, relevancy, completeness, reasoning, clarity, and overall quality Why physician feedback trumps LLM-as-judge feedback in high-stakes medical contexts The role of the Healio Innovation Partners in ongoing discovery and validation Resources & Links Healio — Medical news, education, and clinical guidance for healthcare professionals PubMed — Database of biomedical literature Cursor — AI-powered code editor used to build the prototype Chapters 00:00 Introduction to Healio Team 01:00 Overview of Healio's Services 01:57 Introducing Healio AI 03:39 Addressing Physician Needs with AI 05:45 Building Trust in AI Solutions 13:56 Prototyping and Testing Healio AI 18:02 Refining the AI Product 21:48 Technical Architecture and Advertising Integration 25:16 Balancing Speed and Accuracy in AI Responses 26:30 Ensuring Credible and Trustworthy Content 27:41 Challenges in Data Integration and Web Crawling 29:00 Optimizing Search Strategies for Different Data Types 31:09 User Interface and Trust Building 34:31 Human Feedback and Continuous Improvement 35:41 Guardrails and Evaluations for Reliable AI 39:11 Experimenting with LLM as Judges 45:13 Future Directions and User-Centric Design

Guests Daniel Kappler — CPO (Product & Design), Tendos AI Matthias Hilscher — CTO (Engineering), Tendos AI Key Takeaways Start narrow to prove value: Tendos AI began with just radiators for one design partner before expanding to all building products Own the interface: building a web application (vs. integrating into legacy systems) gave them control over UX and the ability to iterate toward full automation Evaluate each agent, not just the chain: per-agent evals make debugging tractable and show exactly where performance changed Use review agents: a separate agent that checks work (like code review) catches errors before they reach humans Let customers pull you: customers asked Tendos to replace their CPQ software—strong signals of product-market fit Topics Covered The tendering chain in construction and why it's ripe for automation How domain expertise (CEO's construction background) helped identify and validate the opportunity Entity extraction from PDFs ranging from 1 page to 1,800+ pages Planning patterns in agentic systems—creating and updating plans based on findings How agents evaluate product fit against customer requirements Building custom tracing and observability tools for complex agent chains The path toward self-learning systems through human feedback loops Links & Resources Tendos AI Chapters 00:00 Introduction to Tendo and Key Roles 01:01 Understanding the Tendering Chain 02:26 Real-World Construction Analogy 03:34 Challenges in the Construction Industry 04:48 AI's Role in Tendo's Product 12:59 Early Prototypes and AI Integration 18:31 Expanding Product Capabilities 28:56 Customer Collaboration and Workflow Automation 33:15 Strategic Partnerships and Technical Groundwork 34:20 Focusing on Specific Customer Segments 36:03 Product Evolution and Current Capabilities 38:17 Technical Workflow and Automation 40:12 Evaluating and Matching Product Requests 47:00 Dynamic Agent Architecture 55:29 Quality Measures and Evaluation 01:02:59 Future Directions and Customer-Centric Development

Guests Elliot Little, Product Manager, Zero Gravity Dan St. Paul, Software Engineer, Zero Gravity What we cover in this episode Zero Gravity's mission: breaking down barriers to elite careers for disadvantaged UK students The "knowing-doing gap"—why students struggle to act even when they know what to do Why their first prototype (a job suitability summary) didn't create the "wow moment" they expected The decision to use text chat over voice input and why guided prompts beat empty text boxes Context management techniques: removing stale tool calls, summarizing history, exposing tools conditionally Using different models for different tasks (GPT-5 Nano for structured outputs, lighter models for quick replies) Safeguarding architecture: moderation endpoints plus external verification with Unitary Building a failure taxonomy through internal red team/green team exercises What's next: long-term memory management for multi-year student journeys Links & References Zero Gravity Unitary – AI-powered content moderation Blue Dot Impact AI Safety Course – free AI safety course Elliot recommended Chapters 00:00 Introduction to Dan and Elliot 00:45 Zero Gravity's Mission and Impact 02:14 Introducing the AI Career Co-Pilot 04:01 Challenges Faced by Disadvantaged Students 06:49 Zero Gravity's Mentorship Program 09:14 Building the AI Career Co-Pilot 12:01 Early Prototypes and User Feedback 17:05 Refining the AI Career Co-Pilot 37:36 Introduction to Career Co-Pilot 38:02 Current Student Interactions 40:22 Technical Deep Dive 42:14 Context Management Challenges 44:43 Tool Call Optimization 51:48 Safeguarding and Moderation 57:52 Evaluating AI Performance 01:04:09 Future Directions for Career Co-Pilot 01:07:52 Concluding Thoughts

Guests** Jack Taylor, Product Engineer, Gradient Labs Ibrahim Faruqi, AI Engineer, Gradient Labs In this episode The iceberg metaphor: why frontline support is only the tip of automation potential How three agent types (inbound, back office, outbound) coordinate on complex tasks like fraud disputes Natural language procedures that let subject matter experts train agents without engineering bottlenecks The "turn" architecture: state machines that orchestrate agent logic across async, multi-day conversations Skills as modular agent capabilities—and how they're scoped deterministically per turn Defining "done" for outbound agents when the customer isn't the one ending the conversation Guardrails as classification problems: balancing recall and precision for regulatory compliance Ask a Human: a tool call that brings humans into the loop for approvals or missing APIs Auto-eval pipelines that flag conversations for manual review and feed labeled datasets Links & References Gradient Labs Incident.io episode – Referenced in the conversation Chapters 00:00 Meet the Engineers: Jack and Ibrahim 00:39 The Role of Product Engineers in Tech 01:21 Introduction to Gradient Labs 02:11 The Three Pillars of Customer Support Automation 04:32 The Evolution and Growth of Gradient Labs 05:29 Building and Refining AI Agents 06:39 Outbound Agent: Addressing Customer Problems 09:12 Defining Success in Outbound Procedures 17:08 Ensuring Compliance and Guardrails 30:17 Understanding Agent Guardrails 31:54 Complexities of Natural Language Input 36:21 Skill Design and Management 39:53 Deterministic Skill Execution 41:54 Customer-Specific Guardrails 44:21 APIs and Customer Tools Integration 46:02 Ask A Human Tool 48:24 Guardrails as Classification Problems 57:12 Auto Eval System 59:12 Future of Multi-Agent Systems

Guests Chris O'Connor – CEO, Mowie Jessica Valenzuela – Co-Founder, Mowie What we cover in this episode How Mowie evolved from a concierge marketing service to an AI-powered platform The "document hierarchy" architecture: how Mowie builds and maintains context about each business Why they moved from structured schemas to loosely structured markdown for intermediate processing Using Simon Sinek's Golden Circle framework to validate early product-market fit How Mowie generates quarterly content calendars and weekly posts across email and social media The three mini-calendars: public events, business-specific events, and recommended campaigns Building traceability so customers can see which context documents influenced their content Using customer approvals, edits, and regeneration requests as lightweight evals Connecting marketing performance back to point-of-sale data for attribution What's next: deeper attribution, omnichannel expansion, and digital out-of-home displays Resources & Links Mowie AI — AI marketing platform for SMBs Simon Sinek's Golden Circle — The framework Mowie used for early validation Chapters 00:00 Introduction to Mowie AI 00:06 Meet the Founders: Chris and Jessica 00:45 Understanding the Target Customers 01:35 Challenges Faced by SMBs in Marketing 04:20 The Evolution of Mowie AI 05:40 How Mowie AI Works for Businesses 08:16 Onboarding and Data Collection 14:56 Content Strategy and Calendar Creation 22:29 Technical Challenges and Solutions 29:41 Iterative Development and Feedback 37:23 Automated Content Calendar Approval 39:19 Content Calendar Creation Process 40:14 Marketing Pillars and Campaigns 42:11 Transparency and Traceability in Recommendations 45:23 Generating Weekly and Quarterly Calendars 48:24 Customizing Campaigns for Specific Events 54:15 Evaluating Campaign Performance 58:48 Customer Feedback and Document Hierarchy

Guests Steven Payne, Product Manager, Perk Gabriel Stock, Senior Engineering Manager, Perk Philipe Steiff, Senior Software Engineer, Perk What we cover in this episode How Perk's team identified an AI use case by connecting prior experimentation with a real operational problem Why they chose Make.com for prototyping—and shipped to production without touching backend code The evolution from a single prompt to structured conversation stages (IVR handling, booking confirmation, payment request) How breaking up the agent's task dramatically improved reliability Building two eval systems: classification for success rates and LLM-as-judge for conversational behavior Why the team still listens to calls manually even with automated metrics The challenge of prompt engineering for voice: numbers, booking references, and text-to-speech markup Lessons learned from expanding to German (prompts in native language improve results) How this project uncovered other operational problems they didn't know existed Resources & Links Perk Make.com – No-code automation platform used for the prototype Twilio – Voice/telephony provider 11 Labs – Text-to-speech provider (used in early experiments) Chapters 00:00 Introduction to the Team 01:54 Understanding PERK's Mission 02:59 Challenges in Travel Booking 07:27 AI Solutions for Customer Care 09:52 Prototyping with AI and Voice 17:00 Implementing AI in Production 25:51 Learning Through Trial and Error 26:40 Prompting Challenges and Solutions 27:58 Iterating on Prompts and Evaluations 30:08 Scaling and Production Challenges 32:43 Advanced Evaluation Techniques 35:32 Real-World Applications and Success 49:07 Future Directions and Expansion 53:53 Conclusion and Team Reflections