
Loading summary
A
Today on the AI Daily Brief, what 1200 professionals tell us about working with AI. And before that, in the headlines, Gemini 3 Deepthink is now available. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right friends, quick announcements before we dive in. First of all, thank you to today's sponsors, super intelligent Rovo robots and pencils and blitzy. To get an ad free version of the show, go to patreon.com aidaily brief or you can subscribe on Apple Podcasts. And of course, if you are interested in sponsoring the show, locking in those 2025 rates before they expire, send us a Note @ SponsorsIDailyBrief AI now, last note before we dive in, we're doing a bit of a switcheroo today. The headline section is actually a little bit longer than the main episode. There was just enough news that we kind of had to do it that way. So without any further ado, let's dive in. Welcome back to the AI Daily Brief Headlines edition. All the daily AI news you need in around five minutes. Although today is a very jam packed episode, so I expect it to be a little longer than normal. We kick off today with an exciting one for you model testers out there. Google has released Gemini 3 Deep Think Mode, which is their most powerful version of the new Gemini 3 suite. Now right now the new mode is exclusively available to subscribers of Google's AI Ultra Plan, which is their couple hundred dollar a month type of product. Now as you might imagine then, with the price tag that high, DeepThink is designed to tackle the most complex math, science and logic problems available. The mode builds on top of Gemini 2.5 deepthink, and as much as I tend not to care about benchmarks, does claim some impressive performances. They claim a state of the art 41% result on humanity's last exam without the use of tools, outperforming GPT5Pro at 30.7%. And DeepThink also achieved a 45% result on the Arc AGI2 test, more than doubling the performance of GPT5Pro to become the new state of the art. Now, it should be noted whenever we talk about ARC AGI that there are two vectors. There is score and there is cost per task. And while Gemini 3 Deepthink absolutely shattered the previous high score, it did so at a pretty elevated cost of $77 a task. Now this might go some way to explaining why they're paywalling deepthink mode behind the most expensive subscription. It's important to note that this is the first time normal users have ever had access to a model this expensive to run. OpenAI never released the preview version of O3 that cost $167 per task to achieve its state of the art performance at the end of last year. DeepThink achieves its state of the art performance by exploring multiple hypotheses at once before delivering a solution, a technique that has been used in research to boost performance, but generally hasn't been available to regular users as a standard feature due to the high inference costs. Now, one thing that's not exactly clear yet is what the use case is actually expected to be. Here they announced it by showing it generating a domino's game with complex physics in one shot. Another Googler showed it producing a complex physics simulation of a rubber vase falling on a hard surface, which I think helps clear up one thing because it is called deepthink and we already have Deep research. There may be some mental overlap between the two, but they are fundamentally different things. Deepthink is not just a souped up version of deep research. Instead it is capable of scientific reasoning. Now, in terms of first reactions, a lot of people had the same experience of Hyper Browser founder Sri Shrikhani, who wrote with all their TPUs and GPUs, how TF is Gemini 3 DeepThink overloaded and unusable? This is also the response I got for the first couple of hours after the announcement, although then it cleared up. Victor Talon writes, for those wondering and as expected, Gemini 3DeepThink solves the stack overflow bug that cost me a few days. The answer is more decisive than Opus 4.5, the only other public model to solve it. Even Gemini 3 Pro fails. It even points the exact location confidently takes forever. Though I don't have harder tests for now, most of my benchmarks are saturated now. I've only had access for less than a day so far as well, knowing that it sort of probably wasn't the use case, I still gave it a recent business strategy question that I had been both genuinely exploring, but also trying to test GPT5.1 thinking versus GPT 5.1 Pro versus Gemini 3 Pro. And I will say at this stage I don't particularly think that the extra reps of DeepThink are worth it for that type of business strategy question. Basically, I don't think that it particularly added anything more. In fact, I didn't even prefer its response relative to the others. So whereas I have recently been finding myself being being willing to take the time for five one Pro on business strategy questions, I think Deepthink might be a little too far and just not right sized for that particular purpose. In any case, I will continue to experiment with it making full use of that Ultra account. Next up, we stay in the Google universe where they have partnered with Replit to bring Vibe coding to the enterprise. The multi year partnership will see Replit expand their use of Google Cloud services, meaning a deeper integration of Google's AI models as well as using Google Cloud infrastructure on the backend to enable fully functional Vibe coded software. Apps coded in Replit will also be able to leverage Google Cloud Marketplace in their go to market strategy. Replit CEO Amjad Massad said, the goal for us and for Google is to make enterprise Vibe coding a thing. We want to show the world that these tools are actually going to transform businesses and how people work. Instead of people working in silos, designers only doing design, project managers only writing. Now anyone in the company can be entrepreneurial. Richard Serrata, the senior director for Google Cloud, added, it may feel like it, but replit is no overnight success. Amjad and team built something over time that became the exact right thing for this current moment with builders. In separate comments to CNBC on the state of the AI bubble, Amjad acknowledged that the honeymoon phase for Vibe coding is over. He said early on in the year there was the Vibe coding hype market where everyone's heard about Vibe coding, everyone wanted to go try it. The tools were not as good as they are today, so I think that burnt a lot of people. So there's a bit of a Vibe coding. I would say hype slowdown and a lot of companies that were making money are not making as much money. Amjad noted that earlier in the year we were getting weekly ARR updates from the vive coding companies and now we're not. That said, new statistics from Ramp suggest that Replit isn't slowing down all that much. The Ramp Economics Lab reported that replit is currently number one for new customer growth across all software vendors. Google is also up there following the release of Gemini 3 and Nano Banana Pro, sitting at number five for new customer growth and number two for new spend growth. Now one of my contrarian somehow takes is that I think we are actually way too bearish on Vibe coding right now. I tend to think that when we say Vibe coding we are having two entirely different conversations at the same time using the same words. There is Vibe coding for non technical people which is entirely different than Vibe coding for software engineers that shine coming off the rose type of phenomenon that Amjad was talking about is I think specific to Vibe coding for software engineers. There is a recalibration happening right now among developers around how best to deploy these tools around the autonomy spectrum. All these sort of questions around how you're going to integrate agent decoding into your processes in a way that doesn't just create new problems. However, for the non technical people, I think we are barely scratching the surface. In particular, I do not think that Vibe coding has significantly made its way into the business world yet. It's mostly still individual hackers and tinkerers who are discovering that they can build and modify their own websites now without having to use WIX or Squarespace or something like that. I genuinely believe that that is going to change and I actually think 26 is going to be a massive increase year for Vive coding, but with a very different market audience. One more Google adjacent story Google's neo cloud partner Fluidstack is in talks to raise $700 million at a $7 billion valuation. FluidStack started the year as a relative unknown, but signed multiple data center development deals to jumpstart their business. Google served as the backstop on a pair of deals pledging to repay debt if Fluidstack defaults. As part of those deals, Fluidstack became one of the first third party vendors to receive Google's TPUs. Now that wasn't massive news back in September when the deals were struck, but now that the market narrative views TPUs as a genuine contender to Nvidia's dominant GPUs, the that is changing. Fluidstack also secured the contract to build a gigawatt capacity data center in France as part of President Emmanuel Macron's push for sovereign AI. They are additionally the infrastructure partner for Anthropic's $50 billion data center investment announced last month. The new funding round will reportedly be led by Situational Awareness, which is of course the hedge fund started by former OpenAI researcher Leopold Aschenbrenner. Moving to our next story, hype around Opus 4.5 continues to build as the model keeps pushing the limits. Saesh Kapoor, who you may know from the AI Is Normal technology blog, announced that his team are ready to declare that Opus has solved the core bench scientific agent benchmark. The benchmark requires agents to reproduce scientific papers. When given the code and data from a paper, the agent is scored on its ability to set up the repo from the paper, run the code, and then correctly answer questions about the result. Functionally, it's a benchmark primarily about agentic code execution. Core Bench uses a common agent scaffold called Core Agent to allow comparison between different models on a level playing field. Opus 4.5 was initially tested using Core Agent and scored 42%, a solid score, but not close to Opus 4.1's leading score of 51%. DeepMind researcher Nicholas Carlini then reached out to the team with a new scaffold that uses Claude code, as well as some issues with the way the benchmark was being scored. Core Bench Team ran the benchmark again using the Claude code harness and found that Opus 4.5's performance almost doubled to 78%. Interestingly, a jump of this size was unique to Opus 4.5. Sonnet 4 and 4.5 saw much smaller improvements and Opus 4.1 actually went backwards, kapoor wrote. We're unsure what led to this difference. One hypothesis is that the Claude 4.5 series of models is much better tuned to work with Claude code. Another could be that the lower level instructions in Core Agent, which worked well for less capable models, stopped being effective and hinder the model's performance for more capable models. The Core Bench team also manually went through their benchmark, weeding out grading errors that Carlini had pointed out. Eight tasks were being incorrectly marked as wrong due to small floating point errors and one task was impossible to reproduce due to a dataset being removed from the Internet. The team manually scored Opus 4.5's performance at 95% with only two tasks failed, Kapoor wrote. With Opus 4.5 scoring 95%, we're treating core Bench hard as solved. The team now plans to pivot to an undisclosed set of test questions for their next benchmark to ensure the questions aren't included in training data now outside the benchmarks, the personal testimonials for 4.5 opus just continue to roll in. Dan Shipper from every who was very bullish to begin with, wrote a new piece going even farther, he said on Twitter. Opus 4.5 blew me away. This week I built a fully featured reading companion app that I now use every day in between meetings without looking at the code. Two things that are important. We just reached a new level of autonomous coding. You've been able to one shot an impressive app demo for a while now with any frontier model. Opus 4.5 is the first model that just keeps coding and coding without running into endless loops of errors. Second prompt native apps are now possible. Opus 4.5 can now act as a general purpose agent inside your app to power many of your features. This turns building features into an exercise in writing prompts instead of writing code. The NYT's Kevin Roos is also finding Opus 4.5 great for non coding purposes. He writes, Claude, Opus 4.5 is a remarkable model for writing, brainstorming and giving feedback on written work. It's also fun to talk to and seems almost anti engagement maxed the other night I was hitting it with stupid questions at 1am and it said Kevin, go to bed now. As for me, I have not yet found myself switching away from GPT5.1 or Gemini 3 to Opus4.5 all that often. But with all of this chatter, it seems clear that I'm going to have to give it an even bigger swing. Couple more stories Like I said, we are on an extended Headlines today Little bit of Market and Adoption News Salesforce has delivered a strong revenue forecast on the back of Agent Force adoption. Salesforce said their Q4 revenue would be between 11.1 billion and 11.2 billion, outstripping analysts forecasts of 10.9 billion. They also said that remaining performance obligations, a measure of future bookings, would increase by about 15% compared to analyst estimates of 10%. CEO Mark Benioff credited their AI focused products, stating, Our Agent Force and Data360 products are the momentum drivers. Active customer accounts for Agent force have grown 70% quarter over quarter, with many customers now transitioning from the pilot phase to active deployment. Benioff said that they now have over 9,500 paying agent force customers. He said, we've delivered incredible results with AgentForce. It's really exceeding our expectations. This is our fastest growing product ever. Now, one interesting sub wrinkle that I'm watching with the Salesforce story. A big question that many have is to what extent models get commoditized in the future. Salesforce, for their part, has primarily built on top of OpenAI models since they launched Agent Force in late 2024. However, last week Benioff posted, I've used ChatGPT every day for 3 years. Just spent 2 hours on Gemini 3. I'm not going back. The leap is insane. Reasoning, speed, images, video, everything is sharper and faster. It feels like the world just changed again. Then on Thursday he posted, LLMs are the new disk drives commodity infrastructure you hot swap for whoever's cheapest and best. The fantasy that the model is a moat just expired. So interesting things to watch to see how Salesforce thinks about model switching and what that means for the rest of the market. An even bigger market story yesterday, if only a little tangentially related to AI is that Meta could be giving up on their namesake technology. With rumors of deep cuts to the Metaverse division, Bloomberg reports that the Metaverse Group could see budget cuts as high as 30% next year. Their sources said cuts of that magnitude would most likely include layoffs as soon as January of next year. They did caveat that no final decisions have been made, but deep cuts to the Metaverse Group are on the agenda for end of year budget planning sessions. Sources said that Zuckerberg has asked for 10% cuts across the board, which has been the standard request for the past few years. However, the Metaverse Group was signaled out for deeper cuts due to the lack of industry wide competition over the technology. Now, for most public market investors, it's hard for them to see the Metaverse as anything but a massive disappointment, especially relative to the pitch in 2021. Zuckerberg presented the Metaverse with such conviction that he changed the name of the company. Since then, their Metaverse Group has been nothing short of a cash incinerator. The group has lost more than 70 billion since the metaverse strategy was announced, and thus, unsurprisingly, markets responded well to the idea that Meta would be slashing that particular category of spend. The Stock jumped by 5.7% in its largest intraday move since July. Now, while the Metaverse Group is being slashed, that doesn't necessarily carry over to the parent division, Reality Labs. That broader division is focused on Meta's various AR and VR products and has been going from strength to strength in recent years. The Meta ray bans have been a surprise hit and now define their product category, which is presumably a product category only becoming more important as LLM capabilities catch up to the promise of AI wearables. A Meta spokesperson suggested this strategy pivot is underway. Commenting within our overall Reality Labs portfolio, we are shifting some of our investment from Metaverse towards AI glasses and wearables. Given the momentum there, we aren't planning any broader changes than that. The reallocation of resources also aligns to Meta's poaching of veteran Apple UX designer Alan Dye earlier this week. On Wednesday, Zuckerberg announced that Dye would lead a new creative studio within Reality Labs that would focus on design, fashion and technology. In a post on thread, Zuckerberg wrote, we're entering a new era where AI glasses and other devices will change how we connect with technology and each other. The potential is enormous, but what matters most is making these experiences feel natural and truly centered around people. With this new studio, we're focused on making every interaction thoughtful, intuitive and built to serve people. So friends, that is the story from this Extended Headlines edition, but for now, we'll wrap it there and move on to the main episode. Foreign.
Is brought to you by my company, Superintelligent. Superintelligent is an AI planning platform, and right now, as we head into 2026, the big theme that we're seeing among the enterprises that we work with is a real determination to make 2026 a year of scaled AI deployments, not just more pilots and experiments. However, many of our partners are stuck on some AI plateau. It might be issues of governance, it might be issues of data readiness, it might be issues of process mapping. Whatever the case, we're launching a new type of assessment called Plateau Breaker that, as you probably guessed from that name, is about breaking through AI plateaus. We'll deploy voice agents to collect information and diagnose what the real bottlenecks are that are keeping you on that plateau. From there, we put together a blueprint and an action plan that helps you move right through that plateau into full scale deployment and real roi. If you if you're interested in learning more about Plateau Breaker, shoot us a note. ContacteeSuper AI with Plateau in the subject line Meet Rovo, your AI powered teammate Rovo unleashes the potential of your team with AI powered search, chat and agents or build your own agent with Studio. Rovo is powered by your organization's knowledge and lives on Atlassian's trusted and secure platform, so it's always working in the context of your work. Connect Robo to your favorite SaaS app so no knowledge gets left behind. Robo runs on the Teamwork graph, Atlassian's intelligence layer that unifies data across all of your apps and delivers personalized AI insights from day one. Rovo is already built into JIRA Confluence and Jira Service Management Standard, Premium and enterprise subscriptions. Know the feeling when AI turns from tool to teammate? If you Rovo, you know Discover Rovo, your new AI teammate powered by Atlassian get started at ROV as in victory o.com AI isn't a one off project. It's a partnership that has to evolve as the technology does. Robots and pencils work side by side with clients to bring practical AI into every phase. Automation, personalization, decision support and optimization. They prove what works through applied experimentation and build systems that amplify human potential. As an AWS Certified Partner with Global Delivery Centers, Robots and Pencils combines reach with high touch service where others hand off. They stay engaged because partnership isn't a project plan. It's a commitment. As AI advances, so will their solutions. That's long term value. Progress starts with the right partner. Start with robots and pencils@ropotsandpencils.com aidaily Brief this episode is brought to you by Blitzi, the enterprise autonomous software development platform with infinite code context. Blitzi uses thousands of specialized AI agents that think for hours to understand enterprise scale code bases with millions of lines of code. Enterprise engineering leaders start every development sprint with the Blitzi platform, bringing in their development requirements. The blitzi platform provides a plan, then generates and pre compiles code for each task. Blitzi delivers 80% plus of the development work autonomously while providing a guide for the final 20% of human development work required to complete the sprint. Public Companies are achieving a 5x engineering velocity increase when incorporating Blitzi as their pre IDE development tool, pairing it with their coding pilot of choice. To bring an AI native SDLC into their org, visit blitzi.com and press get a demo to learn how Blitzy transforms your SDLC from AI Assisted to AI native.
Welcome back to the AI Daily Brief. In an episode earlier this week I talked about how I thought that heading into 2026 we were likely to see a lot more studies and research and experiments that were trying to figure out just how much of the current slate of human work AI was actually able to do. We got a McKinsey study a couple of weeks ago that said that up to 57% of tasks could be automated. More recently we got that MIT iceberg report which said that 11.7% of value generating tasks could be automated. Now of course those things have been translated by mainstream media into headlines that 57% of jobs or 12% of jobs are going to be lost. If you want a reputation of why 12% of tasks being able to be automated doesn't mean 12% of jobs going away, listen to yesterday's episode. But alongside those types of studies, what I hope for is that we're also going to get more research around how these things are playing out in practice. There is a seismic gap and massive difference in what AI can theoretically do and what it is actually doing in practice. Now, one of the companies that is most on the spot right now when it comes to providing some amount of that real lived experience information is anthropic. Yesterday we looked at some research that they did around their own team where they had interviewed researchers and engineers in August to figure out how Claude and Claude Code were impacting their work. And today we're looking at an Even more expanded look at how AI is working in practice with the introduction of Anthropic Interviewer. The TLDR is that Anthropic launched a new research tool and tested it by asking professionals about their experience working with AI. Now I want to focus mostly on the results and what the professionals actually said, more than the tool itself. But it is worth mentioning the tool itself a little bit because some are understanding that holding aside this specific use case for it, this potentially represents a broader pattern in how research happens in the future. Now, in their introduction, Anthropic points out that while they recently developed clio, which is a privacy preserving system for getting insights from real world AI use of Claude, there were inherent limits there. As they write, the tool only allowed us to understand what was happening within conversations with Claude. What about what comes afterwards? How are people actually using Claude's outputs? How do they feel about it? What do they imagine the role of AI to be in their future? If we want a comprehensive picture of AI's changing role in people's lives and to center humans in the development of models, we need to ask people directly. Such a project, they noted, would require us to run many hundreds of interviews. Here we enlisted AI to help us do so. Now Google's Tao Dong got that there was something interesting about the form factor here, he writes. After reading the project blog post and a few transcripts, my initial impression is that we're seeing a new genre of user research, a crossover between surveys and interviews. I'm tempted to call it semi structured surveys. It acts like a survey with predefined open ended questions, but with the ability to ask decent follow up questions on the fly. While these 10 to 15 minute sessions weren't particularly deep, they combined the scale of a survey with the flexibility of a moderator. Pairing this with AI analysis allows the team to identify quantitative patterns and actually explain why they exist. It seems like a fascinating experiment. What's your take? My take is that this is pretty much exactly what we built, or at least a version of it with superintelligent in superintelligence audits and assessments. Whether they are our agent readiness assessments or our new plateau breaker assessments. One of the key ideas is that surveys are great for scale but bad for context. Interviews are great for context but bad for scale. But with AI, particularly voice AI, you don't have to make that trade off and rather than inferring from a small sample, you can just go ask everybody. Now I don't think that this is some crazy novel insight, nor do I think that Our technology is some stratospheric leap, but I think that this particular pattern of using AI to radically scale information gathering and then speed up information analysis is something that is absolutely going to become de rigueur for all sorts of current research processes. And in so doing, it is going to open up totally new types of research that weren't possible before because of the scale of information that you can collect and analyze. So before we move on to the specific results, I would say if you are thinking about interesting research projects across basically any domain, if they involve talking with people, I believe that you can radically increase your ambition thanks to the new tools that are available. All right, so back to this actual survey of these 1250 professionals. For our purposes, what I'm most interested in is what they said about working with AI. Now, presumably this group is probably going to be more enthusiastic than a random sample of 1250 people. And so I think that that caveat is important. But within that, some of the high level insights from Anthropic are that one people. People are optimistic about the role that AI plays in their work. Positive sentiment characterized the majority of topics discussed. However, as we'll see, there are a small number of topics that have more relatively pessimistic outlooks. A second insight which I think is really valuable about how we design systems and think about displacement. People from the general workforce want to preserve tasks that define their professional identity while delegating routine work to AI. They envision futures where routine tasks are automated and their roles shift to overseeing AI systems. Again, I don't think that that's particularly novel, but it's interesting to see that that is how people are starting to think about their role as well. We talk a lot as insiders about this idea of shifting to a model where humans manage AI agents and AI systems. But it's interesting to see that start to come out as an expectation or a goal from individual professionals as well. A third insight which absolutely resonates with what I'm seeing is that despite creatives facing pure judgment and anxiety about their future, they are turning to AI to increase their productivity. As Anthropic puts it, they are navigating both the immediate stigma of AI use in creative communities and deeper concerns about economic displacement and the erosion of human creative identity. Lastly, number four, Anthropic writes that scientists want AI partnership but can't yet trust it for core research. Scientists uniformly express a desire for AI that could generate hypotheses and design experiments, but at present they can find their actual use to tasks like writing manuscripts or debugging analysis code. So what's interesting here, as opposed to some of these other areas, is that it sounds like scientists want AI to do more, or at least be more helpful with their core functions, not just those routine tasks to be automated. So let's look at the visualization. If you're listening to the show, I'll go through this pretty fast, but if you are watching it, the blue gray represents more pessimistic, the muted yellow represents more optimistic, and you can see across almost every category optimism mostly beats out pessimism. The one area it appears to me, where there's relatively more pessimism, at least among the general workforce, is in career adaption, which makes sense. Now. Among creatives, there are a few areas where again, pessimism takes a little bit more root. In particular, artist displacement shows actually people more pessimistic than optimistic overall. Same with writer displacement. Among scientists, the biggest area that saw actual more pessimism than optimism is around security concerns. And when you dig into the examples they shared, a lot of it reflects broader sentiment that you hear day in and day out on social media. For example, a lot of folks are trying to figure out what parts of their jobs won't be automated, which parts of their skills will be valuable in a future where they assume AI is ubiquitous. For example, a trucking dispatcher said, I'm always trying to figure out things that humans offer to the industry that can't be automated and really hone in on that aspect, like the personalized human interactions. However, that is not something that I think will be necessary in the long run. I'm still trying to figure out what skills would be good to work on that AI can't take over, obviously way bigger than just that particular job role. This question is particularly pertinent, I think, because it's not only something that people should be asking individually, but it's also something that people who are designing upskilling and retraining systems need to be hyper conscious of. It is not going to be particularly useful if we design a bunch of training programs that just get obviated by GPT7. Another thing that comes up on the pessimism side is the stigma of using AI. A salesperson, for example, said, I hear from colleagues that they can tell when email correspondence is AI generated and they have a slightly negative regard for the sender. They feel slighted and the sender is too lazy to send them a personalized note and push it onto AI to do it. I think one really interesting question is to what extent. That is a temporary transitional feeling where in the future people will feel like of course they used AI to write an email or if that's going to be something that's more persistent. On the optimism side, however, you see tons of reflective comments people looking to AI to help them manage their time, expand their creativity, reduce their stress by allowing them to focus on the best parts of their job. Overall, 86% of professionals reported that AI saves them time and 65% said that they were satisfied with the role AI plays in their work. Across different categories of work there was a pretty similar distribution of frustration satisfaction with slightly expanded worry in categories like art, design and media. Among creative professionals there is a much bigger band of responses. Designers, for example, see much more frustration than filmmakers and in many cases you just see complication even within individual categories. Worry and satisfaction and hope and frustration all sitting alongside one another. In my estimation. This is the type of survey that needs to happen not once in a while but on a very regular basis. I want to see these questions tracked over time and I want to see that data available to policymakers of all stripes. Now one really cool thing about this as we wrap Anthropic is making all of the data available in a public dataset that you can download from Hugging Face of course with all the participants approval. Meaning that if you are so interested in, you can go interact with and run your own analysis on this research as well. Overall I think these 1200 professionals tell us a lot of the same story that we've been seeing for years now. A future that has so much opportunity but is fundamentally different and in that difference somewhat scary as well. Good job to Anthropic for digging up this real information, but that's going to do it for today's AI Daily brief. Appreciate you listening or watching as always and until next time peace SA.
Ep: What 1,250 Professionals Say About Working With AI
Host: Nathaniel Whittemore (NLW)
Date: December 5, 2025
Nathaniel Whittemore (“NLW”) delivers a rich, analytic episode centered on a new Anthropic survey of 1,250 professionals. The survey explores how workers across domains feel about AI’s impact on their jobs, revealing nuanced attitudes toward automation, professional identity, creative work, and trust in AI. The episode also touches on how Anthropic used an AI-driven interview tool to scale qualitative research, reflecting broader trends in research methodology. The host underscores the importance of understanding real experience—beyond theoretical studies—and calls for regular tracking of workforce sentiment to help shape policy and adaptation strategies.
“There is a seismic gap and massive difference in what AI can theoretically do and what it is actually doing in practice.”
— NLW, 19:30
“Surveys are great for scale but bad for context. Interviews are great for context but bad for scale. But with AI… you don’t have to make that trade off.”
— NLW, 23:10
“People from the general workforce want to preserve tasks that define their professional identity while delegating routine work to AI…”
— NLW, summarizing Anthropic findings, 25:08
“They [creatives] are navigating both the immediate stigma of AI use in creative communities and deeper concerns about economic displacement and the erosion of human creative identity.”
— NLW, summarizing Anthropic, 26:40
On adaptation and human value:
“I’m always trying to figure out things that humans offer to the industry that can’t be automated and really hone in on that aspect, like the personalized human interactions. However, that is not something that I think will be necessary in the long run…”
— Trucking Dispatcher, 27:05
On stigma of AI-generated communication:
“I hear from colleagues that they can tell when email correspondence is AI generated and they have a slightly negative regard for the sender. They feel slighted and the sender is too lazy…”
— Salesperson, 27:35
“This is the type of survey that needs to happen not once in a while but on a very regular basis.”
— NLW, 28:45
NLW closes by emphasizing that the survey validates a familiar dual narrative: immense opportunity layered atop anxiety and uncertainty about the nature of future work. He commends Anthropic for providing real, ground-level data and encourages ongoing iteration and openness.
"A future that has so much opportunity but is fundamentally different and in that difference somewhat scary as well."
— NLW, 29:03
Tone:
NLW maintains a thoughtful, analytical, and accessible tone, balancing optimism and caution as he stitches together research findings, personal experience, and broader trends in AI's impact on work.