
Alibaba launched Qwen3.8 Max as Moonshot paused Kimi K3 signups amid demand. The Trump administration weighed a slow squeeze on Chinese AI, Hugging Face used China's GLM-5.2 after US guardrails blocked its breach forensics, and Google built a Gemini chip.
Loading summary
A
Athletic Brewing Company crafts award winning non alcoholic beers for those who want to be part of every round. With over 185 flavor awards, they're exceptional NA beers that fit your lifestyle and any social occasion. Summer's full of good times and Athletic fits right in. Go to athleticbrewing.com to have brews delivered to your door or find them at a bar, restaurant or store near you near Beer Athletic Brewing Co. Fit for all times.
B
Welcome to the tech we write home from Monday, July 20, 2026 I'm Brian McCullough. Today, Alibaba launched Quinn 3.5 Max as moonshot pause Kimmy K3 signups amid demand the Trump administration weighed a slow squeeze on Chinese AI hugging face used China's GLM 5.2 after US guardrails blocked its breach forensics and Google is building a chip with Gemini baked in. Here's what you missed today in the world of tech. Every day, shareholders meet to discuss important matters about the companies you invest in. Now you can make your voice heard too. Vanguard Investor Choice makes it easy to set your proxy voting preference for your eligible Vanguard index funds. Whether you hold a Vanguard fund directly or through another brokerage firm, all it takes is a few clicks to select your proxy voting preference and be heard on important shareholder topics like executive pay and director elections. Visit vanguard vanguard.com investor choice to learn more. It's your shares. It's your voice. It's easy. Vanguard investors own shares of their index funds, and those funds own shares of the companies they invest in. Vanguard Marketing Corporation Distributor There was another one over the weekend, Alibaba launched a 2.4 trillion parameter model called Quen 3.8 Max into preview that it says rivals Frontier AI models and is second only to Fable 5 on various benchmarks. Meanwhile, Moonshot AI was forced to pause new subscriptions, saying that over the past two days, Kimik2 demand nearly passed the limit of its capacity, and they're also reworking their plan tiers accordingly. But back to the new QEN model first quoting Bloomberg. The Sunday release came only days after startup Moonshot AI unveiled a powerful new offering that's roiled markets and triggered concern in the US about China closing the gap on global leaders like Anthropic and OpenAI. Quinn 3.5 Max has 2.4 trillion parameters, joining Moonshot's Kimmy Kim K3 as the heavyweight in the open weight class with 2.8 trillion parameters. K3 rivals top offerings, and Alibaba is setting similarly high expectations. Developers can now access Qin 3.8 Max through Alibaba's coding platforms, including coder. Alibaba plans to make the model open wait soon, expanding access beyond the preview release. Interest in these made in China artificial intelligence systems and models is so high that Moonshot was forced to pause, taking on new subscriptions late on Sunday to manage overwhelming demand. While optimism around Alibaba is growing, other contenders in China's hotly contested AI race have suffered a drop in the wake of the new Kimi release. Rival Zhipu declined nearly 30% on Friday and added a further 14% to the losses on Monday after being one of the star debut stocks in Hong Kong for much of this year. Alibaba, China's E commerce leader and one of its biggest investors in AI, recently scored another victory after Beijing approved Apple Intelligence, the software suite for iPhones, iPads and other Apple gear, which will use Alibaba technology in that country. Quote so again, Chinese AI and open source AI is having another moment. But quoting the decoder, the Trump administration is considering measures against Chinese AI models that could amount to an effective ban, according to Axios, the Department of Commerce, the NSA and the White House have explored several options since 2025. These include placing Chinese AI labs on a sanctions list, issuing security warnings, and using an executive order to impose security requirements and liability on US Companies that host Chinese models. The Commerce Department reportedly drafted rules as early as the summer of last year to protect domestic supply chains from Chinese open source models. Advisors who favored a lighter regulatory approach initially blocked these efforts. But the release of China's Kimi K3 model and personnel changes in the White House have helped supporters of tighter restrictions regain influence. A direct ban wouldn't even be necessary. A source close to the government told Axios that what's actually happening is slower and more durable, pointing to procurement rules, sanctions, threats and public pressure campaigns against U.S. companies that use Chinese models. Rather than ban Chinese models outright, the administration could focus on potential backdoors and security flaws, another source said, that closely matches the FUD strategy, short for fear, uncertainty and doubt that OpenAI strategist Dean Ball recently predicted soft guidelines and public warnings could deter companies without imposing binding rules. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models. This will just just drive startups to sketchier providers. There's a happy middle ground here, ball wrote. Commercial and economic interests may also be driving the push. US Companies are increasingly using Chinese open source models because they're cheaper and nearly as capable. Restrictions would protect the market dominance of Google, OpenAI and Anthropic. The AI sector is also driving much of the US Stock market's gains under President Trump. If Chinese models threatened the business of major US providers, the fallout could hit markets hard. End quote. You know what? Let's turn to Ben Thompson to get a read on this new China AI moment. Quoting Ben Nvidia CEO Jensen Huang has described what Nvidia is building as token factories, and from Nvidia's perspective, that framing makes sense. Nvidia GPUs are model agnostic. They generate tokens and do so in the fastest and most efficient way possible. That leads to measurements like tokens per second, time to first token tokens per watt, token cost, etc. And Huang argues that these metrics will be the basis for decision making. This is a framing that definitely made sense during the first paradigm of AI, the ChatGPT era, where tokens were delivered straight to the end user. The second paradigm of AI, however, the reasoning era, confounds this measurement. Reasoning entails an explosion in chain of thought tokens and different models need different amounts of reasoning tokens to arrive at the right answer. Kimi, for example, reportedly uses significantly more tokens than sol, rendering its price advantage moot. Agents introduce a similar dynamic. Some models are more efficient than others in terms of the number of tokens they need to execute agentic workflows. What this means is that tokens are not a commodity. The defining characteristic of a commodity is that it is fungible. A gallon of oil is a gallon of oil. A ton of copper is a ton of copper. A bushel of wheat is a bushel of wheat. A token from one model, however, is not the same as a token from another model. What is fungible is what is constructed from tokens, which is to say intelligence. In other words, if both Kimmy and Saul generated the right answer, then that answer is fungible. The difference is in tokens generated to get to that right answer is a contributor to a difference in the costs of goods sold. The reason this matters is that we are rapidly approaching a state in which intelligence for many economically beneficial tasks is in fact a commodity. Anyone building a basic CRUD app, for example, can likely do so using models from multiple providers. And in a commodity market, the route to profitability is not through charging higher prices. Again, you can or you will soon be able to make the exact same app using multiple models, but rather through having a superior cost structure. Right now, none of the above analysis applies because demand exceeds supply for frontier models and supply is limited by a lack of compute. This compute shortage doesn't just mean that a compute supplier like Nvidia makes very large margins, but also that Nvidia's customers like SpaceX AI can turn around and resell compute at high margins as well to a company like Anthropic. Anthropic, meanwhile, can pay the markup because they can sell tokens with a higher markup. Still, it's not just excess demand that gives Anthropic great margins. However. Anthropic and OpenAI likely have among the lowest costs per unit of frontier quality intelligence thanks to model capability serving scale and token efficiency. They are serving models at a particular capability level for months before their competitors and are simultaneously applying the best models to optimize those costs. It's also worth noting that the market is not yet treating intelligence like a commodity. Demand is for Anthropic and OpenAI specifically, and much less for models that aren't as good. Thus SpaceX AI and Meta selling capacity to Anthropic. One way to think about the push for optimizing costs is that it is a function of defining jobs to be done by intelligence levels such as intelligence. Buyers can create a market where intelligence is commoditized in the long run. However, whoever is on the frontier is the best place to dominate non frontier markets as well, which are just the frontier minus n months, I.e. months in which the frontier model makers have been optimizing their cost of serving. All of this is to say that I think the reaction to Kimi and Chinese models generally is pretty overblown, at least from an economic perspective. Right now there is a price umbrella that is downstream of the lack of compute. I highly doubt that Chinese models are cheaper to serve on a marginal cost basis. They just seem cheaper because Anthropica and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intell. I think the frontier labs are anchored in a world where training costs dominated their financial modeling. As long as training costs consume more GPUs than inference, it was critical to maximize inference revenue to help fund the next training run, which meant charging very high prices for inference. Going forward. However, I expect the inference market to grow much faster than training costs, and that includes the assumption that training costs will continue to skyrocket, which means they really can make it up in volume. It wasn't clear this would be the case as recently as eight months ago, but the agent paradigm unlock is so massive that Frontier Labs should have more confidence that they can not just survive, but thrive with lower prices once they have sufficient compute. Second, intelligence isn't in fact a perfect commodity, in part because applied intelligence makes itself smarter. Specifically, whoever is running inference is also collecting data, and that data goes into making the next iteration of the model better. This is, on one hand, all the more reason for the Frontier Labs to lower prices, increase usage as more compute comes online. On the other hand, this is also why companies like Microsoft are increasingly obsessed with helping companies run their own models. This is much more viable if Chinese models are a viable alternative. End quote.
A
Your team just added its 67th AI tool and also your 67th security blind spot. The good news? The Vanta Agent works like a GRC engineer in the background, finding every app your team uses, scoring the risk, and drafting fixes for you. Vanta is the platform used by over 16,000 fast moving companies like Ramp, Cursor and Harvey, who are shaping the future with AI and staying ahead of AI risk. Get started@vanta.com
B
when critical company knowledge isn't documented, there's a major ripple effect. Work becomes inconsistent, tools don't get adopted, and knowledge walks out the door when someone leaves. Thankfully, our sponsor, Scribe, was built to fix that. Their workflow AI platform is trusted by nearly half of the Fortune 500 to capture workflows in real time. Here's how it works. You turn on the Scribe browser extension or desktop app, do a process as you normally would, and Scribe will build a guide as you go. It automatically redacts sensitive information and even suggests improvements to your existing workflows. To see what Scribe could look like for your org, head to Scribe How Ridehome and mention Ride Home for your first month of Scribe capture free on select plans. That's S C R I B E How Ride Home
C
A burst pipe? A dead water heater? The AC calling it quits? Who do you call? HomeServe is an easy way to handle unexpected home repairs with plans covering stuff basic homeowners insurance usually won't. Instead of scrambling for a contractor, you make one call to get the repair process started. Join the millions of customers who trust HomeServe right now. Go to HomeServe.com podcast for 50% less your first year. That's HomeServe.com podcast savings compared to renewal price void in Florida so get this
B
Hugging Face says an agentic AI system hacked into its data pipeline, accessing several internal clusters and credentials. But here's the rub. Hugging Face's own AI based triage is what caught the breach. Hugging Face says it used the Open Weight GLM 5.2 model hosted on its own Compute for breach forensics and after US Frontier model safety guardrails, Reed Anthropic blocked requests by IT to do similar forensics. Quoting the stack, the platform security team were initially stymied in their incident response by unnamed US LLM Frontier model guardrails, quote which cannot distinguish an incident responder from an attacker, they said. So Hugging Face's defenders turned instead to the open source GLM 5.2 model from China's Z AI lab, running it on their own infrastructure to analyze the more than 17,000 logs or footprints that the attackers left behind. That's a striking public admission for the New York headquartered Hugging Face, which lets users collaborate on models, data sets and applications, and which this summer hit the $100 million ARR mark. In an incident report, the company recommended that Defenders have a capable model you can run on your own infrastructure, vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. The attacker the unknown attacker abused two code execution paths in our dataset processing a remote code dataset loader and a template injection in a dataset configuration to run code on a processing worker or compute instance, said Hugging Face in a detail thin July 16th incident report. They then escalated to node level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend using what the firm said was a swarm of short lived sandboxes with self migrated grading C2 staged on public services. The company's lesson for defenders in their incident write up on July 16th stood out to both infosec practitioners and tech investors. When we started the log analysis, we first used Frontier models behind commercial APIs. This did not work. The analysis requires submitting large volumes of real attack commands, exploit payloads and C2 artifacts, and these requests were blocked by the provider's safety guardrails, Hugging Face said. It then ran the forensic analysis instead on China developed GLM 5.2, an open weight model on their own infrastructure, Hugging Face added. This had a second benefit no attacker data and none of the credentials it referenced left our environment. Hugging Face's incident report was published the same day that Chinese AI startup Moonshot's Kimi K3 model rocked global markets. The 2.8 trillion parameter model is the largest open weight AI model to date. Blind developer testing by Arena, a platform created by researchers at UC Berkeley for its front end code evaluation test, put Kimmy K3 ahead of Anthropic's Fable 5 and OpenAI's GPT 5.6 last week. Chinese frontier models are also notably cheaper than their US Counterparts, as data from artificial analysis shows below. Hugging Face did not say which commercial frontier models it had first tried to use for its ir, nor was it clear what model the attackers used. End quote. Sources tell the information that Google is developing a specialized server chip, informally dubbed Frozen V2, that integrates its Gemini AI model blueprint right into the actual silicon. Quoting the information, Google intends the new chip, informally dubbed Frozen V2, to help it address a major shortage in AI computing capacity that has fueled internal tensions and compelled Google Cloud to turn down deals with outside customers. Google employees working on the chip have projected that it could be 10 to 6 times more efficient than the newest version of Google's existing line of homegrown AI chips when it launches, based on the number of tokens, a basic unit of AI consumption it can serve per unit of power, according to the people. The Frozen name comes from the idea of permanently etching parts of the model into the silicon. Engineers are still deciding what the new chip's main features will be and how its different parts will work together. Google plans to deploy the chips as soon as 2028, the people said. The Frozen project is intended to create a new branch of homegrown chips that's different from Google's Tensor processing units, rather than to replace the TPUs like Nvidia's Graphics Processing Units, the most widely used AI chips today, TPUs are made to work with many different AI models. As a result, those chips must make a lot of time consuming decisions as they interact with whichever model they're running. Whereas Frozen V2 would have some of the decisions for Google's Gemini models built in, reducing the number of steps the chip must make and take and the amount of data it has to move around. This could make the frozen chip faster at responding to queries, potentially enabling new AI applications for Google, one of the people said. But it also means the new chip represents a bet that Google will stick with its current approach to building its Gemini models. The chip would be usable for later generations of Gemini AI models only if they are based on the same underlying model architecture in use when it was designed. Google would be able to make changes to the chip, including updating the model to take advantage of new model weights, essentially the settings that determine how models respond to questions, one of the people said. A range of companies, including smaller startups such as Sambanova and Dmatrix and giants like OpenAI and Microsoft are developing similar chips for AI inference, the computing work to run AI models. The goal is to handle inference more efficiently than Nvidia's GPUs to alleviate shortages of computing power that have driven up costs. Nvidia itself struck a $20 billion deal in December to license technology from Grok, a startup that designed inference chips. End quote. Nothing more for you today. Talk to you tomorrow.
Episode Title: It's All China AI All The Time
Date: July 20, 2026
Host: Brian McCullough (Morning Brew)
Theme: The rapid escalation of China’s AI capabilities, major new Chinese AI model releases, US regulatory responses, and the knock-on effects for the global AI market. Discussion also covers cybersecurity incidents, and innovations in AI hardware.
This episode focuses on an eventful few days in the world of AI, with China’s tech giants launching record-shattering large language models (LLMs) that are making waves globally. US reactions—from policy hesitations to protective trade moves—are discussed, alongside ramifications for big tech companies, cybersecurity, and AI infrastructure innovation. The host breaks down what this China-led momentum means for global competition and the future of AI.
Alibaba’s Qwen 3.5 Max Launch
“Alibaba launched a 2.4 trillion parameter model called Qwen 3.5 Max into preview that it says rivals Frontier AI models and is second only to Fable 5 on various benchmarks.” [00:33]
Market Impact
“Rather than ban Chinese models outright, the administration could focus on potential backdoors and security flaws… soft guidelines and public warnings could deter companies without imposing binding rules.” [04:30]
Ben Thompson’s analysis (cited at length on the show) dives into the shifting economics of AI and the new Chinese competition:
On Tokens and Intelligence
“Tokens are not a commodity. The defining characteristic of a commodity is that it is fungible… A token from one model, however, is not the same as a token from another model. What is fungible is what is constructed from tokens, which is to say intelligence.” [07:00–08:00]
State of the Market
“Demand is for Anthropic and OpenAI specifically, and much less for models that aren’t as good.” [08:55]
On Chinese Competition
Hugging Face Security Event
“We first used Frontier models behind commercial APIs. This did not work. The analysis requires submitting large volumes of real attack commands, exploit payloads… and these requests were blocked by the provider’s safety guardrails. It then ran the forensic analysis instead on China developed GLM 5.2, an open weight model on their own infrastructure…” [12:41–14:22]
Broader Implication
“The Frozen name comes from the idea of permanently etching parts of the model into the silicon… the chip would be usable for later generations of Gemini AI models only if they are based on the same underlying model architecture…” [14:22–16:02]
Brian McCullough on China’s AI momentum:
“Chinese AI and open source AI is having another moment.” [03:00]
On US regulatory strategy:
“You just create enough regulatory risk that every regulated enterprise backs off. There’s a happy middle ground here.” [04:55]
Ben Thompson’s insight on intelligence-as-commodity:
“Anyone building a basic CRUD app, for example, can likely do so using models from multiple providers. And in a commodity market, the route to profitability is not through charging higher prices… but rather through having a superior cost structure.” [08:20]
On the power of self-hosted, open-weight AI for security:
“Defenders have a capable model you can run on your own infrastructure, vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.” [13:53]
The episode delivers a concise but in-depth view into how fast the global AI arms race is moving, why the US is treading carefully, and how both countries—and their tech champions—are rethinking everything from pricing to hardware and security in the age of trillion-parameter models.