![[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model — SemiAnalysis Weekly cover](https://d3t3ozftmdmh3i.cloudfront.net/staging/podcast_uploaded_nologo/45491366/9cf0c89041603386.jpg)
Loading summary
A
All right, Max, we're going to do a podcast. We're going to Talk everything about Kimik3 and maybe some other models that just came out. How you doing?
B
Doing great, looking forward to it. And thanks for having me, Jordan.
A
I'm not having you, Sauce. Thanks for having me. Okay, on the docket, let's say Kimmy K3 hot takes. Is this the third best model in the world? Impact on OpenAI, anthropic architecture changes, personal usage that we've had so far, what we think about their open source strategy and maybe more. All right, Max, quick hot take. Is this the third best model in the world right now?
B
I think the answer is a clear yes. People love shitting on benchmarks. I think benchmarks definitely have their problems, but I think sort of if you take a composite of all the main benchmarks and just look at them all rankings, they have been directionally correct over time. And I think if you look at that composite today, there's like a pretty clear top three with Fable, Sol 5.6 and now Kimike 3. And there's sort of just like always above everyone else, which includes of course other open source guys like, you know, deepseek and GLM whoever. But also like, very notably it includes Google and Meta and SpaceX, which I think is honestly, it's an extremely impressive and a very remarkable feat from the Moonshot guys. Google in particular, I think should feel incredibly embarrassed right now that at one point, guys, remember, as recent as November December 2025, everyone thought that the clear AI big three was Google, Anthropic and OpenAI. And even when I talk to boomers today, they still seem to think their top three is Google, Anthropic and OpenAI. And it's just like clearly not the case anymore. So yeah, I'd say definitely third best mall in the world. I do think it's overall still worse than Fable and Sol 5.6. Kind of funny that they explicitly said that in their mod release blog post. Maybe it's some like old fashioned Chinese humility. Maybe it's like they don't want to incur scrutiny from the US government or anything because obviously there were some delays to the Fable 156 release. Um, but overall very impressed with the model.
A
Yeah, they. In the limitation sections of the blog post they said despite being a highly competitive model, overall K3 nonetheless exhibits a noticeable gap in user experience compared with Fable 5 and GPT 5.6 goal. So my experience using this personally is that it is good, it's really slow, which is really Annoying. It's motivated me to try open source harnesses for the first time. And so I feel like I'm learning more about the harnesses than I am about the models because frankly, all these models are like good enough to do the basic work that I've been doing so far. I can't really find a lot of complicated stuff that it can't do, which in and of itself is a bit of a feat. Here's, here's my hot take. For me, this might be the second best model in the world right now because every time I try and do something meaningful with Fable, I get rejected and I get sent down to Opus. And even though I don't know if this is better than Opus, it is less annoying to not get rejected whenever I'm trying to do something, however I am getting. I'm not getting rejected when I use my API key and pay per token, but I am hitting limits whenever I try and you know, just use the web console or Deep Research or like the coding plan. I haven't used a coding plan, but some other guys at somebody else's have. And so it leads me to be like, what, what is the strategy here? Because these guys clearly just do not have enough GPUs to serve the demand that they're got for this model. And previously that was solved by an open source strategy where they just drop the weights and then other people serve it and serve that demand, but they haven't dropped the weights yet. So I think they said weights in 10 days or something. What do you think the strategy is for the delay between announcement of the model, the API being available, and no waits yet?
B
Yeah, I mean, I think to be clear, this is all just pure speculation on my part, but I think one big reason is they need to give the VLM and SGL guys enough time to make sure they can serve this model performantly. Because if they just dropped it today, you have all this hype, but then everyone else serving the model is one giving you like 20 tokens per second or something. That's probably really bad for their brand. Doesn't really sort of like capture. So they have this incredible opportunity to get a bunch of huge PR and a bunch of adoption. And that was sort of like kneecap it a little bit. I think another possibility is that they are actively talking to the Together Fireworks, Nebius's core use of the world to figure out how they can sign some sort of licensing deal and have them serve, you know, the incremental capacity on, you know, GV3 hundreds or whatever. In my mind those are sort of like the two main reasons why you would wait 10 days to actually draw the model weights.
A
Yeah, makes sense. And I think functionally this is really interesting because just to talk about the model architecture for A second, it's 2.8 trillion parameters. This does not fit in a B200. So you need to have B300GB300 or I guess MI355X in order to be able to serve this model on a single system like a single 8 way HTX server. Of course you can do unique things where you have pipeline parallelism across multiple nodes and stuff, but that's going to really impact performance. So I think there's a lot of recipes being cooked up and only people who have the latest and greatest chips are going to be able to serve this model. Um, just a. Let's, let's go back to your comment on, on Google for a second. The idea that this model is truly competitive in at the frontier at 2.8 trillion parameters kind of gives us some insight into how big the closed source frontier models are. Right. Like it would be even more embarrassing if they're hitting these levels of performance and they're being compared to 10 trillion parameter models with a lot more active. Right. We have to assume that this is in the same range as what Soul and Fable are.
B
Right? Yeah, I think that's a great point. And you just have to be correct. I think I still believe in sort of the competence and the correctness of all the OpenAI anthropic researchers. And there are some people on Twitter who like to claim that current closed source is like 10 trillion total parameters or something. If that's actually true, guys, it's time to pack up the bags. Spy's probably should crash 50% tomorrow. It's over. I'm pretty confident that Timmy K3 cannot be much bigger or sorry, much smaller. If anything, it might even be slightly bigger than the leading open source models today. Sorry, the leading closed source models today. And if that's true, this actually further highlights a point we've been harping on for a while on semi analysis which is that the margins for these like closed source labs have to be absolutely mind boggling because if you're telling me, you know, Kimi's probably not at negative margins and they're serving K3 at 3, 15, $3 per million input tokens. $15 per million APA tokens. That's sort of the same price as Sonnet. And so if you're telling me that like Fable is probably similarly sized and anthropic. Can charge $10 per million at the tokens and $50 per million APA tokens. Then this should just like it may dispel any of the remaining concerns people have about the AI labs being these unprofitable businesses. Like selling tokens at API prices is just, it might be even better than like SaaS. Honestly it's just, it's an incredible business today.
A
Yeah, makes sense. No, no cost of employees, just the GPUs. So can you compare this pricing strategy to the previous stuff? Because you said it's at 315. The previous version from moonshot directly was at 95 cents and $4. So we're talking about more than 3x pricing increase from 2.7 code to kimik3. Do they have more, even more room to increase pricing? Like what's the curve to get the Frontier open source intelligence or Frontier soon to be open weight yet yet to be determined what the license will be intelligence?
B
Honestly I don't think they have that much more room to push pricing up. Uh, like I would guess that even at like this 3:15 there will be a lot of people who are like this is a little too expensive for me. Uh, my task is easy enough for like a GLM 5.2 or a minimax M3 and I might just like use one of those models instead. And I actually think sort of. So on one end, sort of you have like the semi analysis of the world, right, where we don't really care how much money we're costing. Dylan, when we burn tokens all day, we're very happy using Fable for even a relatively easy task that were pretty confident one of these open source models can do pretty well. And then on the other end you have people who are like extremely cost conscious. You Maybe only get $200 worth of tokens per week as we heard some large companies like Tesla and Uber are implementing and I think pretty much all the people in that second bucket are going to be wanting to use the GLM kind of pricing tier models because they are already good enough for most everyday tasks. And then everyone in the semi analysis bucket is still using like 5, 6, SOL and Fable. So I think there actually is a pretty interesting question of who is the user that is actually going to be switching towards Chemike 3. I think it might just be a lot of people who philosophically love open source and are excited to try this new hype model and support it. But it wouldn't surprise me at all if there's like not actual serious adoption amongst say like large enterprises of this model.
A
Okay, what do you think about where we go from here? Like this is obviously a new base model. It's a fully new architecture for these guys. 2.8 trillion parameters. They've got Kimmy Delta, attention, attention residuals. The stable latent MOE that they keep using the it's like a scaled up bigger version of the previous models. Clearly about two times bigger. But previously with K2.5 we saw cursor train Composer based on just continued pre training as well as some RL and then we saw Kimmy give us 2.5, 2.6, 2.7 checkpoints as they just continued the RL. This is a new base model. It seems pretty complete. Like in my usage it's working pretty well. It's not screwing up anything basic when it comes to writing a PR description or like totally going off the rails the way that we've actually seen Some other models that are kind of raw without a bunch of RL have some rough edges at the beginning. I'm not seeing those yet. So where do we go from here? When does 3.1 come out? How does pricing change over time? Like does Composer? Do we get a composer based on
B
Kimmy, K3, I mean composer based on Kim and K3? Definitely not because I think the Cursor guys are pretty set on training the RO model from scratch now. As for when like you know, K3 1, K3 2, whatever come out, I imagine you probably see like two or three updates within like the next each like a month or two apart or something as they just continue post training this thing. I would guess sort of like pricing stays about the same just because I don't like they're not going to be able to run it on new hardware in the next two to three months so they're not going to get a huge throughput increase there to reduce pricing. Maybe it's possible some really craft, I don't know, kernel engineers figure out how to reduce the cost to serve this thing so it's closer to Deepsea v4 pricing or something. That would be really impressive, but given that it's a 3 trillion parameter model, I think I'm a little skeptical. I would guess that the current pricing we see for the minimaxes and the glms is kind of already pushing the limits of what you can serve a 1T to 1.5T model at and not have just embarrassing bad margins. So I feel like this pricing is probably here to stay at least for the next few months. I Think the most interesting question is whether or not the open versus closed gap is going to continue shrinking and if open will ever fully match closed source true frontier level parity. Curious what your thoughts are there, Jordan. I think it has some serious implications for our whole industry if it actually happened, obviously.
A
Yeah. I mean, my view is that I believe the reason that this gap has closed right now is squarely put on the US Government imposing restrictions on anthropic and resulting in us not getting the actual best models that these guys have. And so they've artificially caught up, basically.
B
Interesting.
A
Clearly we see this with, with Mythos versus Fable. I can't use Mythos. I can only use Fable sometimes if I ask it nicely. 5.6 SOL. I think our house view is that it's not the biggest model OpenAI has ever trained, is not the size of 4.5. To me, they have a bigger model somewhere. And I think the result is that we're going to be able to only access frontier intelligence if government entities allow us to. And that's a very interesting change to the setup going forward because I think it represents an opportunity for many of the players that are in fourth, fifth, sixth, seventh place to catch up to a limit, at which point it's okay to release everything and start to battle for user share without really being able to find the frontiers and have the frontier dominate. I think it's possible that, you know, we see the frontier take another big, big step towards the end of the summer. It's possible the politics change a little bit. Yeah, it's possible that we start to find other, you know, modalities beyond coding at which these guys can really improve and they start exploring those like, areas. We didn't, you know, intend to talk about this right away, but I just loved the release of Inkling by Thinking Machines. Thought that the native audio input would be super interesting, super useful in the future and kind of a sign of what's to come. But anyway, yeah, I, there's.
B
On the topic of Inkling, there's definitely like huge, huge, huge demand for a Western open source model that does not suck. Like, it is. Like, I'm, I'm shocked how markets are still so inefficient and like we haven't had a single American company that's at least on par with like the fifth best Chinese company. But if you just like, I mean one, it's probably just a matter of time until the US government bans Chinese open source models entirely. And maybe that's a can of warrants you don't have to go down on this conversation. But two, even if that doesn't happen, I feel like the average large American enterprise is simply unwilling to put all of their proprietary data through a Chinese open source model. Even though you can make sort of like all the logical arguments of like, dude, you're just loading their weights in your air gap data center or whatever, there's no way the CCP is actually going to see any of your data. But I don't think the executives will actually buy that and I don't think they really care. And I think there are a lot of people who A care about token budgeting and B are only interested in running a Western model or a non Chinese model. And so it is shocking to me that I guess Inkling is the best one that we have now, but it's shocking to me that we're not like actually closer to the open source frontier in America.
A
Yeah, I mean it was Nemotron and then it was Inkling and yeah, it's really inspiring to see Inkling go for it. I think they have two business opportunities there. They've got to be better than the bulk of Chinese open source. Like they have to be in the game there to be considered. But then they also need to be better than Sonnet or better than Terra Luna, the tier 2 tier 3 models from the Frontier labs, because you know, you can build a bunch of cheap applications using close to frontier intelligence, using Bedrock or Foundry or whatever and just get access to the anthropic or OpenAI models and save your money there by going with their second best model. So I never really understand the Western open source angle of saving people money. I think it is real and getting those models into the ecosystem of companies like Fireworks and together and base 10. And anybody who's serving open source is a good thing because it is a market. But to me the bulk of the market is government. One of the views on the Chinese models, actually that's interesting, is that Xi has been encouraging the Chinese companies to keep the models open source. That is the view from their party. And I think the big reason for that is that a bunch of the Chinese government wants to download the weights and run it on servers that they own. And they want the support of the local ecosystem. And I think that the American government should work the exact same way. I think that's a pretty pragmatic strategy. It's like you need to give the people in your country access and support to run this stuff. Maybe the other thing worth commenting on is that in the K3 blog, they mentioned the post training, sorry, the quantization during the SFT stage. Right. And in that they were commenting on natively using MXFP4 and MXFP8 weights and activations respectively for quote, broad hardware compatibility.
B
What other hardware do you think Moonshot cares about? Jordan?
A
I got a list of 11 Chinese accelerators. You should subscribe to the semi analysis accelerator model and learn more. But yeah, Huawei, Ascend, you've got Baidu, you've got Kun Lungson, you've got the More threads. Guys, there's all sorts of different chips that are being, they're showing up in papers. We're seeing code like it is a national priority for China to get these frontier models. These are frontier models now running on their domestic accelerators.
B
Yeah, I mean if, I guess if we were calling Google a frontier lab, you know, end of 2025, we got to call. They call Moonshot a frontier lab now. Kind of crazy, dude, it's like vanity,
A
you know, you're gonna. I'm still, I'm still a 34 waist.
B
Yeah, yeah.
A
And so are the seven other Chinese labs.
B
Yeah. No, no, funny enough, my, my dad is actually visiting China right now and he's telling me how like the hotel he's currently staying at is totally booked because Xi Jinping is going to be in the area soon and he's going to give a speech about how AI is a top priority for China. So I think a lot of what you said is right, circling back to what you said earlier about the US government and how if they keep knee tapping the frontier models OpenAI anthropic halves, forcing them to delay them, forcing them to only have their second best model actually publicly label and therefore giving all the other players, The Googles, the SpaceX metas, whoever, time to catch up. Do you think that just completely destroys the Frontier Lab business model? If you're open eyed anthropic, you just lose all pricing power at that point. Right. I don't see how Anthropic can still accelerate net new ARR if their model is on par or comparable with the meta model, the SpaceX model, the Google model, the Moonshot model, the deep seq model. Like what happens to our industry at that point, Jordan?
A
Yeah, I mean first of all, no, I don't think that's going to happen and I think I can explain why. But second of all, I, I don't, I don't know for sure. So we'll have to see it play out. Interesting to think about. So I think the, the biggest thing that I've realized in my personal usage of this stuff is one, how hard it's getting to differentiate between using the Absolute Frontier model and the Max thinking mode versus high versus medium effort on those models. It's really, really hard for me to find day to day tasks that these models can't figure out. And I'm just. My behavior is default to the biggest and hardest thinking because I don't care about Dylan's budget. But yeah, when it comes to actually using this, there is an aspect of the harness being part of the product. So testing Kimik3 requires me to take a serious look at open code and Hermes and PI. And so the harness is totally part of the product. Still, simple, simple things can cause me to want to use one model over the other. Like, can I install it on my remote SSH server? How easy are the keystrokes to get stuff in? Can I edit previous commands? I mean, you know, it's, it's silly, but these, these little tiny features in the harness actually impact where I'm going to send my tokens, which results in where I'm going to send my budget. Right?
B
Um, yeah, that's an interesting point because I think a lot of people, they talk about token budgeting, right? But I think, and correct me if I'm wrong, but what I'm hearing from your description of your own workflow is that like, even for tasks where I'm pretty confident that like a GLM could successfully do it, like, I'm happy routing it to Fable, doing it on Max Intelligence because like the ROI of that task is still worth like the Fable price to me. And so, and then like, there's obviously there's like always going to be some risk in the back of your mind where it's like, oh, if I, you know, use GLM instead, even on like on Medium thinking mode, it's way cheaper. Like maybe it's not actually like as high quality as Fable actually would have been. Right? And so even if the benchmarks claim that like, oh, a lot of your tasks can move to glm, you're still fine, like keeping them on anthropic models or OpenAI models for the foreseeable future.
A
Mostly, yes. But I use a lot of Slack bots right now and I actually don't know what model's running behind the scenes on that Slack bot. So specifically at computer with perplexity. And I think if it starts routing it to Kimik 3, if it starts routing it to GLM, starts routing it to Sonet, and I know it's doing it today because I looked at my usage a few weeks ago and found how much of the OpenAI models I was using because it was making that decision first cut at a PR before I go in and actually fix some stuff up. I don't really care which model they're using. Right. And that is about the quality of the harness there for what I'm using. The justification, this might actually be a,
B
this might be a pitch for outcome based pricing. If anything, one of these labs could potentially just get 95% plus margins if they do outcome based pricing for you. Because as you said, all these tasks you're happy to pay even stable pricing for when they could probably get it done at a fraction of the price even today.
A
Yeah, a hundred percent. Certainly with the dialing in the thinking mode, which is where a ton of the expense ends up going, I can totally imagine them building a router. The second thing though, just on the competition thing you said earlier is like, I don't think that we're out of use cases or ideas for these guys to work on. I don't think, I think they can continue to train incredible models to try and hit RSI on the coding side without ever releasing it to us in the proletariat and keep their permanent underclass. Yeah, keep their bourgeoisie models training each other and keep distilling them and you know, whatever, giving us the little tastes of it while still pursuing a research objective that includes all sorts of other uses of AI. Like we're really exploring coding right now, but we do some video generation stuff. We do a lot of audio to audio stuff. We do lots of like deep research that really doesn't look like coding in some ways. I think there's lots of use cases that they can continue to explore without really encountering like the cybersecurity issues. Robotics and world models is a simple one. Right. What if Anthropic sets their sights on automating away a whole bunch of manual labor jobs instead of knowledge work jobs. I mean the idea that there's no way for them to build a sustainable business with great ROI for their, you know, greatest technology the world's ever seen. I don't believe that at all.
B
Yeah, it doesn't pass the smell tests.
A
No, not at all. But even beyond that, the. I use these models so much every day and I see, first of all how much my friends who work in technology who are Software engineers spend 10 times less than me and use it 10 times less right now. Right. One person using Fable is like, you know, a person using Sonnet and a person using Fable, they use it both the same amount the same day. The person using Fable spends 10 times more at 90% margins. They make up the bulk of the business. Right? As soon as those people, of which I would say there's maybe, maybe I'm in the top 10%, maybe even the 1% of the industry start using the bigger bottles, they use them more. That's just more demand for all of this business. And the models don't even need to get any better for them to discover they can use them for the really important tasks or the bigger ideas that they have. And then I need to go and talk to my neighbors who don't work in technology. And there I'm certainly in the 1%, probably in the 0.1%, maybe 0.01%, and maybe we've got a thousand times more to go from here. And so I, I, I get back to, you know, Masa Son, golden goose exponential chart to the right. You know, sort of this point of just like it's an exponential. Hop on. Why are you okay, just to take this totally off, off the rails, but.
B
Well, actually, before we go there, I want to say that I think what she said about being early is totally right. And this is exactly why I don't think Kim UK3 is going to cause net new ARR at anthropic and OpenAI to decelerate. Because I think even if you want to assume that some non nudgeable portion of people who are using fable and five stick sold today are going to switch to like Temi K3 because it's cheaper and can do their workloads, I think that is like completely overwhelmed by the people who still haven't like seriously tried this technology. You know, the people who've like kind of tried it a little bit, but are every day like discovering new use cases, like new cool things, high roi, things they can do with the models, I think all those people are going to be using like 56 SOL or Fable 5 as a default to unlock these new use cases. And that is like you're just not going to see a ARR slowdown, ARR growth slowdown, because that is going to continue ramping so fast.
A
Yeah, I think we're in agreement on that. Think about how many people there are left to subscribe to this podcast and follow semianalysis.
B
Man, dude, it's crazy to me. So I went to ICML last week. I went to Air Engineer the week before that. These are like, you know, nominally AI conferences, right? I thought people would be pretty plugged in there. Like I would say 80 plus percent of people had never heard of semi analysis before. I was like, guys, what are we doing, man? Like, you claim to work in AI, but you haven't, you haven't read and you've never even heard of semi analysis. Like, we're still so early.
A
It's, that's a, that's an ego check. Max, come on, man, calm it down.
B
I mean, maybe tune our own horn a little bit.
A
Okay, man, I, I, I think we could keep talking about this all day, but probably good to wrap here. Anything you think was left upset, left unsaid. Any burning questions?
B
I, I just if the stock market crashes because all the investors, you know, have deep seq R1 movement part 2. Buy the stocks, guys. Not investment advice, though. Do your, do your own diligence, not invest in advice.
A
Love it. Let's end it on that clip. Clip out, Max. Saying anything about stocks and finish what you say. Do your own due diligence. Good job, man.
B
Okay, yeah, cool.
Hosts: Jordan Nanos, Doug O'Laughlin (with guest Max)
Episode Date: July 18, 2026
Main Theme:
A deep dive into Moonshot’s new Kimi K3 model, its position among global AI models, implications for the semiconductor and AI landscape, open-source trends, pricing, and broader industry and geopolitical impacts.
This episode focuses on the surprise release of Moonshot's Kimi K3, a highly competitive large language model (LLM) from China. The hosts, joined by Max, analyze its technical achievements, market impact, architecture, open-source strategy, and what it means for the broader AI industry—especially in relation to U.S.-based labs like OpenAI, Anthropic, and Google. They discuss usage experiences, pricing, hardware requirements, and predictions for the future landscape of AI model development.
Timestamp: 00:11–02:25
Timestamp: 02:25–04:25
Timestamp: 05:30–06:48
Timestamp: 06:48–08:28
Timestamp: 08:28–11:10
Timestamp: 11:10–14:08
Timestamp: 14:08–16:10
Timestamp: 16:10–19:55
Timestamp: 20:54–27:24
Timestamp: 27:24–31:34
| Timestamp | Speaker | Quote | |-----------|---------|-------| | 00:36 | Max | "There's like a pretty clear top three with Fable, Sol 5.6 and now Kimik3." | | 02:55 | Jordan | "I can't really find a lot of complicated stuff that it can't do, which in and of itself is a bit of a feat." | | 07:08 | Max | "If that's actually true, guys, it's time to pack up the bags." | | 10:20 | Max | “Who is the user that is actually going to be switching towards Kimik3?... It wouldn't surprise me at all if there's not actually serious adoption among large enterprises." | | 14:08 | Jordan | "I believe the reason that this gap has closed right now is squarely put on the US Government imposing restrictions..." | | 17:12 | Max | "It is shocking to me that we’re not actually closer to the open source frontier in America." | | 22:30 | Jordan | "The harness is totally part of the product." | | 25:31 | Max | "This might be a pitch for outcome based pricing...they could potentially just get 95% plus margins." | | 27:47 | Jordan | "There are maybe a thousand times more to go from here." |
Final Advice:
"Buy the stocks, guys. Not investment advice, though. Do your own diligence, not investment advice." — Max (31:10)
For more on accelerators and hardware, subscribe to the SemiAnalysis Accelerator report.