
Loading summary
A
Hello and welcome to SED News. As I think many of you know by now, this is the monthly format of Software Engineering Daily, where we dive into the tech headlines, we go into a deeper topic in the middle, and then we just do a spin around our favorites from Hacker News. Highlights and spoiler. There's also a fun one that didn't appear on Hacker News, but we'll get to that at the end as well. But yeah, as usual, I think I saw you last, Sean, as opposed to spoke to you, saw you in Singapore, which was fun. Yeah. You've been traveling as you often are. But traveling in my neck of the woods was fun to see you there.
B
Yeah, that was great. It was my first trip there, so it was great to hang out. And that's two times in a few months. This might become a regular thing. We might just have to start doing these in person.
A
Yeah. And spoiler, I'm coming back to SF in a couple months, so, yeah, it's good.
B
And I'm off to not quite Singapore, but I am off to Australia here in the next week.
A
Nice. But, yeah, what else has been keeping you busy over the month?
B
I mean, summer, I feel like, has just flown by. Like, I feel like my kids were out of school and then suddenly it's like August and they're going to be back to school in a few weeks. So things have become really hot and fast this summer. I think we've done a lot of traveling and then just stuff moves quickly. But been a big summer for my. I usually don't talk that much about my family, but big summer for my son who learned to swim, learned to ride a bike, and also now has gotten significantly better at Reading. So I'm very proud of the amount of work that he's put into this summer, too, learning new skills. How about you?
A
That's like a whole model change there for your son.
B
Yeah, it's the new Kimmy model.
A
Yeah. On my side, just. Yeah, a bit of traveling. I'm recording this on my side from Scotland. Scotland. I try and come back here a couple times a year. So, yeah, nice to get out of the city. I'm up in the Highlands of Scotland, so lots of nature around. Yeah. As I was flying, just back to something we've talked about a few times, Starlink on planes. But yeah, that used to be super seamless with no login screens. And this was, I believe, dictated by Starlink. But they've apparently had to cave in to airlines wanting to put like login interstitial for that. So, yeah, I mean, it's such a trivial thing, but expecting Starlink to just connect. And then suddenly I found the airline saying, hey, if you need to log in with your membership number, And I thought, oh, what's happened here? And I read up on it. Yeah.
B
Do you still like, do you have to be up in the air or does it work immediately?
A
That's a good question. For some reason I don't think it was working kind of gate to gate. On the flight I was on, this was Qatar Airlines again, but I think it's still supposed to. But yeah, still you just have to log in with your Qatar membership number now, which you don't need to be a special tier or anything, you just literally have a.
B
It seems like United is doing some experimentation, but I think I've only had like one or two times. I got like a text message saying that there's Starlink on this flight. And then I think both times they ended up having problems with the plane and having to switch planes and then we lost Starlink.
A
Yeah, that's the one thing about casar. They really rolled it out like big time across most of their fleet now. So if you're on a Airbus A350, you know you'll have it, which is always nice, but right out of plane chats, which I can easily head into at any time, and onto software. So onto the headlines. Yeah, so there's been a sort of flurry of we're just sort of calling it runaway AI stories over the last couple days. And the first one is actually the most recent, so that's just disclosed today or yesterday, which was Claude. And this is actually in reaction to OpenAI, which we'll also talk about, but basically in reaction to something that OpenAI disclosed. Anthropic has now also disclosed that it's that Claude hacked into three organizations whilst they were just testing their cyber capabilities, or at least that's what they said. And they basically said that Claude had gained unauthorized access to outside companies during an evaluation of its cyber offensive tasks. And it was a misunderstanding, apparently, that Claude had access to the Internet in its testing environment. We're going to see this as a kind of theme here of like, is this really a runaway model or is it just a human who forgot to do something? But yeah, I mean, this one feels like human just forgot to place it under strict no Internet access controls. But yeah. What did you think of this one?
B
Yeah, I mean, I think that all this really ends up coming back to some human decision making.
C
Right.
B
Like that. Perhaps the intention was not to have the model, like hack something, but essentially something got lost on the human side of we made a mistake here and gave it access to the Internet, or we made a mistake here or whatever. The issue is, it really comes back to some human level of control and what guardrails is in place. And then the thing, though, with all this that has me thinking from both the OpenAI attack on hugging face and then also this latest one with Claude, is that if you can just apologize and say, oh, it was unintentional, it was an accident, and stuff like that, when this stuff happens where there was intention for abuse or misuse or something like that, like, does it give people basically license to excuse the fact that they are doing something potentially malicious and just say, oh, it was an accident? I mean, we've seen that with viruses, too. I remember one of the viruses. I can't remember if it was one of the. Back in the early 2000s, there was all these viruses that were like these scripts that went around in emails, and then you click on it and we copy your contacts and then fire off to them, and it would take down people's mail servers. And I believe, like, one of those started as a student who was testing something in an MIT lab, and then it accidentally got out. And I don't know what happened to that student, but this is certainly not something new. But I just wonder where does it kind of stop with being able to provide an excuse of this was an accident versus there was intent behind it?
A
Yeah. And just to sort of give context for anyone who hasn't been keeping up with the news, the OpenAI one was the fact that basically, yeah, it hacked Hugging face. They set it off to do something, and it ended up repeatedly trying to get into hugging face. But what was the secondary part to that story, courtesy of TechCrunch, was the fact that actually Hugging Face could have detected this earlier, but they had a sort of human failure on the security
B
end that, yeah, they caught it, but they didn't escalate it to a human fast enough.
A
Right, exactly.
B
So their systems actually worked the way it is. So in some ways, there's a part of this that is not really an AI problem. It's kind of the boring problem of a human wasn't in the loop at the right time.
A
Yeah. And I think that's the piece it's hit mainstream headlines as, like, AI is. This is exactly what we all were saying, that AI can be catastrophic. It can go off. It's got a mind of its own. It's hacking Things left, right, center. But especially in Both cases, the OpenAI and Anthropic cases, someone, a human did set it up to do something along these lines. But the human did not a keep a watch on just the repeated actions it was taking. And clearly there was no fail safe for it to stop at any time. Again, these are all human determined things that can be set, but it's not like it broke out of the cage. Exactly. And it seems there's some quite fundamental things that could have been done by a human that just weren't. And that's actually, yeah, it's a quote boring story unfortunately, because AI isn't sort of this wild animal quite yet.
B
Yeah, I mean you want to sensationalize the headlines, right? Like even my sister, who's not in tech at all sent me the headline and was like this is scary. So that you get the attention there. But I think one of the things that was interesting about the hugging face attack was when they tried to investigate, they couldn't actually use CLAUDE or GPT because those models have safety guardrails in place where you can't tell whether are you an incident responder or an attacker. So they have essentially mechanisms in place that like I can't go to CLAUDE and tell say like help me break into the Pentagon or something like that. Like it's going to prevent me from doing that. That means that you also are limited in using those models to investigate. So they ended up using an open weight model out of China for the forensics. And then I think that also we're going to talk a lot about the open weight models from China later in the episode. But that's interesting consequence of this is that we have these guardrails in place. But there's always ways to I think manipulate the models even with the guardrails in place to do stuff. But then when you want to use it for something intentional incident response, you might not be able to do that because the protections are in place there in the first place. And then you have to circumvent it by going to a model where maybe there's less safety guardrails in place.
A
Exactly. And the sort of final one on the runaway story is actually Amazon not security related, but cost related. We have touched on this the last couple of SED News episodes. Just where's the tipping point of cost overruns making it not viable for businesses to be allowing employees to sort of quote token max and this kind of thing. But yeah, Amazon basically said that they had a ton of unplanned spend. They called it catastrophically expensive. This is all courtesy of the FT Financial Times. Apparently they had about 860% budget overrun over five months and this was basically in their word, caused by bad agent loops. Just didn't crash loudly enough and they've just kept being billed. So yeah, I mean it's pretty interesting for Amazon to come out and actually say that.
B
I just don't understand how they could be surprised by this. I feel like we've been beating this drum for months now and I think this is just the beginning of these kind of stories that we see. But if you build a leaderboard to encourage people to use AI and that's the metric you're optimizing for, but there's no connection to the value of the use of that AI, what do you think is going to happen? This is really good heart smoth essentially showing up on some sort of schedule. You're rewarding people for tokens consumption. So it's like you're giving people a license to be wasteful and not looking at the productivity metrics of those. So that's like the epitome of token maxing.
A
So yeah, they kind of have themselves to blame in that sense.
B
But yeah, it's kind of similar to the initial stories we were talking about with hugging faces. It's still not AI necessarily just running rampant. It's a human decision that in both cases, even though one is this attack vector and the other is spend, but it's still a person making the decision at Amazon to say, hey, we're going to just have a KPI where we're just going to reward people for maxing out tokens. The other thing that's interesting about this that people have to think about is that with traditional software, when you have a bad for loop or code that runs forever, whatever it is that ends up usually resulting in a crash or setting off some sort of alarm, and you typically know fairly immediately that something's gone wrong. And I think the challenge with things like agents and so forth is that you can have a bad agentic loop, but it doesn't result in a crash. What it really results in. Just keep calling that model over and over again and trying to make adjustments and then giving you sort of plausible outputs or updates, but the whole time you're getting built. So even outside of the wasteful like token use of maybe me using my company's token budget to do my grocery
A
shop list for the next your laundry, that'd be great.
B
Yeah, that'd be great to do my laundry. But then there's also legitimate excess spend where you might just end up having your agentic harness spin for some period of time where it's just churning against tokens and you don't even know that something wrong is happening. And that kind of goes back to the earlier stories as well of just we ultimately need a lot more observability into what is happening. Presumably someone at OpenAI, if they hadn't been really paying attention to what was going on in this experiment, would have saw that this agent with the right observability tools in place is hammering hugging face and trying the same and that should set off certain alarms. And I think it's similar in this case where if you have an agent loop that's out of control and spending excess tokens, there should reasonably be some guardrails. I mean you have that with other services like if you spin up a particular elastic instance or something in the cloud, you're typically setting your top line provisioning of those types of things and you have some controls over it. It just feels like we haven't thought through all those controls for AI right now. And I guess part of it's just this race to try to out compete everybody and everybody kind of feeling like they're behind.
A
Yeah, I'm going to just jump ahead for a second on. I won't say what it is because that's a spoiler, but one of the things I'm going to bring up, Hacker News highlights. It's a very reliable source as you'll find out at the end. But basically this was someone who'd done a bunch of stuff with Claude and we'll get to that. But he points out towards the end that Claude doesn't provide reliable methods of counting tokens, despite live showing token counts, reporting token consumed per sessions and billing for tokens. And he said, but I'm sure this is temporary and this will be fixed. It's just crazy that we do actually have a system at the moment where you literally just don't know what is happening and like exactly what it's going to cost and why. And as we're going to get into on the open weight side of things like this is really feeding into the rise of open weight models as well.
B
You think of a lot of stuff in cloud or even what we saw with ride sharing where they give you somewhat like a prediction model of what the spend will be for certain actions. So it's like okay, well I want to go from here to here in Uber or Lyft and they'll be like oh, that's going to probably cost you X number of dollars so you have some visibility into what the cost would be and and you can do similar things with certain cloud calculators and stuff and it can get kind of do we need that for AI? If I'm saying create an engineering plan for some sort of feature, can I get an estimate of the budget required to do that and then based on what that budget is, maybe try adjusting the plan or something like that to try to optimize it down?
A
Yeah, that's interesting thinking. Sort of. Could you effectively put in your ask your prompt and I say prompting, that almost sounds like a year ago or something. We're talking about spinning up agents, et cetera. But yeah, try and get some kind of estimate before it sets off slightly digressing. So let me get us back to the headlines, which the next one we have is just the fact that Microsoft has sort of gone back on a bit of a tear when it comes to its valuation, which is interesting. It's like one of the largest jump of shares or sorry for their fourth biggest jump on record. So it's up 16% and we don't often cover just like pure financial news of tech companies. But to see Microsoft making these strides is pretty interesting. Some people might then think oh well this is partly to do with OpenAI, but it does own still a quarter stake of OpenAI and claimed that that had contributed $24 billion of revenue, which was about 7% of the $332 billion in sales it reported. But a lot of it was really just AI driven revenue and massive investment in data centers. But actually that's been completely in theory vindicated by the amount of revenue they're also making.
C
You're building agents that can write code, summarize documents and automate workflows, but they're missing one awareness of the world around them. XWeather combines enterprise grade weather intelligence with agent ready APIs, natural language capabilities and an MCP server built for tools like Claude Codex, Copilot and modern ides so your agents can adapt workflows, automate responses and make better decisions based on real world conditions. Backed by Vaisala, whose instruments fly on NASA missions to Mars, Exweather delivers trusted data and unique insights that go beyond conditions to actual impact. From real time lightning strikes to road surface forecasts. Start with 15,000 free API calls every month and pay only for what you use as you grow your full weather stack for developers by developers start building for free today@xweather.com
A
Think about your mobile app's source code. Once it hits the app store, it's out in the wild and without the right protection. Decompiling is easy for malicious actors looking to steal your IP or tamper with your software. That's where Guard Square comes in. Guard Square provides the highest level of mobile app security for Android and iOS applications and SDKs. Their advanced tools integrate seamlessly into your CI CD pipeline. We're talking polymorphic multilayered code hardening techniques and automated runtime application self protection paired with mobile application security testing and real time threat monitoring to deliver the highest level of mobile app security without compromise. Don't leave your hard work exposed. Secure your mobile applications today. Go to guardsquare.com to learn more.
C
You're shipping faster than ever with AI coding agents, but those agents don't vet the packages they pull in and they don't have security context built in ORI by Endor Labs fixes that it plugs directly into your editor via mcp, catching vulnerabilities, blocking malicious packages and flagging exposed secrets in real time. No separate tool to switch to, no dashboard to babysit security that fits how you actually build teams using Ori. See 10 times fewer security tickets and 6 times faster fixes. Free for developers. Get started at www.endorlabs.com Auri I think
B
if you look at also Nadala's quotes related to the announcement of their quarter performance and so forth and also the recent thing that we covered also in the last episode where on X he had written about how the AI companies or the model companies are charging you twice and so forth. In here he says that every model is substitutable. He's telling essentially investors that Microsoft is deliberately building its infrastructure so it can swap out things like OpenAI for its own models or for anthropic. And I think that it seems like they're moving towards a view where they're trying to allow essentially their customers to be very flexible and adapt, which I think makes sense. I think the average enterprise now is using at least five different models and you probably want to be able to do that so that it's a little bit like being hybrid cloud, although it's easier to be hybrid model where you get power essentially in the negotiations if you're not wholly dependent on a single vendor. And it seems like I think Microsoft's kind of leaning in that way. But I remember back in March Microsoft stock dipped and then everyone was like freaking out and essentially calling for the death of Microsoft. And now it's back and I just think that the overall the market is extremely volatile right now. There's these huge swings constantly from. I mean IBM had its biggest drop recently, biggest single day drop in like 50 years or something like that recently and that was coming off like a huge pop just two months earlier. So I don't know what goes up. Must come down. So who knows, we might be talking about Microsoft in another quarter or two, how they had the largest single day loss in a day or something like that.
A
Yeah, for sure. It is quite hard to predict as markets should be, I guess. But yeah, we're seeing just huge swings when it comes to AI related stocks when one minute chip makers up, one minute chip makers down. And just on a tangent there, yeah, Apple briefly hit 5 trillion valuation which is pretty insane. But just before recording I double checked and they had actually reported numbers very recently in the last I think couple of hours and they've gone back to 4.9 trillion. So I don't feel too sorry for them. But only 4.9? Yeah, only 4.9. But interesting to see that they still notched above 5 and yeah, what's driving their revenue? Well, still they've got very good iPhone sales and they've got still very impressive services revenue and even greater China revenue as well. But both of those services in greater China were a little bit less than what was expected, which is again just what's kind of driven that. But very interesting to see that they can still, we've talked about this on previous SED news, still manage to stick on these lines of business that are not that AI driven or even really AI adjacent to be honest. Very interesting. But they are having issues with memory chips, which is another sort of topic. But even in Supabase we've been told by one of our suppliers, we have different suppliers depending on where you live to get laptops. And some of our new employees are being told weeks before their MacBook will arrive because we do custom specs, not just like off the shelf. But yeah, now we're being told weeks, which is quite exceptional.
B
Oh wow. Yeah, I just got a new laptop and this is my first time recording on this. But it's interesting, like with Apple too, a lot of their lines of business which are like incredibly successful, they're still like minority in a particular vertical. Like iPhone is wildly successful, but it's not the most dominant phone I guess maybe from a single vendor, but Samsung might actually be bigger. But then obviously from an operating system standpoint like Android is, there's more Android devices than there are iOS devices similar with Computers as well. I haven't used a PC in a very long time, but it's still the dominant machine overall, I guess. But I think they have incredible brand loyalty and they do make fantastic machines overall. So clearly the things that they're doing is working for them.
A
Yeah, absolutely. So then moving on to. We've also covered Waymo and Driverless from a few different angles, but this one's interesting because we did talk about Waymo and Uber being partners in ways and then obviously frenemies in other ways. And yeah, we're now actually seeing an official split of that partnership. So if you've used one in San Francisco, this might sound confusing because it's a Waymo app and it's a Waymo car. So where does Uber figure in that? But actually this is for some of their other US territories. So Waymo had like first partnered with uber apparently in May 2023, and that was to launch in Phoenix. Then it was followed by Austin, Atlanta, and that's like, so Waymo cars available through the Uber app. And then Uber managed the vehicle fleet apparently with another partner called avomo. But then in May, Uber and Waymo parted ways in Phoenix, and then the two companies have also clashed over the quality of Waymo services, apparently in Austin and Atlanta. So, yeah, it's kind of interesting to see they tried, but I think it was always going to be challenging to see how Waymo being owned by Google, how is this actually going to net out at the end? Could these two actually be true partners? Or was this just always going to be, again, a tipping point of where that partnership just kind of had to end?
B
Yeah, I mean, they say all partners are meant to be broken at some point, especially where they're both in ride sharing. Clearly at some point their mutual interests are going to become too competitive to each other, essentially, in a lot of ways, I think very similar to how a lot of partnerships work, where Uber was essentially supplying demand and operations while Waymo was weak on those particular spots. And then as Waymo has grown and raised more money, essentially they don't need those training wheels anymore. They can build out their own and build their own network and own it end to end. So I feel like this was probably always going to be the ultimate end of that relationship.
A
Yeah. So, yeah, it's sort of not officially over yet, but just people familiar with the matter apparently, again via Financial Times. But yeah, it seems pretty unsurprising, really, that this was going to probably break apart at some point. And then, yeah, just to wrap up on the headlines, it was an acquisition and we were just talking before we started recording. So let me try and get this right. I believe it's Any Scale was acquired by N Scale and when we were talking earlier I said is that scale? And it's like, no, that's not Scale. AI is different again. So we've got amazing naming these days. But yeah, what's this one about?
B
Yeah, so Anyscale, which is known, they were the creators of Ray, which a lot of inference infrastructure and fine tuning and model infrastructure runs on very well established open source project. Anyscale was the company that tried to build or built essentially like a managed version of Ray around that and they raised like a billion dollars so in 2022, but I think. And then they just sold to n scale for 1.65 billion, which is a Neo cloud. So for those that aren't familiar with NEO Cloud, Neo Cloud is essentially what the term is used for companies that offer primarily GPU as a service versus kind of general computing. So there's all kinds of these Neo cloud companies are now available. I had actually talked to Anyscale at a variety of different times. Like I think this was probably like a decent outcome for them because I think it's probably hard to really grow that managed Ray as an independent company into like a really big company. But as part of a GPU infrastructure company, it's probably a good pairing. So I think that it makes a lot of sense. But I do like, I always confuse any scale with scale AI and then the event scale and I don't know the history of how those company names came together, but generally the thing that people a lot of times strive for with naming companies is you want a name that you can say and people can remember and I'm not sure they hit the mark there where it's like N scale, Scale, any scale. It's hard to remember who's who in that Venn diagram for sure.
A
I mean scale AI, okay, that's a great name to have AI. You can basically have anything that is your company and it's going to sound good and it's five letters but then any scale and then N scale, that's pretty confusing. And then yeah, just a sidebar piece of news, Scale AI now have a new CEO as well, which is interesting because that was founded led by Alexander Wang, not the fashion designer, if anyone knows that one. But this is Alexander without an E at the end of the Alexander and Meta had taken us virtually just under 50% stake. That was kind of the point. So Skale and Theory are still running their own show, but when you've got, I think 49% stake from Meta, you're quite beholden to them. But yeah, in that transaction, Alexander went to head up the AI side of Meta Total. Now Skale have a new CEO, so that'll be interesting to see how that all nets out.
B
Yeah, they raised like 14 billion round, led by Meta. So pretty significant.
A
Yeah, I mean, actually just the new CEO of Scale is actually a former Google Cloud executive called Francis de Souza. So definitely interested to follow along with that one and see sort of how that all nets out.
B
Yeah, I saw also like Fireworks who just raised a huge round and made a lot of news. Their new head of engineering just came over from Google. So I think you're starting to see it's a common pattern though. You have people who reach executive positions at large companies like Google and they get maybe a little bored with that and then want to go back to building and moving faster.
A
Yeah, for sure, yeah. Fireworks. Super interesting. We do have an episode with Fireworks with one of the co founders, Benny Chen. So, yeah, go check that out. I think that came out around March this year, so well before this fundraising was confirmed. But yeah, they're definitely having a moment. They are sort of the infra for open weight models and that's becoming increasingly interesting to many companies for many reasons. But yeah, that's actually quite a nice segue into the main topic for today, which is we're kind of calling it the Kimi moment. Kimi, which is an open weight model, the arch company is called Moonshot AI. So if you've heard of Moonshot AI, that's Kimi and vice versa. From a Chinese. Moonshot AI is a Chinese original company. They do have offices, I believe in the Valley and in Singapore and that kind of thing, but very much seen as a Chinese company, which sort of frames a lot of why this is quite interesting. But yeah, when we think of open weight models, there was this deep seek moment first which we can sort of touch on. And then really in the last almost like three months, we've just seen this sort of especially rapid adoption by many companies of Kimi and really forcing companies to sort of think differently about using foundational models from the big players or the big names rather, and getting some quite interesting and very high quality results out of Kimi. The Deep Seq moment, I think. Like Sean, how did that sort of kick things off when we think about open weight?
B
Yeah, I mean, I think the Deep Seq moment ended up having sort of more Market impact in some sense, because I think prior to that, the sense was that you couldn't really get these really powerful models except from these frontier labs like anthropic and Google, OpenAI and so on. And so it was kind of shocking when Deepsea came out and Nvidia's market cap cratered as a result of that. And of course it's certainly bounced back since then. And then when you look at the Kimi launch, which is the largest open weight model ever released, beating some of the top closed models on particular benchmarks, it seemed like nobody really panicked. And I think part of that is not because it's less impressive, but is because we've kind of gotten used to the idea that these Chinese lab open weight models are competitive. And so it's less of a shock. Essentially, we've been desensitized to it. But I do think that what the market reaction has been more around and I think this is something we're already coming to where this was somewhat on the back of what happened with Mythos and Fable and the reaction to that, where people got scared that they could invest in a model and then have potentially a foreign government say, you can't use this model anymore, then I think that has increased the interest in diversification of models and then also having an open weight model strategy, along with having a sort of closed weight model strategy.
A
Yeah, and we'll probably touch on this sort of the banning of models because of course it comes into this as well, if you want to say ironically, because yeah, when you ban at least pause banner foundational model, a closed source model if you like, then open wait becomes very interesting. But then a lot of open weight comes from China, so we're going to kind of be interested in that. Why sort of the last couple of months has Kimi really just exploded? I think in interest and popularity, it's kind of moved from oh, it's good enough for professionals to this kind of actually competes with GPT and Claude and it's kind of closed that gap in basically weeks, at least From Port has been sort of released.
B
Yeah. And they also released a ton of models back to back. Essentially, they're moving super, super fast.
A
Yeah, exactly. So I mean, yeah, if you sort of look at the actual lineage, I guess they've been doing sort of quarterly drops, if you like. July 2025 was K2 and then within 12 months you've gone through 2.5, 2.6, 2.7 and now K3. That's an incredibly rapid cadence. If you look at Those jumps as being as significant as jumps that you might see from any of the foundational players, but definitely not on that cadence effectively.
B
Yeah, I think the model distillation practice has really sped up how quickly people are coming up with models and models that are comparable in performance, which was a big part of the conversation around the initial Deepseek launch. I do think that part of some of the criticism of some of the benchmark results that we've seen from Kimmy is that in particular there was a lot of headlines around their MCP tool calling performance but that beating Opus, for example. But in particular that benchmark is relatively new and there is. I don't know if this is more just sort of like jealousy in the dialogue or whether there's some truth to this, but you can essentially sort of bias towards performing really well on certain benchmarks, whether it's the MCP1 or it's the MLLU style benchmarks as well. And then you can have really good benchmark performance, but it might not actually match reality. That's where some of the criticism has been on even some of the tests of testing a model against LSTAT and stuff like that, and then saying, oh, it outperformed lawyers on the LSAT. But then the LSAT's actually not necessarily a good indication of what a lawyer does on a day to day basis. So these are some of the nuance, some of this is general, you know, model criticism. But this is some of the dialogue that I've heard around Kimi in particular is like, are they building to optimize for the metric? Kind of like what we talked about in the Amazon story is like, you know, whatever the KPI is, people are going to try to optimize for that KPI. So you always have to be careful essentially what the metric is that you're measuring people against.
A
Yeah. And on June 12, K2.7 code was released and this is obviously very much a software engineering tuned model. And as you were saying. Yeah, that became kind of the benchmark moment, especially through mcp. And that was a correct tool. Invocation is sort of how that's measured, apparently scoring in theory 81% and that was versus say opus 4.8 at 76% which is a pretty, if you believe in these benchmarks, that's a pretty meaningful jump. But then the distribution piece was kind of interesting. GitHub actually made 2.7 generally available via Copilot. But there was kind of like an interesting piece there where for enterprise teams on Copilot Business and enterprise, this model was off by default and admins had to explicitly enable it. And GitHub's changelog flagged this and sort of said it may be less aligned than other copilot models. Which is interesting because obviously they have their own interests being part of the Microsoft OpenAI ecosystem, but obviously they couldn't miss not providing that to users given that it's on Azure and they still get inference from it. But yeah, very interesting that this sort of slightly odd warning came with it. That wasn't really clear what that was about.
B
Yeah, I believe it's the first Chinese lab openweight model that's been built into copilot. I do get working with a lot of enterprise customers. They do are really careful about which models they use. So I can kind of understand having some level of governance there. But of course there's a certain biased interest from Microsoft point of view of how do you craft the language around this particular model. But there is I think that fear with the enterprise. So I understand some level of controls there, but obviously there's also a certain bias that could be put into it. I think one of the things that's interesting about this is moonshot as well as this goes for all the Chinese labs, they don't have these top tier Nvidia chips available to them like the US labs use because of the various chip embargoes.
A
We don't think they do.
B
But yeah, even if they did have some, probably not at the scale for sure. So there's a certain scarcity of resources that they have to deal with. And I think that in some ways that is forcing them to innovate in a way that maybe the US based companies or the western based companies where they don't have that scarcity of resource aren't necessarily forced to do that. So some ways maybe we're creating the thing that we fear. But good news of that is it's probably forcing the other companies like the western based companies to react to this to lower token prices and maybe also think about how they keep costs down. It's a little bit. If you look at Google's beginnings, Google started as a research project at Stanford and you have students essentially didn't have a lot of resources so they had to be very, very creative about how they scaled Google and then even that carried over to when Google did initially have some funding. But if that had a project had started out of a bigger, more established, well funded company, they probably wouldn't have been ripping apart cheap machines and wiring them using Legos and stuffing them together to compact the size within the data center as much as possible and they would have just bought super beefy servers. And the downside of that would have been a lot of the innovation that we've seen since then in cloud infrastructure came from some of those forcing yourself to deal with these scarcity resources. My hope with all this is we see a similar thing of innovation driven from the scarcity resources that carry over to the frontier labs in other parts of the world as well.
A
Yeah, I mean, it's interesting. Sort of just sidebarring to who's even sort of behind it. The founder Yang Zhilun apparently was known as Yang the Genius. There was a very good profile on him in the Financial Times actually last weekend. He's 34 years old. He was known at university for being just as much into music and having academic sort of soirees, if you like. And actually Kimmy, when he first released it as a chatbot, it didn't do very well. It was sort of outages and this kind of thing. And so it's clearly as you're calling out, it's like managed to catch up despite not having the same access, or we don't think, couldn't possibly have the same kind of access despite what it still may have. Back to sort of then the turning point here, like July 16, which is for history, that's then 34 days after 2.7 had dropped, K3 dropped. And that was a model with 2.8 trillion parameters, which is I believe the largest open weight model released to date. And I think what's interesting here is then on the July 27, the full model weights were actually published. And this sort of then leads into this whole provenance piece, which is when companies are starting to be asked like, well, you know, you're using AI to generate so much information within your own companies or analyze data, you want to actually know how this was determined and so on and so forth. And here we are, we actually have a competitive model with the closed foundational models with the full weights published on Hugging Face, which is massively impressive to me. That's a huge turning point when you've actually got effectively all the weights right there. This isn't just a model that takes up some tiny little a specific task. It's really competing now.
B
I think overall this is whether Musha and Kimmy kind of win out or not. I think overall this is good for consumers of these models because it will force the other companies to essentially react to this. And hopefully. Actually I saw a headline today that OpenAI was reducing some of their token costs by up to 80%. And we saw a similar thing after Deepsea get drove down token costs as well. And then all the stuff we were talking about earlier of token Max and companies kind of getting sensitive to how much they're spending overall. I think it'll be good for consumers of these things. One thing I was thinking about with this story too. We've covered a lot of Meta over the last year and their decisions in terms of hiring and what they're trying to do with their AI lab. But Meta had such a head start in this open weight model space with the Llama models and it's been a long time since I've heard anything about the Llama models.
A
Funny you said that because yeah, there is something. I didn't dig into it, but I just saw the headline which was that Zuckerberg is basically lobbying to ensure that Chinese models don't get banned. And to me that only meant one thing, which is like, well, they have to keep that door open because Llama's not maybe going where he hoped it was.
B
I think his instinct of trying to own the open weight model space was probably right. The execution was bad. I don't know what happened internally there, but they were clearly on a good path, especially with what we're seeing now between the Chinese labs, between what's happening with Fireworks and the other inference providers that have focused on openweight, there's clearly a huge TAM available for open weight to own a big part in the business. And I think realistically, especially in the enterprise, most enterprise businesses will probably have a mixture of both open weight investments that they've done as well as the closed source models. And there'll probably be the strategy that many, many companies take for some period in the future. I don't know how Meta ended up and maybe they'll bounce back, but they were onto something but they haven't been able to execute.
A
I think that's really good analysis. It is. Right strategy, wrong time, too early effectively perhaps, or just too early in the given they were trying to do it from the US being asked to produce too much too soon and that can distort things. So who knows, maybe Zuckerberg or Meta with Llama, they'll kind of adopt more of a Kimi approach to make this succeed. Who knows? But yeah, I mean, just to kind of wrap up on Kimi and sort of this open weight, I would say resurgence, but like maybe surge model provenance is a real thing that's being sort of talked about now when it comes to it's a business risk. Right. Like you're putting money and basing your company on whether it's like for coding or for other tasks. But you are now having to really decide it's a procurement question on what are you buying effectively and what are the risks with that. Could it get too expensive or could you move it onto your own infra if you really needed to or not your own infra, but exactly via fireworks, you could still own the end to end there. That's interesting pricing. We've talked about that. Pricing is just kind of getting out of control and it was really a case of when, not if. Is this going to stop being possible for anthropic and OpenAI to charged this way? Because it's just not sustainable for many businesses. And yeah, then like the speed as well, the speed of iteration has like clearly been a huge piece here where however moonshot are doing it, they're just iterating like crazy. So onto our favorite parts, Hacker News highlights. Do you want to go first, Sean?
B
Sure. So this headline really caught my eye and made me laugh, which was does speaking to agents like cavemen save 65% of tokens we test? So this was by Jetbrains and they benchmarked the caveman skill, which is a Claude code skill that makes it respond in terse caveman speak to save tokens and they claim 65% token savings. So they essentially put that to the test. It was focused on agentic coding tasks and they found actually in reality it was about 8.5% output token savings versus 65%. But a big part of that is they were focused on coding tasks where a lot of it's going to be code. You can't turn that into caveman speak. You got diffs, you got tool calls. So the caveman skill might not be best for that versus you know, some sort of more conversational chat. But I just really love the. It reminded me of like some IG Noble prize, you know, research out there there. It's just kind of a ridiculous topic
A
that's like yeah, super funny. It's interesting because I didn't actually scan ahead. So that is actually a little bit similar to one I picked out which is called the Economic Benefits of Refactoring. So this is on the Fairly popular Martin Fowler.com website posted by Java User on Hacker News. So thanks for that. But this is not written by Martin Fowler himself, but basically this was someone else at thoughtworks and it was can you basically decrease especially the input token cost if you Refactor your code or what are the consequences of that? And the TLDR is yes, if you refactor, then basically you can dramatically reduce your input tokens. And kind of the theory behind this was that the saving is because the agent has to read less code. But the bit that might not be like that sounds obvious, but it's actually not because there is less code to read. It's actually that the overall code in a certain layer they used to help test this, it stayed constant. But the fact that the agent is then able to successfully identify smaller subsets of files to read is the key bit here. So refactoring into smaller chunks, but also making sure that those chunks are very clearly dri, don't repeat yourself, et cetera, et cetera. So that the code that the agent needs to go and grab is smaller. You can't just sort of chunk it up into small chunks and then go, well, it's smaller then doesn't even know. It has to still get all the small chunks because it doesn't know what's most important. But the refactoring is what then turns it into having that context of having what is most important. Go and find that small chunk, input that small chunk. And the output token and output code was virtually the same in this experiment, which is super interesting. There was a small sort of side note on that, which was him saying that Claude, as Claude is actually not good at refactoring. And I can definitely attest to that. But if you can go through the motions of get it to refactor, then this could be quite a huge saving for anyone who's trying to reduce their input token costs.
B
Yeah. So basically, bottom line, good design leads to also optimized token costs.
A
Funny that. Yeah. We just go back to how we used to design software.
B
I mean, it's the same if you think about the human cost. If you have poorly designed software, then there's going to be more human cost each time someone unfamiliar with it needs to ramp up and make some sort of change.
A
Exactly.
B
And then you have another one on CodePen.
A
My second one was. Yeah, just thanks to user RobinReala posting the fact that CodePen 2.0 has come out. This takes me back. That's why I was quite interested in this. I don't know if anyone else out there. This makes it sound terrible. So if you like codepen. Sorry about this, just I haven't been doing a lot of front end work for a long time, but used to love codepen, used to put up all sorts of things on there and a really useful tool as well. Like I did a little coding class, I went back to my old high school a long time ago and did a coding class and codepen was amazing because you could just get people spun up in a browser writing front end code and just see it do its thing straight there. They didn't have to build with like you didn't have to get them spun up with some sort of repo or anything. So that was really helpful. But yeah, I mean Copen must have come out like probably 15 years ago or something like that. So Copenhagen 2.0. I'll just quickly run through a couple of things that they call out that they have files and folders now. So it's interesting, it's almost becoming a bit ide esque but files and folders and then they have sort of they do, you know, build steps anyway. But they've now like added this concept of blocks which looks quite nice where you can sort of see exactly which bits are in this build process or like the linting process. So that's kind of fun. Yeah, real time collaboration as well. So finally multiplayer on codepen. So for anyone that uses codepen a lot, I'm sure that's quite a huge uplift. So. Or if you haven't checked it out. Yeah, still a fun place to go and just experiment with front endy stuff. So yeah, and then just a special one, I guess. Not technically through Hacker News, but thanks to Ilya Rashetnikov for tweeting at me and Sean. We do often cover these Doom. Can you run Doom on something? And thanks to Ilya, he pointed out that there's Another one called DoomQL which is basically using an SQL query is the frame buffer which is like. I think that definitely rivals Dummon typescript types. Yeah, I think that's probably the closest I can think of that it rivals. So thank you Ilya for shouting that one out to us. That was very fun to read through. So yeah, that's it for another SED news. Have we got any looking ahead predictions, Sean, which we usually get wrong?
B
Yeah, I mean I think the safe predictions here would be that we're going to see more headlines on the token maxing issue as companies start to adjust. I think we'll see more also conversations around open weight versus closed model diversification of models. I think those are both going to be big topics conversation through to the end of the year.
A
Yeah, for sure. I will then say, well, because we just talked about iteration and Kimi, let's assume that by this time next month, we're already on Kimmy 3.2. Or if I'll just push the boat out. Give me 3.5. 3.5 by end of August. Let's see if that lands. So thanks everyone for tuning in as always. And we'll be back next month with another SEG News.
B
Thanks, everyone. Cheers.
Date: August 11, 2026
Hosts: A, B, (briefly C)
Theme: Monthly breakdown of top software engineering headlines, a deep dive into runaway AI and open-weight models (the “Kimi moment”), plus favorite discoveries from Hacker News.
This episode provides an engaging overview of the past month’s major software engineering news. The hosts begin with quick headlines—ranging from AI security mishaps to runaway infrastructure costs—before drilling into the explosive rise of Kimi, an open-weight AI model from China. The episode closes with fun takeaways, including token-cost tricks and notable Hacker News finds.
[03:07 – 13:20]
[13:57 – 26:47]
[27:09 – 41:59]
[41:59 – 47:27]
Friendly, conversational, and insightful, the hosts blend technical depth with humor and practical takeaways. The episode features both high-level summaries and deep dives, with frequent asides and playful banter.
This episode is a must-listen (or read!) for practitioners navigating modern AI risks, cost management, and model strategy—with plenty of flavor for builders at the cutting edge.