
Loading summary
Angela
The last layer of abstraction on top of this is probably the coordination layer. So you have knowledge and you have execution, you have coordination. And at the coordination layer, we're beginning to think of these things called, like, strategies, where basically it's almost like a meta harness. The true low level harness is designed for execution, but the next one is about, okay, if tokens aren't really fungible and you need to give them different jobs, like maybe this token is advising versus this token is executing, you want to start composing these, like, these kind of orchestrated strategies that go together and, and they should sit on top of all these things because at the end of the day you still need to execute and the execution still needs to know what to do. So everything in theory should kind of like ladder together. And so I think if you were to look at our roadmap and maybe kind of project forward a little bit, where you kind of expect us to go, we'll move more and more from the knowledge layer to the execution layer, from the execution layer to the kind of coordination layer in terms of the abstractions that you can see us put out.
Interviewer (possibly Lauren or another host)
Caitlin and Angela, thank you so much for joining us today. Lauren and I are thrilled to have you here. You are responsible for building Anthropic's platform, and so you are responsible for building what I think is one of the most important, if not the most important, developer platform in the world. And we are really excited to interview you today to understand more about what's ahead. And so maybe just to get started, can you give us the context of what is Anthropic platform and where do you sit within Anthropic?
Caitlin
Yeah, so platform is both our externally facing APIs, our developer platform that people build on top of when they want to build applications and systems that access Claude's intelligence as well as internally. We run our product infrastructure and basically we're the layer that our apps build on top of internally as well.
Interviewer
Awesome.
Interviewer (possibly Lauren or another host)
What's your North Star as a team?
Angela
That's a great question. We actually, because we have both internal and external, we actually kind of have like two North Stars, which is probably like, you know, like why usually one North Star? But no, we have a planetary system. Yes, exactly. There's separate solar system, so it's fine. But on the internal side, like we really want to provide is like, literally as much leverage as possible for our internal teams to be able to ship like AGI pilled products and we want them to be able to move fast. We will have reliable, like, great platform to be able to build on top of. But I think that key bit about speed is like really intentional for us and we really, really care about that internally, externally we actually have a lot more like complicated set of things. But one of the true norths that we have there is to be able to basically give any builder the tools to be able to work with Claude, to build whatever they want to build. And so it's a bit of a broad statement, but as a result that boils itself down into being wherever that business is, like, we really care about bringing our platform really, really close to that business. This is why we spent a lot of time with the hyperscalers, integrating really closely directly with them, like aws, Google, so on and so forth. And it is a lot of like primitives that we end up creating. We want people to be able to express what they think their product should be. We want them to be able to almost do like custom software in their own way. You know, like in this new world with AI, what used to be probably economically impossible was that last mile of custom software. Now in theory should be like very, very achievable. And we want to give them all the tools and all the capabilities to go and do that. And so sometimes that comes that form of primitives and APIs and higher order abstractions, and sometimes that comes in the form of just like standards. So for example, like Skills and mcp, those are things just like Claude needs them to be useful and we can just give them out to the rest of the ecosystem, work with everyone to help you create those things and get the best out of Claude. So I would say externally we really are oriented around just helping you just be able to build. But internally that orientation, while still existing, is probably more specified towards speed and being able to move really quickly.
Interviewer
How do you decide what goes into the platform, what gets externalized and what doesn't, to decide what products should be available?
Angela
Yeah, I mean, we generally try to have a philosophy that we try to be consistent across the board. It's actually one of the reasons why we do internal and external. There's plenty of other platform businesses and constructs where you actually bifurcate these two things for us. We kind of try to intentionally keep it equal and then as a result, we try to hold this philosophy as much as we can around for any builder, internal or external. Even though our internal builders might have some slightly different requirements in the same way any user would have slightly different requirements, we want to have the same primitives that are available to everyone. Maybe the overarching thesis for that is that we've just seen the capabilities of these models just grow and such as exponential and it's really hard to figure out a long lasting form factor. I think two years ago we were all like, everything's chat. And now everyone's like forget chat. And he's just like agents and there's going to be another form factor, another form factor. And we kind of imagine that constantly evolving. And so the best way for us to kind of enable that for everyone and also ourselves is to actually build a really robust platform that gives people those kinds of tools to figure out what those form factors are. And I don't think we by any means feel like we're the only ones capable of figuring out that form factor. Like not at all. In fact, the more democratization we can do on that and help people and allow people to experiment, I think the more those form factors will actually kind of naturally come out of the market.
Caitlin
Yeah. And I think within our team we've had moments where we're experimenting even with just like a packaging up of our primitives in a different sort of higher order way. And we've thought about, okay, cool, we've solved this exact type of problem with this product that we've built into the world and so we can go and dog food it for ourselves. But we never want to fall into this trap of like we, we're over indexed on the problem as it needs to be solved for an internal user like, because exactly what Angela said, internal users have very specific requirements, external users have very specific requirements. And so if you over index on one or the other, you fall into a trap. So a lot of the time what we'll do is dog food something internally at the same time that we open up early access of some sort with external customers so that we can kind of get a range of feedback and bring those things back into the platform.
Interviewer (possibly Lauren or another host)
I'd love to talk about the higher levels of abstraction that you discussed. So I guess at the base level this is just raw access to Claude Opus or whatever tokens. How do you think about the layer cake of abstractions above that?
Caitlin
Yeah, if you look back. So when I joined Anthropic around a year ago, the platform was basically just the Messages API. It was a Messages API. We had come out with standards like MCP. We obviously have developer tooling around our SDKs and our docs and our console and things like this, but for the most part it was a stateless API. And what's interesting to Angela's point on form factors evolving over time is we found a lot of our customers solving the same problems over and over again that we also were solving over and over again around as the models got better at running for longer and working with more context at a given time. You want to build agents that can succeed in a kind of long running context and even a remote context that doesn't necessarily have a human in the loop. And so we found that we could piece together our primitives and stand up all the same infrastructure that we're finding ourselves standing up internally to power our own products and arrive at some higher order abstractions that let you do more agentic work out of the box. And the problems that we're solving for you are infrastructure being kind of a hard thing to deal with. Like how do you figure out spawning sandboxes that are going to have the right governance and security and like spin them up and spin them down when you need to, or the storage around transcript sessions so that you can resume a session if you stop it and pick it back up later. So that infrastructure is a big thing that we wanted to be able to provide more of out of the box and we do more of that today. And then the second thing just being harnesses and harness engineering. There's a lot of thought and energy going into how do I do my prompt caching and how do I manage my context window as well as how do I actually just get more intelligence out of the model and how do I manage my costs and things like that. So we've kind of packaged up our primitives a bit more in tune with the problems that we found ourselves solving to provide more of these things out of the box for people so that they can, if they're building systems for themselves internally, if they're building products, they can just be more focused on the problems that they want to be solving. And if they want to offload some aspects of those problems to us, they can. And that's kind of the ethos.
Interviewer (possibly Lauren or another host)
And are your customers generally choosing to opt from the grub bag of stuff that you offer or are they like how often are they opting into the. Just the managed agents offering, I guess just take care of it all.
Angela
For me, it varies by the user group. So for I would say really AI native startups, the ones who are tinkering and experimenting at a really low layer, they're just going to go for the primitives and then for everyone else, these are kind of classic more enterprises or areas where it's like the purpose of the startup or the philosophy behind the startup isn't necessarily to optimize on some kind of hill climbing piece. It's more like stringing together a bunch of workflows and providing unique user value to that user. For those people, you know, it's just kind of not their core competency. It's not where they want to focus their time and resources and they reach much more for these kind of like higher order, like package offerings.
Interviewer
What are some examples of the primitives you've released at different layers in the last few months? We've seen a few of them. Would love to hear.
Angela
Yeah, I think maybe one framing I would give for some of the constructs that Caitlin was talking about is like, and this is a bit of an oversimplification, but effectively there's approximately like three layers of this cake. At the very bottom is just kind of like, like knowledge. And so at this layer, like in many ways it's knowledge about the model, it's knowledge about the things that the model needs. And it's just like the ability to know how to actually do something with Claude is maybe the way I'd phrase that. And so there. The primitives that we have spent more and more time on have been actually things of the past because like, we still evolve them, but they tend to be a little bit more baked. Like for example, there's very specific shapes and parameters we put on the Messages API. And it's more like trying to expressly like showcase Claude's like, design. Like Claude, the model's actual design, the way it thinks, the way it respects certain parameters, the way it kind of like will do tool calls, like all of those different pieces. And then we started standardizing like tools. And then we started standardizing bits and pieces of like context that you could put in at different moments in time, which is concretely like skills and like memory. And so those are like the kind of like knowledge layer type of abstractions that we've put out over the past, I guess like year plus plus a bit. The next layer of abstraction that we've actually started to spend more and more of our time on is once you know stuff you then need to execute. And so at the execution layer, that level of abstraction is the part that Katelyn was talking about around, like we're doing these higher order pieces, but what are we putting higher order there? It really is because you're now getting Claude to execute work. It's not just to know something. Right? I can give it a question and give me an answer. You can put string a lot of that stuff together. But now if you need to execute, do work, give me the output, edit files and a bunch of different systems that becomes a lot more complicated and requires infrastructure to handle. And so that layer is basically, I would say a low level harness plus managed infrastructure as like the set of abstractions Today we just like our high level product for that is called Quad Manage Agents. And so that's like a piece, but we started to wrap more and more pieces in that. I think there's going to be a layer like on top of that. We have like some inklings of it if we started to build towards. But the last layer of abstraction on top of this is probably the coordination layer. So you have knowledge and you have execution, they have coordination. And at the coordination layer, we've started to expose some of these in ways that aren't very obvious, but we're beginning to think of these things called strategies where basically it's almost like a meta harness. The harness, the true low level harness is designed for execution, but the next one is about, okay, if tokens aren't really fungible and you need to give them different jobs, like maybe this token is advising versus this token is executing, this token is dreaming versus this token's executing, so on and so forth. You want to start composing these like these kind of orchestrated strategies that go together and they should sit on top of all these things because at the end of the day you still need to execute and the execution still needs to know what to do. So everything in theory should kind of like ladder together. And so I think if you were to look at our roadmap and maybe kind of project forward a little bit, where you kind of expect us to go, we'll move more and more from the knowledge layer to the execution layer and from the execution layer to the kind of coordination layer. In terms of the abstractions that you can see us put out, that's a really cool way.
Interviewer
How do you think this all comes together into a broader ecosystem beyond just the things that you guys are building? How do you help support people building products on top of it? And how do you help them get the most out of all these pieces?
Angela
Yeah, I think this is super top of mind for us. We really want to find a way to support as many people in doing this as we can. I think we're still learning. A lot of the industry has evolved. We've seen a lot of different pieces get spun up and spun down and. And I think the operative part for Caitlin and I has been in the category of like making sure, at least at the Base layer that we provide as many primitives across the board as possible. So you know, this kind of like, yeah, like knowledge execution coordination layer, we want to give all of that out to everyone so that people can start to compose and create on top of that. And that's just from a, I think pure builder kind of point of view. Then there's a point of view around like, how do you kind of like plug in with us, right? Like we're also building first party products of our own. We've also created some ways to embed natively with us, like for example connectors which are built on top of the MCP spec. And we try to be more open about those types of things. And we're starting to figure out like what are the right bits and pieces. But what we're really trying to do is get to a place where, you know, a company is able to get created and built on. They can build whatever products that they want. They can build agents if they need to. And then those agents and those products could be things that could point plug into other agents. Some of those agents could be cloud agents, some of those agents could be other people's agents. But we want to be able to enable that kind of transactability across the board. And then I think in order for all of that to kind of ultimately be true, there is a bit around standard setting and I think there's the traditional standard setting which is around how do systems interoperate. And that's things that you've kind of seen us do with like skills and mcp, but they're at again like the builder layer, I think at a higher order layer. There's also a bit around interoperability and standard setting, around how do we all kind of treat safety together? And we've talked to a lot of these companies and this is less from philosophies aside, just more like no one really wants to have technology that's for example doing negative things on their service. So Cyber, I think is a great example of this. You want to protect your own systems from negative actors or bad actors. And so these kinds of standard settings of how can we find ways to partner with more and more people to be like, yeah, we all kind of want to make sure our critical infrastructure is good. We all want to prevent like fraud or any of those things from happening. And how can we work better with each of these members? I think on the last layer we're still kind of like we're still evolving and I think we're still very much like trying to find ways that we can be better and work with the rest of the industry to bring people along and work with them. But those are kind of like, you know, the higher order primitives or pieces that we wish to kind of like be in place so they can work with folks to ultimately solve this. I think if I were to take a step back at the end of the day on all of these things, this technology is so transformative. And if it's a little bit like electricity in the sense before electricity there was just like you had to have a candle and it was like you can only do so many things but with electricity. The reason why it's such a transformative technology for all of us and so greatly of a utility is because you can actually wire it into everything. Everyone is able to actually access it. We also have standards and ways to plug in and do all the pieces that we need. And that's not something that anybody can do by themselves. They always have to work with the ecosystem and work with partners to figure out a path forward.
Interviewer (possibly Lauren or another host)
How do you think about the philosophy of building an open ecosystem versus a walled garden? And how do you think about what products are really important for you to own first party versus where you're perfectly happy to plug into other components of the ecosystem?
Caitlin
Yeah, so maybe in using Angela's kind of layered cake that we talked about a little bit earlier, you'll see that on some pieces of this, like execution, for example, what we've done within something like cloud managed agents. And I think over time you'll see us try to make this a little bit more modular. We actually aren't precious about. You should run these things on our infrastructure. Like it should be sandboxes that we control or it should be a storage layer that we control. Well, we actually like, for example, we launched self hosted sandboxes and we partnered with Modal and Vercel and Cloudflare and a bunch of other folks even like Amazon's new micro VMs to have a first class offering where you can go plug any of those things in. We launch MCP tunnels so that you can call out to your MCP servers that are behind your firewall. Right. And be able to punch through there. And so for some of these things we, you know, the weather, whether it runs on our infrastructure versus somebody else's infrastructure is actually not important to us because the thing that's important to us is more that the architecture of how you put together these agents in a way that will be powerful, in a way that will be reliable and Scalable. We have strong opinions on that. And you can kind of just conform to the interfaces that we put out there and plug those things in. And we think that that generally is a thing that works really well.
Angela
Yeah, I think on the kind of, like, verticals where we might build products, you know, I think we kind of have like, two frames here. The first one is we are always trying to figure out a form factor, like an evolving form factor. We, by the way, don't think form factors are static. It's like a dynamic thing. So what might be awesome for one year's worth of AI development will probably not be awesome for the next year's worth. And we just kind of try to have that mentality. We tell the team just overall around anthropic, everyone's always trying to be like, is this AGI pilled enough? And then we also have this mentality of, like, you know, we built something, it works, it was cool for a year, and maybe it's not the right next thing and so throw it away, try again. And we tell, like, platform users the same thing. I just think that's probably just, like, you know, attached to the technology. But so, yeah, one. One principle is, like, trying to always, constantly find this new form factor. So sometimes we'll, like, launch products in certain areas to try to showcase a new type of form factor. It's not necessarily because we think it's like, the biggest TAM or the most important thing to go after, but sometimes, like, okay, this is, like, always been a really difficult thing and people have always communicated this way or tried some things this way. And can we show that maybe there's a slightly different way? And because the model capabilities are so advanced now, can we try to express it a bit differently?
Interviewer
What's an example of that?
Angela
Yeah, you know, like, cloud design is a little bit of that way. I think depending on how you squint, you might see it as, like, a way that we kind of are going into design as one of the verticals. But more often than not, it's like, if you take a look at what we're trying to do with that product, there's a couple of, like, decisions that were made in there. The first one is that, like, you can actually try to offload more and more and more to Claude. And so it tries to be kind of opinionated on, like, you know, just. Just, like, talk to it and, like, let it really try to figure out. And yes, you can still edit it and do these kinds of things, but kind of, like, discourage a little of that and more just like let just talk to Claude to go figure it out. The second thing was it was really trying to express that actually like code is a, is a way to solve for things that you wouldn't normally think would be the way. So a lot of people who have built kind of generative, you know, like slide decks or designs or whatever will pick the way of like they have like some kind of design system you integrate against the design system. It's almost the traditional like classic WYSIWYG style of designing something. And with like Claude design it was like, okay, can we try to just like use code purely have Claude generate that code and would it like do a good job? And we found through some experiments early on, it's like actually it looks like it can kind of do that and how can we kind of showcase that to the world? So that's like an example. We have a lot of other internal projects and this kind of falls in the category of like expressing form factor. We'll all try it out internally, it'll be super cool for like two weeks and then we move on to the next thing. We never even ship the thing frankly. But yeah, we actually do a lot of product experimentation in that area. And that's like our labs team. And then there's like the second category which is that we actually do look at tam like we're a business. We do look at tam. We do look at areas that we think, you know, there'd be reasonable agentic like operations that would happen in those areas. We do tend to have an orientation towards things that are more token heavy and by token heavy or token hungry maybe is the way I would say that is like what we mean is like, you know, you for spending, once you spend a like, call it like one turn, you look at the end of that turn and you say like, am I done or am I actually so glad that I did that thing? I want to do more of that thing. We like industries where it's like the answer to that question, you say, I want to do more of that thing. So coding is obviously the one that we all know. And the great thing about coding is that what it's actually doing is that once you've finished a turn, you look at that and you're like, that was incredible. I'm unlocked, I'm going to do more, I'm going to build more, I can do more. And there's other services where it's like actually when you finish that turn, you completed the job and you just move on. You Know what I mean? And so we tend to go into the ones that are a bit more like there's this kind of iterative flow, you're going to build more, generate more together. And then the last angle that we kind of take a look at is just sort of like there's going to be certain business functions that we're like, they are the buyer that we like to go to. We want to help them optimize their workflows, help them create better products there. And I think we've been pretty transparent with some of the verticalization. Like we've done like finance, we've done like legal and we've tried to kind of like narrow on into specific areas where we feel like by having the right context and the right tools and putting it together in a good form factor is probably useful for us to be able to do.
Caitlin
And in each of those areas we do, we're trying to do a bit of like showing the art of the possible across all the different ways that you would accomplish those outcomes. And so for, you know, like finance, for example, is a good one. You know, we, you could be a company that solves problems in finance and you could build directly on the messages API and you can just get some tokens and you can build everything else on top. Or you could be someone who builds on cloud managed agents. You can get a lot more out of the box. Or you could say, I'm going to build a plug in that or like a connector, right? That's going to sit within one of our products and within those form factors. When we did recently, we launched like cloud for financial services is like, okay, cool. We've got packages of skills and things like this you could choose to use within our product, within other people's products. We even launched like cookbooks on. Here's how you would use cloud managed agents to go and do these things. And so I think for us it's all kind of an experimentation around like, you know, we provide people all these different pieces and see kind of where they run with it. And then sometimes we put together products that are just packaging of all of these things. Like Claude Tag I think is a really good example. Like we had been seeing people in the industry go and say like Shopify did this with River Square Block, recently did this with Builder Bot. There's like a few of these examples where people said, I'm going to, I'm going to build like an agentic platform internal to my company and I'm going to try to give it all the right Context and I'm going to make it accessible from Slack or from various other, you know, platforms that you'd want it to be accessible at. And I think Claude Tag was very much a packaging of all those same things that anybody could choose to build something similar. But this is how we're kind of like, well, this is how we're doing it internally. And if you would like to just kind of plug in and go, here's what that looks like.
Interviewer (possibly Lauren or another host)
What do you think people misunderstood about Cloud Tag because there's all this like ruckus about oh my gosh, it's just a Slack bot, like tell us what the magic of Tag is.
Angela
No, I think it's a great question and I do think it actually showcases a little bit of where maybe the future could be going. Yeah, I think like the, I think if you look at products in the past, people are like, oh, you really attached to like the form or the UI almost right. Like it looks like this. So it's like super cool. And I think when you look at like tag it like, yeah, like the way you interact with it is that you like literally tag it in Slack. And so yeah, that is like the interface. But that's not really the important part. The important part is all the kind of like context engineering and like architecture that we put underneath the hood so that tag just works. It really should just like just feel like a coworker, like a co. You know, if you go to a company and you onboard, the coworker comes into your channel and then you can chat with it. It's proactive, it figured out like what's like useful and it just gets stuff like done for you. And so if you think about, you know, especially like non technical audiences, this is like, it's a huge unlock. You just, you literally create a channel and then you Claude, or sometimes you don't even Claude and you're like, hey, I want to be able to do this and do that and I can't figure out this and how do I actually like submit an expense report again? And traditionally you think about how to solve that workflow, you are going all over the place and you're talking to your manager and you're talking to your spin up buddy and it's really, really complicated. And today now you just like go talk to CloudTag and we do a lot of the hard work on doing the context engineering, the proactivity, a lot of the harness pieces. I think Andrej Kaparthy said it really well. He's like, it's like an org level harness. There's a lot of complexity baked into that. Like Caitlin mentioned, you can use our APIs to go and construct that. You can do a lot of the experimentation yourself obviously. But this is an opinionated take from Anthropic on how you can have this really awesome, always on kind of agent for your entire company. And the bit that's futuristic I guess is a lot of that complexity is actually like, it's like an iceberg. It's like all the stuff underneath it that's actually becoming the harder and harder and useful part that we're trying to push through. And I think we'll see more and more like that kind of like tidbit that's like outside in the water. It's just like the interface can actually constantly swap. Like today, right. Like Slack is a place where a lot of people collaborate, a lot of business collaborate, but also a lot of people collaborate in teams. And some people collaborate by a WhatsApp group or they text each other or they may. Some people still email each other. And like those could be the form factors that actually completely. You can imagine agents just going there and being. And they're almost taking up the same form factors as humans have taken up. It was almost like a very, almost like boring take. But it's actually like I feel like the most like forward one because you want the agent and you want AI to basically be like another person and it's helping you, but it's like, you know, very intelligent, can figure out all the context and you can always have it to be a really helpful assistant.
Interviewer (possibly Lauren or another host)
Totally. You talked about context and then harnesses quite a bit. And so your team has such an opinionated point of view on what it takes to build an exceptional agent. I imagine a lot of that comes down to the context engineering and the harnesses.
Angela
Totally.
Interviewer (possibly Lauren or another host)
Maybe what best practices or advice would you share with people about what you need to get right on the harness and what you need to get right on the context?
Caitlin
Yeah, I think so. It's interesting because we've kind of talked about, you know, we launch cloud managed agents as this like very generic but high performing harness because we've done all the nitty gritty work that's actually like really boring and not super interesting around. How do you deal with prom caching? How do you deal with context management? You like clear old stuff out of the window. Sometimes you like call tools programmatically so you don't pull everything into the context window and you can keep it clean. There's a lot of Those sort of details on the lower level harness layer, and I think honestly, like, best practices are just stuff like prom caching. Do it. You're going to save a lot of money and token costs. Obviously, try to keep your context window clear and then putting those things together in a harness that will be performant is sometimes specific to the task that you're trying to accomplish. Right. And then, of course, evals. I'm surprised we got this far into this thing before one of us said the word evals. We're like, you need evals to make sure that what you're trying to accomplish is performance. But I think where we're starting to go, and Angela mentioned this a little bit earlier, is more of a concept of strategies or meta harnesses, because I do think that, yes, you can again make this lower level harness. It's going to be performant. And maybe that's interesting for you to do yourself or maybe not and you offload it to us. But this concept that you can take any given token and spend that token on just executing, or you could take that same token and choose to actually reflect on your past agentic sessions and write learnings to memory so that the next agent does a good job. Or you could take that token and advise with a bigger model so that a smaller model can execute and do a better job. Or you can say, execute, execute. And then like a grader comes in and is like, did you do a good job? No, you didn't try again. Right. And so I think the, like, interesting innovation is going to come more at that higher level on like the meta level.
Interviewer
Right.
Caitlin
And I think optimizing within those strategies is something that our team is really excited about and we're starting to do a lot of work there. And I think a lot of other people are starting to feel really excited about this concept of strategies and like the jobs you give to tokens because again, like, yes, there's best practices on stuff like your prom caching and exactly how you clear stuff out of your context window and how you write your evals and like a lot of things like this. But I don't know that there's necessarily so much juice to squeeze in a lot of cases out of that layer as compared to a layer higher than that.
Angela
Yeah. And one of the reasons for that, I think, is it has to do with the generations of the models. If you look like two years ago, a lot of the harness was like a scaffold to kind of like tell the model to go from point A to point B and you had to like, you really had to like build in a lot. You practically build one wall here and one wall here. So like the thing would go in a straight line. And now the models are actually very, very steerable. And so a lot of that steering you could just put in the prompt like, go, do go from point A to point B and the model like will go from point A to point B. So a lot of if you have harnesses that are like designed to kind of do that kind of like steering, you can delete that part. Like that part. We actually frequently encourage you could delete part of those harnesses. I think various people have said things along those lines and that's, I think what people oftentimes mean when they're like, either the model will kind of consume some of the scaffolding. And like, in that sense, like for sure, if your scaffolding is telling it to go in direction that it can just intelligently figure out, like, that I think will increasingly continue to be so. But as a result of this, what the harness needs to start doing is more allow it to run longer. And so that's where that execution bit tends to be. I think it sounds like a maybe somewhat silly point, but I do think it results in a lot of differences because you can go in the direction that you tell it to go. You obviously don't want it to stop at B. You're going to be like, okay, now go from B to C and then go to F and then go to Z and then come back to me on A, something funky like that. In order to be able to do a lot of those things, the kinds of harnesses that you do are less the steering harness and it's more like these kind of strategy harnesses that Caitlyn's mentioning, which allows you to operate at a slightly higher level of thinking, which matches, I think, a lot of the intelligence gains that we're starting to see with the model.
Interviewer (possibly Lauren or another host)
Do you think task specific harnesses make sense or vertical specific or task specific harnesses?
Angela
I think people have different opinions on this. Our opinion is yes, I don't think there's a general harness. I think there are some capabilities that are obviously very general and they tend to be very useful. Like coding is a capability that is very useful because you use it across so many things. And software has just eaten so much of what is capable. So our ability to write software is therefore useful. I think when you think about very, very specific types of domains, they were going to require a couple of pieces of the harness to Be sort of customized. One of that, I do think, is how you choose to kind of handle errors between when you do something and you hand something off to the model. So in domains where you require an extreme level of verification, that logic of how you handle that, again, I think it sounds small, but I totally understand why some people feel like they really want to own the harness, because tweaking that last bit will give you a ton of juice. And especially domains like legal and finance, where there's a lot of consequences to not getting it perfectly correct, is really going to matter. And that's going to be the difference between your product and someone else's product being the thing that the user ultimately uses. And then there are other domains for which I would say it's not going to matter as much because you're able to compress it into a general model capability. So the tweaks that, I guess where we feel like the domain specificity is really going to matter is the specific verification logic between the model and your execution. And then I think it's going to be about some of these kind of higher order strategies on how well you're able to actually allocate your token budget. I think the context bit is actually a little overdone. Yes, you're going to throw in context, but any harness can actually handle a lot of context. And so that's just more like you have the data, and if you have the data, then obviously you're uniquely qualified to do something useful.
Caitlin
Yeah, and I think when people say harnesses, they often mean a lot of different things. And I think this is why, in part, there's so many different opinions on this. You can think of a harness as literally just a loop. Um, that's like, okay, cool, like user model, user model tool, you know, like that sort of thing. Then you could think of the harness as also all of the tools that are packaged up with the harness. Right. And. And there's just like a lot of different definitions of these things. And I think the stuff that can be pretty generic and like less interesting to own and deal with, kind of saying earlier is like getting your prompt caching rights. Right. Like maybe that is not the world's most interesting thing. Choosing to clear out old tool calls from the context window and things like that. Right. Are like maybe a little bit less interesting. And you like go a layer higher into some of the stuff Angela's talking about. And then you get into like, okay, yeah, these are things that I might want to own and control. And so it's interesting with Cloud managed agents, like the thing that we built today, we call it higher order, but it's not really like that high order in the sense that you can choose to define all of the tools that you want to bring in as custom tools with the harness, right? And like, we give you a lot of knobs to control. You can define skills, you can do your system prompts, you can do a whole bunch of different things, MCP servers and things like this. And I think where, you know, we want to get to you is a point where you can literally just tell an agent, here's the outcome I want and here's the budget that I want to spend, like, ready, set, go. And you may be like, don't think about any of those things underneath. And so I think there's just a few different layers of this, right? That for certain things, like you might want to sit at a different layer of what you actually go and control. And you can probably get better outcomes within some of those layers by doing a little bit more optimization work.
Interviewer
Very cool. One of the things I'm curious about and one that I love about infrastructure and platform teams, is that you get to see what the most advanced users in the world are using and learn from them. I'm curious, what are some things that you're seeing and learning from the people building on your platform?
Angela
There's some people that have been doing some really funky ways of handling context. We ourselves explore this a lot. That's actually one of the reasons why Tag is such a great product, is there's a lot of really awesome context kind of engineering that's happening. We've seen some teams be really clever about how they do that, and they are able to kind of think through, like, okay, if I have all these contexts in a bunch of different places, how can I proactively go reach out to them? How can I try to generate enough permissions across each of them and then feed that all into an agent? And it's interesting that I guess this is kind of the level of innovation that we're actually very excited by. It doesn't express itself as like a completely different product form factor, but what it actually does express itself as is like maximally useful to users. And we've been seeing this more and more with actually internal use cases instead of external ones. So companies who are becoming more AI native, basically, they're the ones we're seeing increasingly more and more innovation out of. And so we've had customers try to do this for. They've built their own custom sdlc Kind of set up in very, very innovative ways. We've had ones who do that for like their entire back office. And just like the kind of nuances of how they like stream in context I think has been like actually really interesting in terms of like how they've been putting together the pieces. So that's been like one category that's been like really, really like fascinating. Another category that's been like really interesting has actually been with companies that are dealing with like really old school software. And so there's a lot of like healthcare companies that we kind of engage with. And you know, like, they're like the systems I'm working with, they don't even have APIs. Like that's, that's a dream. And so how can they use computer use and things like this to be able to start to kind of automate and create more connectivity with our systems. And that area of innovation, I think has been really exciting. It's been really interesting to see people try all sorts of crazy stuff from like taking a laptop and trying to like run a bunch of things on it, to auto generate a bunch of things that then their agents can go and use. And this has actually been probably like an area of, I think, a lot of innovation coming from a lot of our customers that we want to find ways to support better and see. Like, okay, maybe there are like, how can we make this easier for you? How can we help you with some standardization? How can we get it so that, you know, like you can just have a spec and then Claude can then respect it. And so it's much easier for you to organically connect a lot of these things. But yeah, maybe the general theme I would just give you is like, interestingly, a lot of the innovation that's most exciting out there right now has been this kind of like context and connectivity layer, which has been really fascinating.
Caitlin
Yeah, like a good one in that we were working with a customer who they built some agents on cloud managed agents. They also have some agents that they built on other models and other platforms. And they've kind of optimized each of these agents to be good at the things that they want. They want these agents to all be able to work well together. And they kind of were like, wow, Galaxy brain, like, what if I expose an MCP server on top of this agent so that it can then go and like have this other agent call a tool on that agent, right? And have these things just be more modular and be able to work together. And we were like, yeah, totally. And we sat down with them and worked through it and it worked perfectly and it was pretty cool. And so we're seeing a lot of again that connectivity layer that I think is one of the cooler areas where people are innovating. But outside of that, one thing that has been cool is just seeing the shift in I guess like industry trends of where we're seeing a lot of our usage come from. Like we talked a lot about coding, like coding as a category, like of course, absolutely exploding. There's so much going on there and we're starting to see some of these emerging trends. Like more recently we're starting to see manufacturing really pick up as just a category where people are building with AI and like one of our PMs, like getting on a flight to Detroit to go like figure out what these customers like, what they need and what's going on. And so I think we're going to start to see a lot more just kind of like outside of the box of what people think about today, sort of use cases, which we're really excited about.
Interviewer (possibly Lauren or another host)
It seems like there's now there was a. We went through a token maxing moment of history and now there's like the token rationalization moment of history. What are your thoughts on that and what should companies be doing and then how does the platform team think about enabling that?
Angela
Yeah, I mean it makes sense. It makes sense from the high you start to rationalize. I really like that framing. And I think there's a couple things that are top of mind for us on this front. I think again, it makes sense. And as these models get more and more capable, you're going to hit levels of intelligence max maxing that are there, that then you want to do the next kind of dimension. And the next dimension after intelligence will either be cost or it will be speed. And you just kind of go through that across all possible task complexities in the distribution. And as we kind of see that happen, something that's really top of mind for us that we kind of try to spend some time with users on is what you don't want to do is stop AI usage. That's kind of the wrong move. And we do actually see some of our customers do that. Oftentimes the way that AI spend has erupted inside their company has been through some kind of like shadow it, you know, like their employees just like want to use it. They find a way, they end up procuring it themselves. And before you know it, like half your org has like found some way to install cloud code. And in that world it is kind of hard to manage because these things are again, like, they're very token hungry ultimately. And so what we try to kind of encourage our customers is like, you don't want to, like, stop the innovation. Like, if you are getting returns on top of this, you are shipping faster than ever before. You can, like run more operationally, like efficient. Then those are gains. And so the area that we actually try to encourage people is like, if there is a way for you to kind of construct again, like a strategy that allows you to design an architecture that says, like, given a task, assesses level of complexity. I mean, I'm effectively describing a router, but, like, there are ways to do this that are like, I think a bit better now. And so, like this task comes in, has a certain level of complexity. For that level of complexity, like, you can define some rules, but for the most part, right, if it's like a hard task, you should probably route that to like a big super smart model. And if it's not a hard task, you can wrap that to like, cheaper models. Designing that I think has a little bit of like, there's a lot of technical complexity in that, but it's like very, very doable. And we actually, like, encourage people to try those kinds of things. I think ultimately, I think within the quad space, it will like, make sense. It's actually one of the strategies we imagine, like, designing because the way that we are kind of thinking a lot of these things is like, it almost feels like every month there is a new era of something. And if we just take a step back, we're like, okay, and this seems to be really fast. And so what are the different ways that are recomposable so we can redesign very quickly for any new. Whatever the cool thing is that month kind of like bit. And so this is like in that category of things where we feel like we can actually just recompose a lot of our primitives and then design it. I think the bit that we do feel really strongly about on the model routing front is like, we are designing our platform for Claude and we want to make sure that Claude is great at solving all these things. So we'll restrict to that space rather than. I don't think we're that interested in saying, okay, and then you should route to a different model or whatever.
Interviewer (possibly Lauren or another host)
Makes sense.
Caitlin
Yeah. And, well, some of that too is just like, I think we have a strong belief that harnesses and just like the agentic layer should be tuned to the model family that you use it with. And so I think there was a period where people were kind of like, yeah, cool, I can build a harness and build an agent and then just like, plug in a different model underneath. And they were excited about routers from that perspective. And I think we started to see, like, Vercel just did this with Harness Agent, for example. Some of these players in the space come up a layer of abstraction and say, actually plug in the whole harness and the whole agent that's tied to a model family, which makes a lot of sense. And so what we could provide is a little bit better, smarter. Like, how do you mix and match the right models within the model family underneath that thing, if that makes sense. But, yeah, on the general question of token maxing costs and these sorts of things, I think we're just kind of going through what feels like a normal, natural cycle for companies and figuring out how to make the best use of this technology and run their businesses really well and really effectively. And it's interesting, like, before working at Anthropic, I was at Stripe, and we were kind of in the very reasonable era of, like, we paid a lot of attention to our AWS bill. And so if someone were to have built some background job and they didn't quite configure it correctly, and this thing's, like, burning through, like, CPU or whatever it is, right? Like, at any given moment and causing a big increase in spend, that's not actually worth it, right? Like, we have put in place the guardrails to find that and then go ask that engineer very nicely to please turn off their background job. That's not, like, within the bounds of what they should be spending for the thing they're trying to accomplish. I think those are the things with AI that people are going to start to go and figure out. And I think to Angela's point, the thing that gets dangerous is when you're kind of just like, here's a cap, and you're stuck within your cap, like, ready, set, go. But I do think that encouraging innovation, encouraging people to create really excellent outcomes with this stuff, and then coming in from the side and looking and saying, like, okay, well, there are a few different ways that we probably could have accomplished that outcome, right? And one is like, you take Opus and you run it all night and you do something crazy. And another is maybe to get a little bit smarter with the strategies that you put together in order to create that same outcome within a lower cost. And I think that's the, like, next layer of thinking that everyone's going to
Interviewer
start to do Very cool. Is there anything that you guys are excited about building over the next few months that you can share a hint at what might come next?
Angela
Yeah, I mean, I know we said this word like 20 million times, I apologize. But like, we really are trying to build ways for you to compose strategies. And so that is an area that, that we're like trying to move into that kind of like coordination layer of the abstraction. And we want to start at this front because the types of problems that we see people building, they are at a layer where it's like, in order to get the most return on this, you have to be a little clever about what is the nature of the problem that you're solving. So to give you something like concrete when you try to solve for, let's say you want to build an agent that's trying to do bug hunting and you could just send one off to go and do that, and it's going to give you a certain level of return of possibility. And then people kind of get stuck at that and they're like, okay, my next options are I can like make a bigger. I can just like swap the model for a different, probably bigger model, or I could like let it run like longer. And that's pretty much like the only two, like levers that you have to like try to make this like bug hunting agent from. A lot of experimentation when we do these kinds of things, there's like actually the thing like those two, those two things are still true, but you actually have like a third lever and tends to actually do a lot more than you think it does, which is that actually if you were to best of end the thing, it would give you a lot more returns. But just to be just saying those words are fine. There's plenty of papers and people have published it. To actually build that thing and put it into production so you can actually test it on users and see the results for yourself. That's really, really freaking hard. And you end up building all these custom harnesses, so on and so forth, all that stuff. But we're seeing this is where the alpha is and it's hard. And so in the same very simple philosophy that we talked about at the beginning, like, if it's like gives you the return that you want and it's hard, we're going to go try to just make it easy for you so then you can use it to then run the experiments you actually need to run.
Interviewer
It reminds me of when people are talking about agent swarms a year ago. It's some version of that.
Caitlin
Yeah. In a whole year. Yeah, I know.
Interviewer
We're finally there.
Angela
Yes. Yeah. No, I think that that's like, that's a type of strategy. Exactly in the same way that you have like, you know, one big one that separates a bunch is another type of strategy. And I think people have thought about this maybe in the way of like human organization. I guess it could be similar, but if you take it to kind of its end state, it's actually more just like the token has a job. And I think it's this job piece that we're really indexed on and we see a lot of returns too. And that's the thing that we want to spend time with users and the rest of the ecosystem on. How can we just make that easier for folks to then experiment? We can give you five jobs off the top of our head and we're probably like, that's what we have internally. And if we give this out to the rest of the ecosystem, there's probably going to be like 100,000, 200,000, who knows what other combinations that people could put together.
Caitlin
Yeah, we want to be able to keep doing this hill climbing on like, how do you get the most value, the most intelligence per dollar and just put that power in people's hands. But around the edges of that, we have these Personas that have kind of just like things they have to work through in order to be able to like really deploy AI either within their companies or within their products. And that's like the sort of enterprise ready security and compliance controls and things like this. But really even just like making the platform more modular in the right ways, like being able to plug in different pieces of the solutions that we're building. Like I want to use memory for this thing over here. Right. Or whatever else it is. And having a truly excellent developer experience around that because we spend a lot of time with enterprises who are like, okay, I have this like walled garden. I need to figure out exactly how I can plug these solutions in. And so we're, we've got a part of our team that's innovating on things like strategies and jobs and trying to help you maximize intelligence. And they're like, that's really cool. But I can't actually use any of that for X, Y, Z reasons. So I think solving those problems is really, really important to us. But then the other Persona is, you know, the like weekend developer who's like, I want to go and build something useful for myself. Right. And they're often doing that on top of our platform and on top of many other just pieces of developer platforms in the community. And I think for some of those folks there's more that we can do to be provide solutions that are maybe more open or more hackable or whatever it might be for those folks to kind of just like go wild with what we can offer them and have this really excellent developer experience. And so I think there's a lot of stuff that maybe I would put in the category of table stakes that I'm really excited about because I think those are the things that then unlock getting people to say, okay, yes, this thing works for me and now I can plug in on some of the stuff that you guys are doing that's really innovative and hill climby to get more intelligence and save costs and things like that.
Interviewer (possibly Lauren or another host)
Wonderful. Caitlin, Angela, I feel, I mean you are building one of the most important developer platforms in the world and talking to the two of you over time, I just feel really optimistic that that platform is in very thoughtful hands that care about the ecosystem. So thank you for taking the time today to share what you're up to and we look forward to what's ahead.
Caitlin
Thanks for having us.
Interviewer
Thank you guys.
Caitlin
Sa.
Date: July 14, 2026
Hosts: Sequoia Capital Team (Sonya Huang, Pat Grady, Lauren, and others)
Guests: Katelyn Lesse (Caitlin), Angela Jiang – Anthropic Platform Team
This episode explores Anthropic’s philosophy and strategy for building an open, developer-friendly platform, structured around layered abstractions for knowledge, execution, and coordination in AI systems. Katelyn Lesse and Angela Jiang provide an insider's perspective on how Anthropic supports both internal and external builders, why openness and composability matter, and how the platform evolves to keep up with rapid advances in AI – all while balancing developer needs, ecosystem standards, and responsible innovation.
(01:28–03:53)
“Internally, we really want to provide as much leverage as possible for our teams... Externally, it’s about giving any builder the tools to work with Claude and build whatever they want.”
— Angela, (01:51)
(03:54–06:13)
“If you over index on one or the other, you fall into a trap… We dogfood internally and open up early access externally to get a range of feedback.”
— Caitlin, (05:15)
(06:14–12:13)
“If you look at our roadmap ... we’ll move more and more from the knowledge layer, to execution, to the coordination layer in terms of abstractions.”
— Angela, (00:00) & (12:10)
(12:13–15:39)
“It’s like electricity ... transformative because you can wire it into everything. That’s not something anybody can do by themselves. They always have to work with the ecosystem.”
— Angela, (14:58)
(15:39–21:25)
“We are always trying to figure out a form factor... we tell the team, is this AGI pilled enough? ... Sometimes we build something, it works, and maybe it’s not the right next thing and so throw it away, try again.”
— Angela, (17:12)
(21:25–26:16)
(26:16–34:07)
“The interesting innovation is going to come more at that higher, meta level... optimizing strategies, the jobs you give to tokens.”
— Caitlin, (28:22)
(30:27–34:07)
(34:07–38:28)
(38:28–43:31)
“What you don’t want to do is stop AI usage... Instead, construct strategies that route tasks to the right ‘level’ … that’s where a lot of technical complexity but also real value lives.”
— Angela, (38:56)
(43:31–48:01)
"If we give this out to the rest of the ecosystem, there’s probably going to be a hundred thousand, two hundred thousand... combinations that people could put together."
— Angela, (46:02)
Angela (01:51):
“We actually kind of have like two North Stars... internal is about speed and leverage for our teams; external is about empowering any builder to create with Claude.”
Caitlin (05:15):
“We never want to fall into this trap of over-indexing on internal or external users… so we dogfood and open up early access at the same time.”
Angela (14:58):
“It’s like electricity … transformative because you can wire it into everything… you have to work with the ecosystem.”
Angela (17:12):
“We’re always trying to figure out a form factor... what’s awesome for one year’s AI development may not be for the next; so throw it away, try again.”
Caitlin (28:22):
“The interesting innovation is going to come more at that higher meta level… optimizing strategies, the jobs you give to tokens.”
Angela (38:56):
“What you don’t want to do is stop AI usage... Instead, construct strategies that route tasks to the right ‘level’ … that’s where a lot of technical complexity but also real value lives.”
Conversational, pragmatic, transparent, and highly informed by real developer/user feedback. The guests emphasize curiosity, willingness to experiment and discard, and humility—they’re clear that they don’t have all the answers and value what emerges from the ecosystem. Both Katelyn and Angela repeatedly stress empowering builders rather than dictating solutions, advocating for an open, collaborative AI future.
Anthropic’s approach is to advance the state of agentic AI not by locking down users, but by providing thoughtfully structured primitives, embracing interoperability, and learning from the world’s most creative developers. Their focus on composability, best practices, and modular infrastructure positions them as a pivotal force in shaping the next generation of AI-powered products—firmly rooted in ecosystem and user needs rather than proprietary walls.