
Loading summary
Host (Interviewer)
Okay, we're here in the studio with Akshay from OpenAI. Welcome.
Akshay (OpenAI Core Product Engineering Lead)
Thank you.
Host (Interviewer)
And with our trusty co host, Vibu. So you recently launched ChatGPT work. You lead core product engineering. You know, it's been a long journey into all this. I find it very interesting that you started with no code or low code with Walrus and Airtable. And to some extent, ChatGPT work is kind of like the super app of super apps of. Well, here is the ultimate no code. You just write a prompt.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, yeah. It's funny how things come like full circle. I mean, I think for a long time, my career. I mean, I started my career working consumer fintech, but then after that there's this hypothesis that the things that we were able to do with code as engineers, if we could bring that to many more people in a more accessible way, then that would be truly magical. We were working on a startup. It's actually funny, before LLMs, before vision, LLMs on how to do automated testing with AI, and it was just kind of jank back then, but doing what we can. And then worked at airtable for a while on the same thesis that if we can bring a database or the primitives behind a database to people, that'd be really useful to them. But once I think LLMs came onto the scene, it became clear that this was the missing piece, the missing technology required to bring the magic of code to everyone without them having to know what's going on underneath the hood. And so I think this launch and a lot of the stuff that we've been up to is the manifestation of that.
Co-host (Vibu)
How was stuff when you joined? So you joined OpenAI 2023. Now we've got, you know, so much more stuff, so. ChatGPT Codex app, ChatGPT for Work. Have things changed?
Akshay (OpenAI Core Product Engineering Lead)
Actually, I think the more interesting thing is how things haven't changed. Like, I guess like one I. I joined. I remember when I joined it was like 500 people. One thing I was worried about was like, I was looking for something, you know, more early stage and like, was it going to feel startup enough? And I joined, I was like, this feels even more startup y than I could ever imagine. And. And that really hasn't changed even till now. I mean, I think the level of bottoms up ambition and the ability of anyone to do anything or have an idea and ship it is really cool. But on the sort of mission side, I think what was really compelling to me is this mission of bringing frontieral intelligence to everyone, building AGI and then bringing it to everyone. And I think acknowledging back then that that vision is going to not be a linear progression. We're probably going to try different products and have different things that succeed and don't. But the vision has stayed the same and the mission has stayed the same and we're starting to see the pieces fall together and that's really cool.
Host (Interviewer)
You worked on enterprise. A lot of people never touched ChatGPT for enterprise. God. What is something that you learned from there that you're bringing into your work now?
Akshay (OpenAI Core Product Engineering Lead)
I think how there's no one size fits all solution in enterprise. I remember in the early days of ChatGPT Enterprise we would talk to customers and everyone. That was like when I think it was a year after ChatGPT was released and everyone was so excited to bring AI into their enterprise. And there was all these teams that were being stood up as the AI deployment team, these enormous budgets. And if you asked anyone what were they excited about, what were they excited about solving? At first you'd get kind of the baseline answers of we have all this context and data and all this stuff. But then if you ask them what was a discrete use case that they want AI to enable in their workplace, you get such a different variance explosion of different types of answers. And it's interesting using these models and these products, you have this box and you can say anything to it, which is the magic. But on the flip side it also means that you don't know what to do with it. And in the enterprise, I think a big part of that is actually meeting the users where they are, what use case are they trying to solve and then actually teaching them how they can use AI to gain leverage there.
Host (Interviewer)
Do you meaningfully differentiate that from forward deployed engineering or.
Akshay (OpenAI Core Product Engineering Lead)
I think there's like the like go to market side of it and then there's like the product side of it. I think you need someone to product side. Yeah. And I think like however good we get at FD motion, like I think at the end of the day, if we have a user who's like looking at their computer or looking at their phone, like it's our job in the product to like be enabling them and showing them where to go. So really excited about that.
Co-host (Vibu)
Do you think there's been changes, you know, over the past three years of adoption? So the there have been, you know, step function changes, you have reasoning models and whatnot. Is there still the same problems of enterprise has black box, don't know what to do with it or how things changed?
Akshay (OpenAI Core Product Engineering Lead)
I mean we're Seeing now that like, there's this huge update, right, everyone's extremely excited about it. It feels like, you know, many people, like millions, hundreds of millions of people are using ChatGPT. They understand, like, how generally to work with AI, but then like, every time, like a new capability gets unlocked. So now, like we're seeing with agents, like, there is probably a contingent of early adopters still who truly get it, who are like, you can do anything, you just have to make sure the right context is there, it's connected to the right tools, and then that you're supervising it, but anything is possible. But then there's this 10x or 100x bigger market, or they don't yet get that, or they don't yet see that. And so I think that's the next stage here. So I guess to answer your question, I think the adoption is there and growing fast, but I think the opportunity is far, far bigger than that. That's where we want to play, especially with ChatGPT work.
Host (Interviewer)
Yeah, well, let's skip ahead to ChatGPT work. Only like a month ago or so announced what was the sort of decision process that led into it. There was this overall merging of the super app. Is that what we're officially calling it? You deprecated the browser as well? Just, I guess summarize your last couple of months of working on this thing.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, it feels like forever now, but I guess it's only been a few months. I think maybe the one impetus that is most salient is when we released codecs or even internally had codecs. It was really surprising to us. I think we recently put out some stats on this, that there's this real inflection of adoption among non developers at OpenAI and through this product development process, would go to these UXR sessions to talk to people internally. And the thing that stuck out to me is like one like, you know, you go talk to like strategic Finance or marketing or whatever, and they're all using Codex for, you know, their use cases. That, that, that part's cool. But the thing that really stuck out to me is how proud people were that they were using Codex. Like how like, like I'm not supposed
Host (Interviewer)
to be using it, but it was
Akshay (OpenAI Core Product Engineering Lead)
that it was like that they were, you know, early to this, like, new thing. But it was also this thing of like, they felt like they had a superpower. Right. And what we recognize then is that like the power of Codex, the power of agents, we already have this massive distribution base of people who have come to know and love chatgpt how do we show that to them? How do we bring it to them? Which is a hard product problem and it's a tricky thing. There's many ways you can go about it. And so that's what we called the merge and the super app over time and ultimately launch in ChatGPT work is how do we do that? But it came from that initial realization that the power was not only for developers much, much earlier than probably even we thought, like it could be extended to everyone.
Co-host (Vibu)
How do you see the products differently? So like, who is it for? Right, so Codex started out even CLI10 app. Now there's a merge of ChatGPT Codex and ChatGPT work. So is it the opening for the average user for enterprise for work? How do you position it?
Akshay (OpenAI Core Product Engineering Lead)
I think we want to get it to position it for if you're doing quirky related things, for lack of a better word, I think productivity is like is actually what, like the pillar that I support, like that's the name of the team and the reason for that, the reason we call it productivity and not like, you know, enterprise or like work or something like that, is because there's also personal productivity. Right? Like I think ChatGPT work is. I've seen people do things in their personal lives that you wouldn't classify as like work technically. But like these agents are, you know, super capable for like one. One recent example that someone posted about on our Slack is like someone had like a missed package, like they didn't receive it and then they got like the picture of it, you know, from Amazon or wherever the courier was and they like asked ChatGPT work to like find out where that package is. And like the agent, you know, is extremely tenacious. And like, like took the image, like looked at a bunch of like listings around their neighborhood and figured out exactly the apartment complex in which the package was like, gave them some information. And so like I think there's all these things that like you, you know, work E or productivity related things. I think that's what we want the product to be. Because asked about Codex, I think we think Codex is a durable brand. But we have a principle that the user, we don't want a user to get stuck in a tab or an experience where they don't get the power of the product. And so basically everything that you can do in the Codex portion of the product on desktop, you can do in ChatGPT work and vice versa. But we made some opinionated product decisions on how much of the Git state if you're in a git repo. Do we want to expose to the end user or how much do we want to make the experience of seeing the agents thinking diff forward so that you get exposed to the diffs out of the bug? And then on the safety side, how do we want to think about sandboxing and making sure that we have the right defaults in one state versus the other? So there's some opinions that go behind that, but we don't want the user to need to choose which experience they're in.
Host (Interviewer)
That is a good goal for AGI, right? People don't want to choose what version of AGI they want, they just want the AGI to decide for them. Can I get an answer or like it's not super clear to me. Is the codecs harness and the ChatGPT work harness the same? Is it just UI affordances or are they actually prompt level or even deeper differences?
Akshay (OpenAI Core Product Engineering Lead)
So the harness is the same? The harness is shared in both of the products. We made improvements to the harness to make it good for knowledge work, especially as it relates to plugins or computer use or artifacts. You get that power regardless of what your experience you're in. On the UX side, there's opinionated takes that we have when you're in Codex mode, what the UX should be, how the UX should behave and some stuff around the sandbox like I mentioned, but the underlying harness and capabilities should be the same.
Host (Interviewer)
Actually, I'm just kind of curious. Maybe we can. Is there a query that we can run that would look different in the two modes?
Akshay (OpenAI Core Product Engineering Lead)
Yeah, try to create like, ask it to create like a retirement calculator, spreadsheet or something in both modes. And in Codex mode you might have to be in a repo for this, but you'll see the diffs of the sheet that it's creating and stuff like that and the file edits. But in work you won't be able to see that.
Host (Interviewer)
I think that's super clear. And then also the other thing I wanted to dive into was the productivity team. What else is there? First of all, what are the top level teams other than productivity? Isn't productivity everything?
Akshay (OpenAI Core Product Engineering Lead)
So you know, we have a team focused on, on ChatGPT, like the, the core chat experience for consumer which is like, you know, not, I think all productivity. Like there's people are using ChatGPT every day for search to, you know, figure out how to write messages to loved ones to think about, how to like learn a new topic, etc. And so there's so much more inside to create images and there's so much more in chat that, you know, the 100 hundreds of millions of users are using that. Obviously, that warrants a very dedicated effort. And there's teams focused on enterprise and infrastructure and API and stuff like that as well.
Co-host (Vibu)
I will bring it up. Yeah. So I have them both running. This is work. There's a Codex version here I picked 5.6 SOL. So this will take a while, I think. We'll just keep it in the background and as they finish, we'll look into some of the differences.
Akshay (OpenAI Core Product Engineering Lead)
Yeah. But immediately, I think if you flip back to the codecs version, you'll see that that assumes git exactly. Like Dynamic island assumes that you're in a kit repo and you might miss some stuff because some of it is, like, in the actual chain of thought. But what those changes and how we display that.
Host (Interviewer)
But is there an unintuitive. Like, is there a thing that you wanted to ship and then you got feedback and you were like, no, let's not do it. Like, what's the thinking behind that?
Akshay (OpenAI Core Product Engineering Lead)
In Tribute of work.
Host (Interviewer)
Yeah.
Akshay (OpenAI Core Product Engineering Lead)
I think one direction we could have gone with this is like, keeping the experiences, like, completely separate. So it's like, why different apps? Exactly. Like different apps or even in the same app, like, different. Completely different experiences. Like, why merge at all? What does Codex, obviously, people love? Why bring these products together? And I think the intuition here is that all of our jobs are changing dramatically with AI. Every few months, I feel like I wake up and I'm doing a completely different thing than I was doing a few months ago. And my hypothesis here is that our hypothesis is that, like, part of what we're building in this technology is giving people leverage. Like, you know, the things. Maybe it's the more mundane parts of your job or parts that, like, if you were able to automate, you'd be able to share more ideas faster or whatever, like you're able to do now. And because of that, like, that might actually blur the lines between someone who's, like, only writing code or creating strategy docs or planning events or helping with marketing or doing podcasts or whatever. Right. And so, like, these things are going to get blurred over time. And so trying to draw a hard boundary based on who you are is going to be tough. And we should enable users to choose, but we shouldn't box them in. And so a lot of the work that went in here, keeping the primitives the same, for example, plugins are unified across this product and chatgpt in the cloud was because of that. It's this thesis that eventually things are going to come together and we don't want to be like. We want to be prescriptive about when to be in either experience, but we don't want to box anyone in.
Host (Interviewer)
I wonder if there's users who very tuned to the old ChatGPT harness that is effectively now replaced by the codecs harness. I can't imagine what that was, but maybe they're more. The more conversational side. Can you compare and contrast the two harnesses because only you've seen it.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, I mean, I think chatgpt, the existing harness, still exists today. It exists in this app, the classic. Right.
Co-host (Vibu)
You just start a new chat and you don't call under work, right?
Akshay (OpenAI Core Product Engineering Lead)
Yeah. If you start a new chat and go to chat, then you're talking to chat.
Co-host (Vibu)
We can technically do another, I guess, you know, on instance.
Host (Interviewer)
Yeah. So this one's not going to code or it's going to be inline. It's not in line in a sandbox.
Akshay (OpenAI Core Product Engineering Lead)
Actually, when you try to push you to go to work, if you're creating
Host (Interviewer)
a spreadsheet decision, sorry, is a router
Akshay (OpenAI Core Product Engineering Lead)
decision, this is the decision that, you know, the model is making and then like, you know, it sees that you're able to, or you're trying to do something that would be better served in work mode. But I think your question was like, what are the advantages of like the, the chat, like ChatGPT chat harness.
Host (Interviewer)
It's more broadly like I want to basically do an oral history of harness engineering.
Co-host (Vibu)
Right.
Host (Interviewer)
You know, the ChatGPT harness lasted us from, let's call it the 01 era until now and now it's being replaced by the codecs harness effectively and they're overlapping somewhat. But I'm curious what changed if, you know, if there is.
Akshay (OpenAI Core Product Engineering Lead)
My perspective on this is like there's is, there's, there's sort of like a, a constant process of like divergence, convergence. Divergence, convergence. And in chat, like many of the use cases I was talking about before, like, you know, search or learning. I think we're really optimizing for latency and optimizing for personality and like different things that over time, like the product. The reason people love ChatGPT is because we've been optimizing for those things and working on them for so long. Codex. What we learned was that if you give the agent access to this infinitely flexible environment as a computer, it can do really, really powerful things. And so when we think about. Okay, well, for knowledge work, which mode should we choose? It felt more natural to us to bring that to this computer environment and maybe abstract some of the details of this computer away from users who might not be used to that, but give them that same power. But ultimately I think that we want the power in all places, right? We want to meet people where they are. So I'm sure there'll be work down the road in order to get things to be equivalently capable in all scenarios. But it's just a question of what we've been focusing on the product on historically and what we're focusing on now.
Co-host (Vibu)
I think alongside that, outside of just Harness and when to use Codex, ChatGPT or work, there's also the new models you've released, right? Any guidance there? So people love to min max what to use. Like only use Terra on high reasoning versus for this, you know, you want to use SOL here ignore there's 32 options. Yeah, yeah, yeah. But that being said, you know, for people that are expanding so productivity, trying stuff for work that don't have the breakdown of what all this is, what's the advice? Right.
Akshay (OpenAI Core Product Engineering Lead)
Well, I mean, I think before the advice, the first thing is none of this would be possible without these models. I think you asked earlier what was the inspiration for work, and early on I mentioned what we were seeing with codecs, but that was also because the models were getting infinitely more capable. That's happening again. I think it's another step function jump now. And to answer the question on advice, we want this default to be the best possible. We want to be opinionated about the default. And so we've chosen a default that we think is going to be the best for everyone and we have for power users, options under the hood. One could argue that there might be too many right now, and we're working on simplifying it. But you can extend the reasoning level and you can change between the different model classes if you need to. But the default should be the best for most use cases. My advice to most people would be to stick to that. And then if you reach a situation in which you think that you want to try a different configuration, if you're not seeing either the efficiency on the cost side or the quality on the intelligence side, then you can change the defaults and see if you can get something better. But we think that the default should be good enough.
Host (Interviewer)
I'm just going to run something by you since you have way more experience than me. I've recently been doing Soul Light, but with Goal with the idea that the goal basically augments the reasoning effort, but with more terminations and turns. Is that a good way to think about it as opposed to Soul Ultra or Soul Extra High?
Akshay (OpenAI Core Product Engineering Lead)
Yeah, it's hard to say because it's
Host (Interviewer)
like an interaction effect.
Akshay (OpenAI Core Product Engineering Lead)
Exactly. It's like there's a preference for you as an individual. How do you like to collaborate with the models? How many of those terminations, as you call them, do you want where you can steer or make sure that it's doing the right thing? I think generally people should try whatever works for them. I think that using Ultra or the multi agent setups are best for when you have tasks that are either incredibly complicated, open explorations or very parallelizable. I think even for tasks using Goal, I think is best for tasks that you know that you'll be able to make consistent progress in a way that's verifiable over time. But I think for most tasks they actually don't fall into either of those buckets. And so like at least when they're starting. And so that's why I think the best first step is like trying it with the default configuration and then seeing like where you want to go from there. Right.
Host (Interviewer)
You guys worked on a slider, which actually is super helpful for reducing the amount of panic.
Co-host (Vibu)
It's nice on mobile at least there's a nice slider.
Host (Interviewer)
It's nice here.
Co-host (Vibu)
I haven't tried it.
Host (Interviewer)
So you have the advanced view there. But if you click Advanced View.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, yeah.
Co-host (Vibu)
Just a nice lighter.
Host (Interviewer)
Yeah. Very, very pretty, Very colorful.
Co-host (Vibu)
Yeah.
Akshay (OpenAI Core Product Engineering Lead)
The idea was here was like, reduce it to like one dimension even though there's multiple dimensions. Right. Try to project it onto a single dimension for the user. You have like, you know, something from that represents like, you know, speed and efficiency on one side. Yeah. And then like sort of like quality and thoroughness on the other side.
Host (Interviewer)
I am just puzzled that it uses Soul so much.
Co-host (Vibu)
Like the lower spider, if I'm not mistaken, is. Oh, it did.
Host (Interviewer)
So they preset Terra to only be the light one. But I think a lot of people, actually more people should use Terra1 because Seoulki is running out of capacity.
Co-host (Vibu)
I'm the reason. Here's 10 minutes of our retirement calculator.
Host (Interviewer)
Oh, that's the Excel thing working. Oh my God.
Co-host (Vibu)
This is work. And then Codex is still cooking, so we'll get back into it. I think it'll be interesting to actually see the thought process, the reasoning and also, you know, I guess this is eight minutes on work. Codex is still cooking.
Host (Interviewer)
Yeah. And by the way, so I. Do you know Gabriel Chua, he's part of the opening Singapore team. He showed me this and I was like pretty shocked that this looks like Excel. It edits Excel files. You never paid an Excel license.
Akshay (OpenAI Core Product Engineering Lead)
Right.
Host (Interviewer)
Like, but somehow this is like kind of workable and it's agentic Excel. Yeah.
Akshay (OpenAI Core Product Engineering Lead)
I mean one of the biggest pushes that we made for this launch was like artifacts, both on the model side. I think if you Compare this with 5.5 and 5.4 before that, you'll see that there's been pretty dramatic improvements in the quality of these artifacts. And then also on the product side,
Co-host (Vibu)
the UX side is also crazy hosted sites and whatnot. No longer needing to host your own little web page.
Host (Interviewer)
Oh, I have a story about that. I can do a separate thing. I'll need to take the, the visuals here, but we'll cut to that later. Was there co training, I guess because you were making this big move and you launched 5, 6 on the same day as ChatGPT work. Was there influence between the model training teams and the harness teams or did they. The launch dates just happened to line up the same day.
Akshay (OpenAI Core Product Engineering Lead)
I think we collaborate heavily with the research team and I think that's one of the most magical parts of the job. The most fun parts of the job. But yeah, I mean just using artifacts as an example, a lot of what you're seeing underneath the hood, there's a lot of work that went into making sure that we had the right infra to be able to train the models to get better at this and then on the product side had the right experience for users to be able to collaborate with the model on an artifact like this. In fact, this whole viewer, the intuition here is that it's not necessarily that you wouldn't need an Excel license. This is stage one, right. This is probably not what you meant when you were making a retirement calculator. You want to iterate and like when you, when you're seeing it and if this thing is high fidelity to like what you would actually see in or what your, your co workers would see if you were to send this to Sean like that. That I think makes it so easier and makes you trust the product in terms of iteration.
Co-host (Vibu)
When you say co workers would see, do you see a multiplayer multi team collaboration with artifacts? Any, any things you guys think about,
Host (Interviewer)
you can already share it, right?
Akshay (OpenAI Core Product Engineering Lead)
Yeah, yeah, it's. It's something that you know. Right. Actively thinking about one thing that we've noticed internally without talking too much about the roadmap is that there's many times when someone will ping me about something and I will ask ChatGPT work the question and then I'll ping them back the answer. And then I'll be thinking the syncless
Co-host (Vibu)
would be the three of us are just all on one hosted.
Akshay (OpenAI Core Product Engineering Lead)
Exactly. I'll think about was I required in this loop? And maybe it was rephrase what they were asking or pulled from certain context or whatever. But when I gave them back the answer, that process was also lossy. I gave them just my interpretation of what ChatGPT work cooked up. But, like, underneath the hood, there's so much context, like in the rollout and stuff that could be interesting.
Host (Interviewer)
So like the answer was preemptively respond to every inbound request.
Akshay (OpenAI Core Product Engineering Lead)
No. And it's just like, literally like this is what I do sometimes with my job.
Host (Interviewer)
I know you copy paste and then you're just a message forwarding service from AI to AI.
Co-host (Vibu)
But I think it's interesting, right? It helps people understand the capability of what you can ask and delegate, that oftentimes people don't realize until they try to or someone shows you and then you're like, oh, okay, okay, I see.
Akshay (OpenAI Core Product Engineering Lead)
Yeah.
Host (Interviewer)
I think there's also like a light security issue where like, basically you're the permissions layer. Like, yes, I could query everything that you query and I could get an automated response, but maybe I'm not supposed to see it.
Akshay (OpenAI Core Product Engineering Lead)
Yeah.
Host (Interviewer)
And there's no way I would know because I'm not supposed to know what I don't know.
Akshay (OpenAI Core Product Engineering Lead)
Especially as like, you know, with ChatGPT work, we're asking you to connect your plugins and, you know, it's pulling from your local files and stuff like that. Like the amount of context that the agent has access to is like deeply personal. And that's something that we need to preserve. So that'll be definitely a challenge.
Host (Interviewer)
There's Excel, there's PowerPoint, there's Docs, the grand trio of work. What other formats of work do you think about? Obviously you worked on Airtable. Is there a future where there's OpenAI Airtable? What does that look like if you ever ended up doing it?
Akshay (OpenAI Core Product Engineering Lead)
It's a really good question, I think. I mean, one that you didn't bring up was sites. And I think that was the core part of this launch. There's one one side of sites that I think people commonly talk about, especially on Twitter and stuff, or X of like, you know, this sort of like prototyping tool and actually like we saw that happen with this launch. Even the, the model slider that you guys were referencing earlier, like that was developed almost fully in a site like, you know, the, the collaboration between design and engineering and product on that was like on a site where we play with the affordance and figure out how it feels and all of that. But the other aspect that I think is a little bit less talked about is cites as an artifact for knowledge work. I was actually talking to someone the other day who was on our corporate finance team and they were mentioning how now when they have these reports that they're working on as a team month to month. Historically those things were in slide decks and in spreadsheets, and now they're just in sites. Sites is the mechanism that they collaborate across the team. And the reason is because I think it's somewhat higher bandwidth. These tools like PowerPoint and Excel are infinitely flexible, but at some point you reach the boundary of either as a human, you may not know how to use some feature or something, or the product itself doesn't support it. But with a site you can kind of do anything. You ask for anything, and you can get that. Once people see that magic, I think it's been really valuable. Yeah.
Host (Interviewer)
Let me show you my case study. This involves all the hot topics, including ChatGPT work, but also 5.6 token billionaires and token maxing and sites and auto research. I'm a fan of this game called Strada. Basically it's like a little board game that you play with physical blocks that come on top of it like that. So over the weekend I took like 30 photos and just threw it into ChatGPT. 1.7 billion tokens later, out comes this site with a fully playable thing with 3D block placement and everything, because it requires physical blocks and I needed friends to train on it so they can get better, so I can play against them. But also I could also do things like train an AI on it. And that's your auto research that gets into auto research. So you want to train your own AIs and then make sure they self play against each other. I need to set both AIs. So this is AI versus AI and they're going to self play. Obviously the AI start out bad and then you want to define a loss function and get good. I wasn't going to supervise all this. I was at Dallas San Mateo attending a conference. What I ended up doing was auto researching on this and creating benchmarks. And there was just way too many parameters for me to read. So I started asking it for a site and it's created this lab panel. Is there a shortcut for a site that is created?
Akshay (OpenAI Core Product Engineering Lead)
You should be able to go in the sidebar to Sites. Top of the sidebar, the left sidebar.
Host (Interviewer)
This one. Oh, left.
Co-host (Vibu)
Yeah.
Akshay (OpenAI Core Product Engineering Lead)
Just scroll all the way to the top. Oh, oh.
Host (Interviewer)
It says sites. Oh, there you go.
Akshay (OpenAI Core Product Engineering Lead)
Yeah.
Host (Interviewer)
So it creates the sites. I don't think this is exactly what I wanted, but let me show you what it popped up. Right. I think as a research artifact, it is very important to communicate exactly what is being done outputs. This thing, which I eventually started publishing, so I moved it off of sites because I wanted more database and infrastructure than sites afforded me. But this is like a research output that you can start to mess with and try to think about what hyperparameters are you tuning for training your AIs. And I was trying to make scaling laws and everything and doing all sorts of game optimization stuff. And the fact that you can just kind of throw this up as a research artifact. Like I no longer need to read ChatGPT output. I read site output. But then there's also a huge sprawl. Like, look at how long this thing is. There's so many numbers. It is pretty overwhelming. So then I have to start pruning from there. But it's an interesting transition from markdown actively that you're putting out to. You're putting out a whole functional site.
Co-host (Vibu)
I think markdown just isn't that optimal for people to read. Right. Might as well just write HTML website. And I don't know, I think you can do a lot with customizing this. Right. You have your skills that explain what you want. Like I noticed they're quite verbose. I don't need a lot of this information.
Host (Interviewer)
It's very verbose.
Co-host (Vibu)
So. And then the nice thing of having a site side by side is, you know, you just iterate on what you want and what you don't. Right?
Akshay (OpenAI Core Product Engineering Lead)
Yeah.
Host (Interviewer)
I don't know if any that triggers any stories for you of how it's run internally. Am I doing this right?
Akshay (OpenAI Core Product Engineering Lead)
Yeah, I mean, I think that this is like a workflow that we're seeing like all different types of teams use where like the. The canonical artifact that was previously a deck or something is now becoming a site. And like with a site you. Because it's just HTML, you can like, it's infinitely flexible. And so, you know, if you. If you want to give more prominence to a certain thing that like in a slide deck would, you know, feel like it was buried like you can do that. You can have it be like the hero image. Right. And so I think that people are starting to see that there's obviously more work to be done to make these things much more easy to collaborate on. You mentioned that they're long and verbose could be broken up. I'm sure there's something super long. Yeah. But I think we're starting to see that there is. This aspect of this is a really interesting format for people to use that's much more flexible than what they were had before.
Host (Interviewer)
I think your job also becomes kind of meta. You're not designing the products, you're designing a product to make products. And I'm curious how you manage that.
Akshay (OpenAI Core Product Engineering Lead)
I think one thing that we've been like when we look at the UX that we've been thinking a lot about, is how can we balance simplicity with capability? If we're designing a product, like you said, that is made to build other things, you can build so many different things, but we can't put that all in front of you because you'll get overwhelmed. And so we had similar problem or similar challenges even with ChatGPT, but especially now when there's so much that can be done. I think the balance that we're constantly trying to strike is how can we give the user enough of a UI surface where they can be expressive, they can tell the agent what they need, they can verify that it's using the right tools, it's pulling from the right sources, etc. But then it gets out of the way and then how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is going to be like, how do they discover the next use case and the next one after that if they really want to be super powered by the AI.
Co-host (Vibu)
Yeah, it's interesting. I feel like everyone also just has a different way to do it. Right. I made a similar version of this same game. I didn't take any pictures of board or rule game. I threw an OAL 18 minutes, 53 seconds later, lot of tokens later, I've got a similar version. Obviously not with all the other research and whatnot, but, you know, you got
Host (Interviewer)
to do all the latest trends.
Co-host (Vibu)
And yeah, I did it with Codex. Not work. But it's interesting, right?
Akshay (OpenAI Core Product Engineering Lead)
Yeah.
Host (Interviewer)
And this is obviously GPT image generating the pro avatars. Very good for game design. Like a lot of game designers were like, really into GPT image.
Co-host (Vibu)
I will say, like, the broader takeaway probably is the reason that we do. This is more so just to test the tools. Right. Like this was also a test for 5.6 came out. I had done the game on 5.5. Right. The ability for me to no longer need it to. I had to feed it the rules. It's a pretty niche game. It couldn't find how to do this on its own.
Host (Interviewer)
Oh, yeah, 5.6 it is auto distribution. That's why. Also very keen on testing the 5.6 capability.
Co-host (Vibu)
But, you know, this is just as work comes out, as new things come out. These are just our sideways to test things.
Host (Interviewer)
Right.
Akshay (OpenAI Core Product Engineering Lead)
Yeah.
Host (Interviewer)
It's some kind of private evo, I guess that is not less private, but
Akshay (OpenAI Core Product Engineering Lead)
also valuable because now you can send this to your friends and I mean, I learned about this game through seeing this.
Co-host (Vibu)
It's a hard game. He's very good.
Host (Interviewer)
It's good when no one is competing with you. But yes, it's a classic RL problem of self play bootstrapping your game AI. Yeah. You see how easily work becomes personal and personal becomes work. Because the thing I do for personal, it actually directly informs people, people I work with. Because I showed it to them, they were like, oh, you can do that with GPT. Which I imagine is the growth strategy.
Akshay (OpenAI Core Product Engineering Lead)
Yeah. The show not tell is a big piece that, you know, I think we're not still not fully cracked of, like, you know, showing people all the things that they can do with the product versus like trying to teach that to them, like, you know, articles or onboarding or whatever.
Host (Interviewer)
Yeah.
Akshay (OpenAI Core Product Engineering Lead)
Meeting them in the moment, it's a
Host (Interviewer)
career risk for me because I used to be in developer relations. Right. Where your job is to show. And then you're like, what do you mean you don't need. Actually your job is to tell. But the product people are like, well, we don't need you. If our product is intuitive enough.
Akshay (OpenAI Core Product Engineering Lead)
Yeah. I mean, that's the magic of the models. You can tailor the telling or the showing to specifically what the user needs, what they care about, what they've done in the past, exactly where they are on the adoption journey. So I think that's going to be a super big opportunity.
Co-host (Vibu)
Seems easier and easier now to tailor custom showing. Right. People have different use cases. As much as you said, you don't want to segment different people into different buckets. Right. It's also not that hard to. For people that are in different categories. But the question I guess is you said your team is more broadly on. What was the term you use? Productivity.
Host (Interviewer)
Yeah, productivity, which is now work, basically.
Co-host (Vibu)
Is it work? Is there another distribution that we're not hitting? Is there a group of people that will have something different than ChatGPT Codex or work? Is there more that the mass isn't targeting?
Akshay (OpenAI Core Product Engineering Lead)
I see it as like a sequencing. The vision is bring useful agents to everyone. We started with developers. Developers historically are early adopters that are willing to put up with more friction, set things up, et cetera. That's where codec started. I think the next opportunity is what we call general knowledge work. All the other functions around developers. I think when you go from developers to this segment, there's inherent challenges obviously with the show not tell thing that we're talking about making the product more understandable, bringing in new capabilities that matter more for this cohort than that for developers. Things like artifacts, things like computer use, et cetera. And then I think the same learnings similarly how we took the learnings from developers and brought it to general knowledge work. The next stage will be taking the learnings from general knowledge work and bringing it to everyone, no matter what they're doing in their lives. And we're already seeing that a little bit like this game example that you have is something that's on the border of fun and personal life to your professional Life. I use ChatGPT work full time at home for everything, for whatever I'm doing. I used it the other day to come up with a meal plan and save that on on the computer environment that it has and something that I can continue going back to. Is everyone doing that yet? Probably not, because the things says work on it, but eventually we want to get people there.
Host (Interviewer)
ChatGPT Life.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, exactly. ChatGPT Cooking. But I think there's a lot of opportunity there. But I see it as we built a foundation in software engineering and we're going to take the same learning. So we take from software engineering to knowledge work, Knowledge work to everyone.
Co-host (Vibu)
Do you have any power user advice? I feel like there's a group of people that will live it, use it for everything, stay on it 24 7. And then there's a bit of a gap between that crew and people that, you know, okay, I use it for work. I use it occasionally, Sometimes I pipe questions. Any advice, any learnings, anything you recommend or just, you know, takeaways that you found that help bridge that gap.
Akshay (OpenAI Core Product Engineering Lead)
I think a couple of things that I've seen is like one that it really helps to broaden your imagination of what's possible. And this has been a learning even for me. The technology has progressed so fast that it was something that even three months ago, no way the models can do this now. It's like, wow. It actually can give an example. We're going through right now our review cycle internally. And people always talked about this as kind of a. A thing that the, the models are good at. And like, you know, there's a cliche of like, okay, like no one wants to be writing reviews and like, we just use AI to do it, but I mean, and then you have to
Host (Interviewer)
evaluate it as well.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, exactly. In all seriousness, before it was like, just like slop basically. And like, I think it was helpful, but, you know, not super productive. Now I've found that like, the model can do a much, much better job than me, especially in this environment of like pulling contacts on, like, what people are up to, how they've, like the things that they've done to make a difference, highlighting wins that they've had that I may not even have seen, has access to everything, the code, things that they've caught, reviews, slack, everything. And so it's incredibly powerful in that domain. And just six months ago, the last time we did this cycle, I tried using it, but it was not at all helpful. And this time it's been incredibly helpful. I think continuing to push the frontier of imagination what's possible, even if you tried something before, I think is maybe my biggest piece of advice. The other, I guess, thing is the more you put in, especially in this environment where the model has access to everything on your computer or in ChatGPT work, you can create artifacts over time and save them in your library, and the model will continue having access to those. The more information you give it about whatever domain you're in, whether it's your life or your work, the more valuable it becomes. And it'll become valuable in ways that might surprise you. It might pull from context in a way that may be proactive and that you might not even have thought about, but. But it needs to have access to those tools or that context first.
Host (Interviewer)
One thing I just want to talk about the review stuff because I still. That's a very sensitive thing. And you're a founder, you've managed people, you've hired people. As manager myself, I'm very reticent to put out any LLM generated things, especially when it comes to people, because it feels like you don't care. Presumably at OpenAI, people are obviously more open to being basically rated by GPT. Uh, but are there any unofficial rules around this? Like, what's the etiquette?
Akshay (OpenAI Core Product Engineering Lead)
Oh, I mean, I think the Etiquette is that like I wouldn't ever write something via like solely via AI and like present it as like a review for someone. What I was talking about is more like gathering context. Yeah, that's the place where it's.
Host (Interviewer)
So it's just search, it's agentic search,
Akshay (OpenAI Core Product Engineering Lead)
like agentic search, but you know that you can tailor and steer much more capably than you could before. And because like the thing is it's all. There's sort of a flywheel happening. Right. Because of Codex people are able to do and because of tragedy or people are able to do so much more now than ever before. And if you're able to do so much more, it's easy to miss things as well. And so like I think we need to use these same tools to keep up with all the impact that people are having and understand, you know, where it can be helpful.
Host (Interviewer)
I think that the thing like obviously I run a small company, so easy to search. But at the scale of OpenAI, with the amount of messages that you guys put in Slack, do you think that it misses things?
Akshay (OpenAI Core Product Engineering Lead)
Probably, but I think that I also miss things like it doesn't matter.
Host (Interviewer)
I think sometimes it needs to be human level.
Akshay (OpenAI Core Product Engineering Lead)
It's all relative, right?
Co-host (Vibu)
Yeah, sometimes it's nice when it finds things you wouldn't. Right. Like right now my codec system prompts, they're set up in such a way that every project I have has a separate notes MD and it just writes learnings to there and then the global one can pull from all these. So sometimes it'll be like, oh, there's this project you did like four months ago, here's a note that we had. And it randomly pulls it back into context I would never do, I haven't thought about. And I'm like, okay, this is quite superhuman, right? Like stuff that would. And you know, it'll save like hours on chunking of stuff or find something that's already been done. I'm like, as much as it might miss stuff, I would too. But it's very useful when it finds stuff. And I have like a very, you know, non super engineered solution to this. It's just markdown files that get pulled whenever they want.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, I actually have a funny anecdote about this. Like recently gearing up to this launch, you know, the team has been, you know, really cooking on it for, for, for a couple months and over that time like there's so much conversation, chatter going on in Slack and docs and elsewhere and one, one of the members of the team set up this scheduled tasks like automation to like, look at everything that's going on and like, come up with the best memes and then post it in one of our shared channels. And like, there's two cool things about this. Like, the first is like, I think the models are over time actually starting to become funny, whereas a year ago that was not at all the case. The second is it was what you were saying. They find things in surprising ways that you may not have thought of and create connections that you may not have thought of. And that really helps with the meme generation because then you can see something that genuinely surprises you and is funny in that way. So, yeah, I mean, obviously that's not the most productive use of the technology, but it doesn't cover this capability that's emerging, which is just to find information that you otherwise would not know of.
Host (Interviewer)
Talk about the launch. I think I have pretty much said this is the most successful launch in a long time. I think even more successful personally than 5.0. And they're announcing 10 million users. Does it feel different? You've been through a lot of launches.
Akshay (OpenAI Core Product Engineering Lead)
I think it feels like a combination. Well, I think two things. One, it feels like a culmination. Like I was mentioning earlier, this vision mission that we've been on for a long time. Like I said, we saw the magic of codecs internally and you're extremely excited to bring this to many more people and to see it working, to see us reach the distribution goal. Numbers that you mentioned, I think that's huge and super exciting. The flip side of that is there's so much more to do too. That's also really exciting. ChatGPT as a whole, this product that everyone almost equates to AI and loves, has hundreds of millions of users. And so 10 million is really cool. But we need to get this to everyone. We need everyone to feel this magic. And so that's the next step from here. But yeah, I think extremely pumped about how it's going so far and the opportunities.
Host (Interviewer)
Awesome. I did want to also, because I've been tracking the number closely. It transitioned at some point from just Codex users to Codex plus ChatGPT work, obviously, because it's same harness. The whole point is that you don't. You can't count them separately. You have roughly a billion ChatGPT users. Why did it just jump to 1 billion right away? Isn't that the default on ChatGPT or.
Akshay (OpenAI Core Product Engineering Lead)
No, we don't default you into ChatGPT work if you're on ChatGPT if you're free. It's also only available to paid users right now. And I think there's a process of educating users of what is the value of this product, having them try it, learning from their feedback and making it better over time. But I mean the goal is to get as many people who love ChatGPT today to feel the power of ChatGPT work. But I think it'll be a journey.
Host (Interviewer)
Yeah, Codex will still be alive as a brand for the foreseeable future and we'll just toggle between them as needed for UI stuff.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, I think it's even stronger point than that. I think we fully intend to treat developer. Developers have been a core market for us for so long and there's, there's so much more that we can do to make Codex great specifically for software development and we'll continue to do that. This doesn't take away from that at all. If anything, it should increase the utility of something like Codex because now you can move seamlessly between writing a diff to creating an artifact or doing a search over your character.
Host (Interviewer)
I do wonder how much this terminology leaks to the non technical user. Do they have to learn to say artifacts if I want artifacts or you
Akshay (OpenAI Core Product Engineering Lead)
know, it's funny, like we call artifacts internally because that's what the teams call externally. Like no one says that, no one calls it an artifact. But I think that people like often like describe things whatever they're used to. Right. So if you know ChatGPT is good at creating slides, let's say ChatGPT is good at creating slides and that's actually what we want.
Host (Interviewer)
One big. Another, I mean It's July of 2026. One big thing that also happens for OpenAI was OpenClaw. And that's I think a lot of people's first time really maxing an agent for personal stuff, but also crossing over to work in same way. As far as I understand, openclaw is still independent. But did you go through your own OpenClaw moments? Were there any lessons that you took from OpenClaw to Codex or back, whatever?
Akshay (OpenAI Core Product Engineering Lead)
I think there's a lot of inspiration. I did go through my own openclaw moment.
Host (Interviewer)
Yeah, tell the story.
Akshay (OpenAI Core Product Engineering Lead)
Me and my, my wife set up an open claw to try to manage everything in our house. Not that there's a ton, but it was actually quite useful. We gave it a calendar, it started creating events for us and stuff. At some point the laptop that we were running on it died and never got a chance to pick it back up. But there's a lot of inspiration there. In chatgpt work in web and mobile, you get access to this persistent computer environment where you can store vials and those files stay around between sessions. And the idea is to be able to enable use cases like this. One of the members of our team actually uses ChatGPT work for what they used opencloth from before and has completely transitioned, which is workout planning and meal tracking, which again, it's a worky thing. It's not work necessarily, but it's in personal productivity space. But it has all the same primitives, it has scheduled tasks, it has the ability to store files on a file system, it has the ability to reference those things over time. And so you start to see the same types of use cases emerge, which has been really cool.
Host (Interviewer)
Is there a point that ChatGPT work completely replaces OpenClaw? Obviously they're independent, so yeah, I mean,
Akshay (OpenAI Core Product Engineering Lead)
I'm not close to it, so I can't speak to the openclaw roadmap, but I don't think so. I think that there's going to be. There's always a need for this incredible open source technology that team has built and I think that we can draw inspiration in the product and ChatGPT. I think many more people have heard about and used chatgpt than have used openclaw. And if we can take the magic from openclaw and bring it to them, I think that'll be a success. I think that one thing on the ChatGPT work side that we feel strongly about is that the core experience is that you come to this product and you have a conversation, start a session, whatever you want to call it, with this agent. And the magic of the product is that you can do anything in that moment. And we would like to create a product where you don't have to click a button or to go to a different place, whatever, and you can get whatever functionality exists in your finances, app or any other product in this one place. And so that's the goal. It's like we want an extensible system with plugins where you can connect to the tools that you need in order to be able to accomplish a financial task where you can, if you're doing science work, we have an ability to extend the system and such that you can like write latex and it performs well. There will always be like, products that we support that are best in class at those things. But we want as much of the magic as possible in that core experience.
Host (Interviewer)
Yeah. Do you think that you can do everything you used to do with Wealthfront in ChatGPT finance.
Akshay (OpenAI Core Product Engineering Lead)
I actually tried it. I mean like ChatGPT doesn't yet custody, cash and assets for me. So that part, no, not yet. But I mean there was like a whole component of like retirement planning and sort of like financial planning and budgeting and stuff that we were looking into when I was there. And like with the finances plugin, like that's all possible today. So I feel like at least that component's replaced for me.
Host (Interviewer)
I haven't really plugged it in yet. I somewhat scared to look at the answer. Like that's honestly like the same reason for health and finances. Like I'm like, no, no, it's really good.
Akshay (OpenAI Core Product Engineering Lead)
I mean it's really cool how. I mean we were talking about the agentic search aspect a little bit earlier, but it's really cool how in conventional ux, the more power you want to give to a user, the more knobs and bells and whistles you need to add. For these finance and budgeting apps there's always a bunch of the different filters and search bars and stuff like that. But now with the right connectivity to the right data, you can have whatever you want. You can ask any question you want into that box and get the answer. And I think that's super powerful.
Co-host (Vibu)
I think it's also nice to just have it centralized in one space, right? You have different health apps. I have one for a smart scale, a watch, all these different things. It's just nice to centrally co locate
Host (Interviewer)
it, which is, you know, part of the whole thing of Openclaw. Right? Like that you would have personal OS, which presumably ChatGPT wants to become. I do think that just relying on like just in time pulling of data for let's say via mcp, cli, API, whatever you, whatever you do. Still not enough. I come from a bit of a data engineering background. You still want a data warehouse or some kind of caching or semantic layer. Do you feel that or do you already have that?
Akshay (OpenAI Core Product Engineering Lead)
I can't speak to all the details on how everything works, but I think it depends on the access pattern. If you want an answer immediately, then yes, it's very difficult to do that. You need to pull from all of these sources. But a lot of the use cases that we want to enable in ChatGPT work aren't necessarily something that you need immediately. It's more like a task that you want the agent to go and do and that's going to take a certain amount of time. And with things like programmatic tool calling and stuff now Some of that time and sub agents and stuff, some of that is also parallelizable. And so it's possible, I think it's very possible that the ceiling on what can be done with MCPS and calling out to these third party services has been raised substantially. So we're really excited about that.
Host (Interviewer)
You mentioned sub agents. I got a double click on that. Ultra is a new mode. You have special affordances in ChatGPT itself to show off the agents. Can't really do much with them, to be honest. Just, just watch. What if. What have been your experiences? Any design issues that you would call out to other builders building with sub agents?
Akshay (OpenAI Core Product Engineering Lead)
I think it sort of goes back to the balance that I was raising earlier about like, you know, showing builders the power of the tool, but also creating enough of an abstraction to not overwhelm them. I think with some agents, the thing that we wanted to show is that you can take a task that, you know, has many parallel tracks or is complicated in a way that, you know, sub agents can handle. And this product is for you. Like the model can, can accomplish those goals or try to accomplish those goals. Um, and so like, that's the point of like showing them in the product. And, and that's where we, we've gone with the design. There's another, you know, iteration of this where like, you can see exactly what they're doing and, and things like that, which I think is like, you know, could, could verge on like overwhelming with information. And so this is like the deliberate trade off that we made for now.
Host (Interviewer)
I mean, you, you do display quite a lot of transcripts.
Akshay (OpenAI Core Product Engineering Lead)
Right?
Co-host (Vibu)
Right. I think it's.
Host (Interviewer)
You can display more than that.
Akshay (OpenAI Core Product Engineering Lead)
No, no, it's.
Co-host (Vibu)
Some, some people could want more. So I'm one of those people that will basically throw a lot of stuff at goal and pretty much every goal I'll tell it to use subagents seems redundant.
Akshay (OpenAI Core Product Engineering Lead)
Right.
Co-host (Vibu)
But every time I'm like, okay, use subagents where possible. And I have a lot of people, a lot of friends that recommend and do the same. Whereas I'll sometimes talk to people that are like, okay, this is where I want you to use subagents for this subtask. And I'm sure they would appreciate seeing into how they're being used. For me, it's primarily like two things, right. One is net time efficiency, so span out across subagents. Two is probably cost. Right. Don't use big expensive model offload to a lot of smaller, cheaper models. And some people want that level of control. So if you have repetition in what you're doing. Right. Say I want something built where I want it to consistently do this every day. I might want to go in and fine tune subagents here, subagents there. So you can see both. But I think if I'm not mistaken, it's hidden by default. There's a dropdown that goes a lot where I'm like, okay, I'm just going to keep.
Host (Interviewer)
You can change the model that they use.
Co-host (Vibu)
I know I tell them to be steered. I'll say my. I know Anthropic offers this in cloud code. You can tell Table to use Sonnet or Opus to use Sonnet as subagents. So pretty trivial thing. You tell it to span out subagents. With Sonnet, it's cheaper, faster, I would assume if it's not there, it could be built there. But I think there's a side of too many toggles. It's not a toggle, actually. It's just utility chat. The way I do it is prompt it. Right. And I think this is something that gets abstracted unless it's something you built for repetition. Right. So if I'm building something, say that's podcast prep. Right. Research into people do a very, very deep, extensive research that I might want to configure to cheaper, faster model just for web search. Right. I can see a world in which you want both. I think the default is actually pretty good right now where it's hidden. But you can drop down and get some more info into what's done. I know people talked a lot about it on 5.6's launch. This thing loves to use a lot of sub agents and causes the ChatGPT app to just crash because it's so processor heavy.
Host (Interviewer)
But for what it's worth, that's not my experience. Yeah, I mean, you know, I haven't had a crash from Saudi.
Co-host (Vibu)
I haven't either. We both have big laptops. I know people brought it up. There was a topic of discussion that we didn't see the same, but it is another vibe eval. Right. People are like, okay, the amount of sub agents Sol this morning is crazy. And I'm like, I think this is okay. I think it's good. But just stuff people bring up.
Akshay (OpenAI Core Product Engineering Lead)
I think when we launched the product too, we weren't as opinionated about who is Ultra for and when should they be using it. And since then we made some changes to require you to turn it on and find it in the advanced setting, because that's who it is for. It's for power users who understand what's going to happen because it also, depending on your use case, can use more of your limits as well.
Host (Interviewer)
Yes.
Akshay (OpenAI Core Product Engineering Lead)
So that's where I think a lot of the feedback was coming from.
Co-host (Vibu)
It's okay. Reset the limits. Always reset the limits.
Host (Interviewer)
Today we're resetting. Because of this, I want to change topics to one last piece of the harness is memory. A lot of people are commenting on memory recently. ChatGPT's new memory system used to suck. It's not very good. And then this guy also basically the same thing. And Samir, who you presumably work with, talking about memory. What can you say there?
Akshay (OpenAI Core Product Engineering Lead)
I think that, you know, Sameer and the team have made a ton of and the research teams have made a ton of updates and improvements over time. I think when I talk to friends, family members about what they love about ChatGPT, like the fact that it knows them, they feel like their ChatGPT is, is their ChatGPT, I think comes up probably number one. And ChatGPT work in the cloud, like by default, all conversations inherit from your ChatGPT memory. So you'll know, they'll know context about you and they'll also be able to
Host (Interviewer)
write back to this memory with a small text write, you tell me when you're writing. Right? Is it.
Akshay (OpenAI Core Product Engineering Lead)
No, it's part of the same memory V3 system that we launched.
Host (Interviewer)
Yeah. Dreaming V3. Yeah.
Akshay (OpenAI Core Product Engineering Lead)
So I think that's been really powerful because going from ChatGPT to ChatGPT work feels like an extension of what I've already been doing with the product for sometimes many years. So that's been awesome and it's awesome to see that people are recognizing the improvements here.
Host (Interviewer)
So it's basically a retrieval problem, right? Like, are you retrieving the right things? Are you over focusing on the wrong things? Is there more false positive or false negative? If that makes sense. What's the bigger problem?
Akshay (OpenAI Core Product Engineering Lead)
So I don't work on memory directly, so it's hard to say what the bigger problem is with certainty. But I think you're right. I think that there's two sides of it. It's making sure it knows things about you, but then also having the EQ to bring those things up at the right moments proactively or surprising you in ways that are positive, not negative. So I think it's a very challenging problem, but something that I think we feel very is a huge opportunity to get right, which is why we've made big investments in it.
Co-host (Vibu)
How do you see the side of okay, when you're building ChatGPT for work different than the regular chat app, different than Codex Manag memory across different projects, collaboration and whatnot. How do you see the side of what's separate from the harness? Right. So if I have four threads on one project, any learnings on how to build memory systems there, you know, for background as well, I guess to steer it a bit is when you do chat style applications I'd say you have a lot of one offs right. When you switch to work it might be something you're doing for a month, something you do a lot. Right now as I add more sessions there's a lot more than just single threaded, right. And there, there might be memory there.
Akshay (OpenAI Core Product Engineering Lead)
I mean I think first I challenge that like the depth of the memory or the value of it is like fundamentally different across chat and work. Like it is true that like you know there are a lot of like shorter sessions on chat, but I think you know the ChatGPT, the product has had like a ton of longevity you know as long as this, this technology has been around and, and people use it for worky, productivity related things already today. And so I think we found that there's a lot of value. I mean I found this with my personal usage like all these one offs add up over time into something like quite durable and quite a good representation of who I am. I know from time to time something will go viral on X about ChatGPT telling you everything it knows about you and people are always surprised how deep that is.
Co-host (Vibu)
The fun roast me exactly.
Akshay (OpenAI Core Product Engineering Lead)
So I think the. That's all to say that I think there's a lot of depth there in the existing ChatGPT product and so that's why I think we think it's valuable to bring into the work product. But the other reason I brought that up is because I think hopefully we can use some of the same fundamental primitives and systems to extend memory here as well. And I know this is something that the team that focuses on this is working through right now.
Host (Interviewer)
I wanted to bring up one element of memory which I honestly don't really use much and I'm curious if you do chronicle which is up on screen right now. It's kind of a super memory or what is it?
Akshay (OpenAI Core Product Engineering Lead)
I think the idea is that it can learn from how you're using your computer and it's another input source into memory and I think it's experimental right now and something that isn't default off but I'd recommend that you try. I think that it's Quite interesting how it goes back to a conversation we were having earlier on. You know, you were asking like, does it. Can ChatGPT miss things? Like does it, you know, on Slack when it's searching, does it miss things because there's such a volume of stuff. Right. And like it's. You can ask the same question about like everything that you're doing on your computer. Like is it going to. Everything that you're doing is going to capture the intent and stuff like that. Probably not. But like it probably will find things that you might not know about and then if it can surface those you in relevant times in proactive ways, like when you're doing tasks. And I found at least that it can be quite helpful. So it's worth trying.
Host (Interviewer)
So mostly for insights and longer term.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, exactly. Like insights and it builds context that can make you more productive on certain tasks. But it's hard to describe without feeling it.
Co-host (Vibu)
I will say you can feel it pretty well the idea of what they're saying here. Right. Just check through my memories or check through my logs and add skills. Pretty underrated.
Host (Interviewer)
Right. But that's automations. You can repeat that using a cron
Akshay (OpenAI Core Product Engineering Lead)
job, checking through your memories and creating skills. But I think the creation of the memories from Chronicle itself is like what's different. It's like you have much deeper memories because you have Chronicle on it's there.
Host (Interviewer)
I don't use it much but maybe I just. I need more examples. I imagine you guys use a lot of it internally, so I'm always fishing for use cases. Yeah.
Akshay (OpenAI Core Product Engineering Lead)
I would just try turning it on and then like it just auto works. Yeah. And seeing where it might start helping you. I think you'd be surprised. Yeah.
Host (Interviewer)
Amazing. I think that was about it. In terms of the Overall coverage of ChatGPT work. I think there's been a lot of good progress and discussion on building and all these things. There's a lot of ex founders in the community in OpenAI as well. Do you think that things have changed a lot? I guess your overall reflection of building pre AI and post AI, I mean
Akshay (OpenAI Core Product Engineering Lead)
I think things have changed a ton. I think it's super exciting to see how quickly you can go from idea to something real today. Whereas even before I think five, ten years ago it was fast if you were scrappy and rolling to build the minimal viable thing. But now the extent of what you can build is much broader. I think that also like what we've seen internally building is like that gives you an opportunity to validate much more quickly to Talk to users, to talk to internal documenters, et cetera, and make sure you're on the right track. And like that loop, I think has become more closed than ever before. And that's like a win for product development. I think it's a win for consumers and users too, because ideally that means they're getting much better products out the gate.
Host (Interviewer)
Does it mean your team is a smaller.
Akshay (OpenAI Core Product Engineering Lead)
I think there's much more to do now. So I think people can accomplish more individually or in a small team than that would require more people than before. But there's at the same time there's also more to do. So I think the teams are much more ambitious.
Co-host (Vibu)
Have you seen any changes in scopes of roles and building teams and how we used to have teams, say a few years ago, versus what ideal teams look like now?
Akshay (OpenAI Core Product Engineering Lead)
I think we've seen a blurring in the lines between the typical product development functions, between em, vm, engineer, designer, et cetera.
Host (Interviewer)
Yeah, I want to bring up this quote. There will be only four jobs left in tech. There's AI slop cannon, people who just burn a bunch of tokens, and then there is SRE people who are more responsible. There's grownups who sell things, and then there's hot people.
Akshay (OpenAI Core Product Engineering Lead)
This is an interesting take. I think my suspicion is that there's everything. Everyone will be T shaped in a way and that AI will enable everyone to become a generalist. I never would be able to come up with a design before and even now I don't have maybe the visual taste required, but I can iterate on something with the help of AI. But then people will have a specialty and that's the straight line and the T or the upper line in the table. And so you can have a specialty that you're interested in with the help of AI, you can go deeper and become better at over time. But then you'll also be a generalist. And so with that foundation, what you can accomplish is almost limitless.
Host (Interviewer)
What are you bottlenecked by in terms of specialties? Do you need more designers? Do you need more slop cannons? Do you need more hot people?
Akshay (OpenAI Core Product Engineering Lead)
I think the bottleneck becomes sort of like ideas and tastes, I guess. I think because anyone can build now, I think it really is the era of bottoms up ambition. And because there's so much to be built, you're always going to be bottlenecked by the amount of ideas and amount of things that you're doing at any given time.
Co-host (Vibu)
What do you think models help solve that?
Akshay (OpenAI Core Product Engineering Lead)
Models?
Co-host (Vibu)
Yeah. I mean I have the example of like, I have a front end design skill that's like, they give me four drastically different examples of what this looks like. Sure, it burns a lot of tokens, but, you know, and then I'll mostly just condense down. Okay, I like this part, I like this part. Let's draw these together. And it's like, yeah, I had a vision, but like, I don't know.
Host (Interviewer)
I would say that the one automation that I would love to work and it doesn't work is bring me new ideas.
Akshay (OpenAI Core Product Engineering Lead)
Right.
Host (Interviewer)
Somehow, LLMs are just not it.
Akshay (OpenAI Core Product Engineering Lead)
One interesting part about ideas is like, they're not like in a vacuum. It's like not. They usually come from somewhere and like, you know, in product development, like, they're coming from talking to users or reacting to friction that you're seeing or feedback, building on some foundation that you already had planned out before, whatever. And so I think that's where there will always be value in these generalists that we talked about closing that loop and coming up with those ideas that are grounded in that feedback or talking to users, whatever it is.
Co-host (Vibu)
Cool.
Host (Interviewer)
You lead the productivity team. How do you define productivity?
Akshay (OpenAI Core Product Engineering Lead)
I think our mission is to make it possible for people to do things that they weren't able to do before. And right now we're thinking about it from the perspective of knowledge work. And so when I look at knowledge work, I think about people are no longer siloed by their roles. They're no longer siloed by maybe the background or training that they have. No matter what function you're in, you can suddenly build things, you can suddenly get access to data that you otherwise might not be able to interpret, et cetera. And then I think that extends to your personal life where we want to give you leverage. At the end of the day, we want the models and the product to be able to give you leverage so that you can, you know, create time for yourself to do the things that you love.
Host (Interviewer)
Does that also translate to a way to measure productivity? Like, what is how you measure leverage?
Akshay (OpenAI Core Product Engineering Lead)
I think we haven't figured this out yet. Part of the reason is it's so diverse. Everyone has different goals and really the true measurement is like their ability to achieve that goal. Did we help you or did we not? And it's very difficult without knowing what that goal is up front and also tailoring it for every individual.
Host (Interviewer)
And the thumbs up and thumbs down from ChatGPT doesn't give you anything, right?
Akshay (OpenAI Core Product Engineering Lead)
Right. I mean, you don't know if they're thumbs downing the content of the answer, the vibe of it, whether or not it helped them with their goal, I think that's difficult. But it's something that I think we will need to figure out and the industry at large will need to figure out because that's how we measure success. If this is what we're.
Co-host (Vibu)
Do you think it's changed productivity and how you measure it? Basically you said there's a more work that can be done, a lot more scope. Has it changed?
Akshay (OpenAI Core Product Engineering Lead)
I think it was always true that what you really wanted to measure is like, you know, was your team, was the individual, were you personally able to hit the goal or are you closer to hitting whatever your goal is? Right, But I think previously we used proxies for this. So like you know, code commits, lines of code. Lines of code or whatever.
Host (Interviewer)
Story points.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, exactly, story points.
Host (Interviewer)
And like they're coming back by the way.
Akshay (OpenAI Core Product Engineering Lead)
Maybe. But that is sort of part of the change. And I think with AI now those proxies starting to fall apart, the number of tokens you use or the number of pull requests you make are no longer maybe as hyper correlated. Was your team able to hit the goal or are they on track to hit their goals? So I think we'll need to come up with new measurements for the managers
Host (Interviewer)
listening, give them one thing to try.
Akshay (OpenAI Core Product Engineering Lead)
I think for me what's important is like at bats, are we as a team building the muscle to have not just quantity of at bats, but quality? Are we able to go all the way from generating an idea, building it out, getting the feedback, reacting to that feedback, actually validating or invalidating the hypothesis, going on to the next idea? Are we able to do that really efficiently that goes to the actual code that's being written or the designs that are being made, or the specs that are being written, whatever. But also the culture of the team. Do we have the humility and are able to go through that process many, many times and stay motivated and excited throughout that? So that's the thing that I think is important now, especially when we're on the frontier of this technology and there's so much to build, there's so much to do. That's probably the most important thing that
Co-host (Vibu)
we look at any traps people fall into around measuring productivity, what your team work on. I feel like there's a lot of, okay, we added a lot of LLMs, we have dashboards for this and that, but not much has changed. Right.
Host (Interviewer)
That is the trap. Yes.
Co-host (Vibu)
And you know, the broader source of the question is for the managers and teams building, you know, how should they approach this?
Akshay (OpenAI Core Product Engineering Lead)
I think maybe the trap is like conflating motion and progress. I think motion is much easier now than ever before because of the tooling that we have. But progress requires you to be very prescriptive and deliberate about what you're actually trying to achieve. And it goes back to our question of measurement we were talking about. Can we OpenAI figure out how to measure productivity for our users? That's a very hard problem because of the diversity. But as a team, you should have a really prescriptive and deliver view on what progress looks like for you and for your team. And if you don't have that, then it's very easy to conflate these two things.
Host (Interviewer)
I think at batch is a really great thing. I'm really glad. I like the discussion between motion and progress. I think that's a quote that we're going to feature on the write up. You've been very generous with your time. Thank you so much. And congrats on 10 million.
Akshay (OpenAI Core Product Engineering Lead)
Yeah, thank you for having me.
Host (Interviewer)
Next one. I had 100 in two months, two weeks.
Akshay (OpenAI Core Product Engineering Lead)
Thank you, Sam.
Date: July 28, 2026
Host(s): Latent.Space (with co-host Vibu)
Guest: Akshay Nathan (Core Product Engineering Lead, OpenAI)
This episode dives into the evolution of OpenAI’s product engineering, focusing on the launch and rapid adoption of "ChatGPT Work"—OpenAI's new super app for productivity that merges ChatGPT, Codex, and agentic workflows for individuals and knowledge workers alike. Akshay Nathan details the product journey from early code tools to a generalized AI productivity platform, sharing insights into design choices, user adoption patterns, product philosophy, technical underpinnings, and future directions. The conversation is full of practical and strategic wisdom for AI engineers, product builders, and anyone curious how AI is transforming work and personal productivity at massive scale.
LLMs as the Final No-Code Platform (00:08–01:33):
Akshay traces his product journey from fintech and low/no-code startups (like Walrus and Airtable) to OpenAI. He notes that LLMs are the missing piece to bring programming power to everyone via natural language, which led to ChatGPT Work's vision:
Startup Culture at Scale (01:44–02:39):
Akshay notes OpenAI still feels like an early-stage startup culturally, with a "bottoms up ambition" and ability for anyone to ship ideas, even as company headcount and scope have exploded.
No One-Size-Fits-All in Enterprise (02:52–03:56):
Early ChatGPT Enterprise deployments revealed every enterprise has wildly different needs and ideas for AI, making productization hard:
Adoption vs. Understanding (04:39–05:27):
General excitement and adoption of AI is rising, but only a small fraction of users truly exploit advanced AI capabilities like agents. Unlocking broader adoption is a key product goal.
Bringing Power Beyond Developers (05:50–07:16):
Internal stats revealed surprising non-developer adoption of Codex, fueling the decision to merge agentic capabilities into ChatGPT Work and make the experience seamless for all:
Avoiding “Product Pigeonholes” (07:35–09:26):
ChatGPT Work is positioned for all forms of productivity, including personal, not just “work.” The UX is designed so users never need to pick between ChatGPT, Codex, or Work— the system routes and adapts as needed.
Unified Harness, Tailored UX (09:49–10:55):
The harness (core AI backend) is shared between Codex and Work, with tailored UX affordances—e.g., showing git diffs in Codex, hiding technical details in Work, UX-based routing decisions.
Artifacts and Collaboration (21:01–24:22):
Artifacts (dynamic, shareable outputs including code, sites, spreadsheets) are a major push in Work, enabling agent-generated Excel, sites, etc. Akshay notes the value of “work as collaborative artifact”—as teams are increasingly replacing slide decks and spreadsheets with sites/artifacts for richer knowledge work.
Advice: Defaults are Powerful (16:59–18:29):
While there are many model/difficulty settings (SOL, Terra, Ultra, etc.), Akshay recommends most users stick to default settings, which are optimized for general performance. Power users can experiment as needed.
Slider UI for Model Gradations (19:30–20:04):
A new UI “slider” abstracts numerous technical model options into a simple control balancing speed/efficiency vs. quality/thoroughness.
Speed of Iteration and Closing Feedback Loops (60:45–61:32):
Product cycles are compressed: "Now, you can go from idea to something real much faster—feedback loops are tighter, and product validation is almost instant."
Team Structure & Scope (61:33–62:08):
Product, engineering, design, and management roles are more blurred than ever—everyone is T-shaped, enabled to be a generalist via AI.
From Siloed Roles to Leverage (64:46–65:25):
The team's mission: let everyone do things they couldn't before, removing silos and increasing personal and team leverage.
Pitfalls: Mistaking Motion for Progress (68:18–68:57):
Akshay warns not to confuse new AI-generated “motion” with actual progress. Teams must be highly deliberate defining actionable progress metrics, not just surface usage data or LLM outputs.
"LLMs became the missing technology required to bring the magic of code to everyone without them having to know what's going on underneath the hood."
(00:32–01:33, Akshay)
"You have this box and you can say anything to it, which is the magic. But on the flip side, it also means you don't know what to do with it."
(02:52–03:56, Akshay)
"We want the model and the product to be able to give you leverage so you can create time for yourself to do the things you love."
(64:46–65:25, Akshay)
"Motion is much easier now than ever before because of the tooling that we have. But progress requires you to be very prescriptive and deliberate about what you're actually trying to achieve."
(68:18, Akshay)
"Our hypothesis is that part of what we're building in this technology is giving people leverage—that might actually blur the lines between someone who's writing code or planning events."
(12:14–13:44, Akshay)
"The show not tell is a big piece that ... we're not fully cracked—showing people all they can do with the product, not just onboarding or docs."
(33:02, Akshay)
| Timestamp | Topic / Segment | |-------------|--------------------------------------------------------------| | 00:08–01:33 | Akshay’s career, LLMs as the ultimate no-code | | 02:52–03:56 | Lessons from enterprise users & heterogeneous needs | | 05:50–07:16 | Discovery: non-devs using Codex leads to “merging” products | | 09:49–10:55 | Technical: Harness unification and UX distinctions | | 16:59–18:29 | Guidance for choosing models, defaults vs advanced configs | | 21:01–24:22 | The rise of “artifacts” and the new collaborative workflow | | 34:27–35:51 | Vision: from developers to all knowledge workers to everyone | | 50:23–54:33 | Subagents, parallel execution, and Ultra Mode | | 54:39–59:24 | Memory systems, cross-product personalization | | 60:45–61:32 | Building & iterating much faster in post-AI world | | 68:18–68:57 | Motion vs progress—core productivity insight |
The episode is conversational and open, reflecting internal product thinking and candid stories—ranging from technical details to personal anecdotes (like using OpenClaw for managing home calendars or ChatGPT for meal planning). Akshay and the hosts blend engineering detail with reflection on how AI is reshaping both the tools we use and the way we design, diagnose, and iterate on those tools.
For further details, visuals, and links, check the show notes at latent.space.