
Loading summary
Kieran
Today we're going to show you how you're using ChatGPT01 completely wrong. You're treating it like a dumb model and it is so much more than that. We're going to show you how it can save you hundreds of hours a month and literally 5 extra output. We're joined by Sully Omar, who's the founder of Autogrid, which is an AI first spreadsheet startup. And he is giving us a masterclass in how to use O1 to get remarkable output for your business. Let's get into today's episode. Hey guys, real quick. You know we love building custom GPTs on the show and we love sharing it with all of you. Well, we wanted to kick that up a notch. We just developed this free guide that teaches you how to build your own custom GPT on chatgpt. We've taken the guesswork out of it. We've got templates, we've got a step by step guide to design and implement custom models. So you can focus on the part that's actually fun, the part we love actually building it. And if you want it, you, you can grab a link in the description below and go check it out now. Now back to today's show. Sully, thank you so much for joining us.
Kip
Yeah, the tool is really cool just to give Autogrid its proper intro. Like it has a bunch of agents to basically make research really easy, right? Yeah, it's like a smart spreadsheet.
Sully Omar
Yeah, exactly. It's like every single cell is its own little agent and you give it a little task and it'll go out and do that for you and give you this nice little output. Actually a lot of people, interestingly enough, use it to enrich their contacts and update their leads in HubSpot. So there's a lot of people using it or wanting to be like, hey, you know, we want to sync it with our HubSpot. So when I add a lead to HubSpot, it'll enrich the contact and then add it back in.
Kip
So yeah, very cool.
Kieran
Agents for data enrichment, man, it's a thing.
Sully Omar
Yeah, it's a big use. Case agents for data enrichment and research. There's a lot of these relatively straightforward tasks that are pretty annoying to do and not too difficult, is like the sweet spot right now.
Kip
Yeah. One of the things I think that you've done, which is what I think people should think about in terms of agents to wrap their head around, is so when you think about your product, you actually have like these micro agents that complete these tasks autonomously, but they're not like doing all of your customer support or doing all of your sales. You have an agent per column to fill out a piece of data. Which I think is a really good example of where I do see autonomous agents being able to do work in 2025. I think when people think about autonomous agents, they're being told it's going to do all of your sales, do all of your customer support, do all of your marketing. And I don't agree with that. I think you're going to have agents with human in the loop doing that. But I think your product is a good example where you do have these micro agents. And like an agent per column is like a really good example where it owns like one piece of data. And so it's a good example, I think, of like where I do think you actually do do really see autonomous agents own the entirety of like that task.
Sully Omar
And the reason we actually did that was because we previously had tried to build these autonomous agents and the success rate was like. The example I like to give is imagine every time you open slack and like 25% of the time it just didn't work. Right, Right. Like you would not use Slack. And for agents it's like there's that threshold of like, okay, if it doesn't work every like, you know, third iteration, it's basically useless.
Kip
People give up.
Sully Omar
Yeah. And that's why we took that approach of like, why don't we just give an agent a very, very specific micro task that there's like almost a zero chance it messes up and then you can kind of take those and kind of get them to work together. So you have a hundred of those or whatever, a thousand of those and it becomes a lot easier to do these sort of like long multi step tasks where each one, like you said, it's like a micro agent focusing on one little thing and then you build a product around that that lets you sort of have them work together to like solve a more difficult, broader task.
Kip
Yeah. Sam Altman talked about this. He said one of the weak points of LLM models is if you ask it to do something 10,000 times, it will likely get a lot of those things wrong. Like the reliability is just not there. And so if an agent is going to replace an employee, it's not a great experience. If you tell that agent to do five tasks that day and then it comes back and says, oh, I did these two and then these other three, I just decided to do something completely random instead. You probably stack that employee.
Sully Omar
That's like the one thing that I think some people when they get excited about Tonos agents, they forget that like, okay, maybe three, four iterations work. Yeah, like you said, like if like the 7th, 8th and 9 somewhat work and then maybe the 10th does work, like you would not hire that employee. You would put them on a performance plan and you'd be like, hey, you gotta fix this. And I think the autonomous agents, they need smarter models is what I've kind of playing with them. And I think the reasoning models that we saw from, you know, OpenAI was a good step forward. And I think we're going to continually see that trend is like the thinking models get better, they'll think for longer and they'll be able to do these sort of more difficult tasks with a higher success rate. But I still think we're a couple of generations of models away from that. But I think even now though, like, there's the things that you can do with something like O1 Pro compared to even six months ago is just like mind blowing. Like the things that I've been able to do at my company, it just like to me feels like, used properly, like a genuine senior engineer. Like I use it to code a lot. I use it to look at my KPIs, my metrics. And I think the one thing I've noticed is that it's used very differently than how you would normally use, you know, like the regular model, like ChatGPT or Claude. And then I started to notice that a lot of people prompt it in a very sort of suboptimal way where they're expecting like, hey, I'm going to ask it a quick question and a quick response, which is not the right way to use it. The right way to use it is sort of like what I like to call is like building context. So grabbing all of the piece of information, sitting there and thinking, like, what am I going to ask this model to do for me? Right? Like as if it's like a person. Right. You're not going to just tell a person, hey. And it's not going to say hey, right? You're going to sit there, you're going to probably be on maybe like a zoom call and you're going to have a very long conversation. And I really like to see these like thinking models. Is that as well?
Kieran
Yeah, can we go into that? Because I think there was a lot of hype around OpenAI's O1 model and especially in engineering and science and academics. And I think the average person was like, oh, I guess this is cool. But it's for stuff that I don't do and I don't kind of know how to do it. I'm just over here doing like my basic stuff in ChatGPT. Could you maybe like demystify this? Give us some prompting secrets like how should the average person, if you're just like hey, I work in marketing or I work in sales, what should I be doing with O1 and how should I be interacting with a reasoning based model versus a maybe non reasoning traditional base model?
Sully Omar
Great question. So what I've done is I've had people ask me this and I came up and it's just my own thing. I called it a three tiered system to like thinking about models. So what I did is I actually created what I like to call internally and we referenced this at our company as well is it's like the three tiered system to using all these models. So there's like, I don't know, God knows how many models over 50 now. But I boiled it down to these core, core ones which is Claude, Gemini and OpenAI ChatGPT. So the system is pretty simple. You have three tiers, you have your working models which are the ones that are really cheap, really fast. So you have things like Gemini Flash, GPT4O mini haiku. These are sort of dumb. They're not like the smartest models and I wouldn't ask them to do a very difficult task. But if I'm like, hey, I have a bunch of pages of context or something I need to get summarized or maybe I'm looking at a transcript or something where I'm like, okay, look, this transcript's like four hours long and I don't want to pay a large amount of money to send it to 01. I like to refer to these as the first tier. Then you have your second tier which is the more intelligent day to day usage model. So if you're using ChatGPT, if you're using Claude, you're using Gemini, the actual products use this sort of middle tier of models. And now this is the one that you're probably going to converse with, you're going to ask it all sorts of questions and the majority of people are going to be in this tier. So you're using Claude, using GPT4 or Gemini, then we have the reasoning models and sort of, you know, these came out I think you know, last month, maybe like six to eight weeks ago. And right now there's only two companies that have them. So you have ChatGPT and they have 0101, Mini 01 Pro. And then you have Google with Gemini thinking. So I split it into these three, because the way that you use them is slightly differently. So I'm not going to dive too much into, like, how you would use this first and second tier because I think most people already, you know, use it. You ask it to help you write an email, summarize things. I wanted to crow in a little bit more into, like, how I craft and generate prompts with these reasoning models, because I do think there is a little bit of skill. And I say this very lightly. It's not like, you know, you need to have a PhD or anything. It's just more about knowing the right thing to say and how to phrase it to these models, especially 01 Pro, where it can just do things that every day I'll be like, I can't believe there's this thing that I can pay, you know, a company for so much a month. And it's like, literally like having like this super, super intelligent, super smart coworker.
Kip
I love this. This is really cool. So for people following along on RSS, you should definitely go to YouTube to check this out. Can I just make sure I'm following? I've used all these models and the Gemini thinking. Is that Gemini 2.0?
Sully Omar
Yes.
Kip
Is that the name? Yeah, yeah, yeah, yeah.
Sully Omar
Honestly, Google has the most confusing naming for all their models.
Kip
But yeah, the pricing and packaging of all these companies. So confusing is like beyond a puzzle. Like, it's really bad.
Sully Omar
Yeah. I think OpenAI, they skipped O2 and they just went straight to O3 and like Gemini and then Claude with like their 3.5.6 V2. I don't know who on their teams there are naming these models. Really confusing. But yeah, Gemini 2.0 thinking, which I think that was the most recent sort of like thinking model that came out.
Kip
Yeah, that came out with their bundle where they had deep research, which is an incredible product. They had Flash, they had the stream Real time, which I think was the first time Google stole OpenAI's thunder because they had the 12 days of shipments.
Sully Omar
Yeah, I think up until that point, everyone had been sort of saying, you know, OpenAI is kind of the company to beat. Was like real time. And then I think Gemini's real time did a really good job.
Kip
So good.
Sully Omar
I think that's most. Because Gemini is natively multimodal input and outputs. For those of you who don't know, it means it can take in video, it can take in Images, but can also natively generate these images and these videos, as opposed to other models where it'll generate text and then they'll have another AI model take that text and convert it. So Gemini is native, meaning it can output text, it can output video. So that's why it's really good at, like, doing these sorts of, like, use cases.
Kip
Cool. So I would love to learn how you prompt then, for the third category of LLMs, like, the reasoning models, because Kip said, like, what is, like, fundamentally different about using those models?
Sully Omar
So the main difference for using ChatGPT versus 01 Pro is that these models have to think, and sometimes they have to think. I think my record is I got one to think for like 10 minutes. So you can imagine if you're waiting 10 minutes between prompts, you don't have that chance to sort of have this back and forth. Right. You could do it, but that's really not what the models are built for. Then you'll actually start to notice if you have a chat with Zero1Pro, it actually will do worse than if you were to just sort of throw all of the text at it. And within prompt engineering, this is called sort of like one shotting a prompt, which means you just sit there. And I've noticed that I do this myself, is I'll sit there for five to 10 minutes thinking of the prompt to ask, and I'll actually go and I'll consult Claude or I'll consult ChatGPT. I'll have a conversation with it and then I'll say, okay, given this whole chat, please create a prompt that I can give to O1.
Kip
This is exactly what I do. I'm glad you brought this up. Yes, this is an incredible tip for listeners.
Sully Omar
Exactly. And I think one thing I will say, and I'm guilty of this, is I'm actually not good at writing prompts. And it's come to a point where I will just have a conversation with voice mode on ChatGPT, or I'll use Claude and we'll just have a chat. And I'll say, okay, I am no longer writing the prompts. And this is actually something that is going to become more common. It's called meta prompting, where your requests or the things that you type into the app no longer get sent directly to the model itself. There'll be a layer, an intermediary layer, which will try to decipher what it is that you're asking and then create a sub prompt that will go to it. And I think this sort of thing that I'm doing right now manually and you guys are as well, is going to be done eventually. Like the companies will just do it themselves. And it's just because you can imagine in the future if a model takes an hour and it, you know, takes maybe a day to think, you probably want to get that right piece of text. But yeah, that's sort of what I do in terms of just getting started is I'll go into Claude and I'll say, hey, when I actually have an example that I can walk you guys through, which is one that I found pretty powerful. But it's, it's a super useful workflow for like maximizing. And it's a lot simpler too because you don't need to sit there and think about like, how do I articulate this prompt? So this is going to be. If you guys are cool with it, this will be just how I actually use O1. I'll walk you through the example and then I'll walk through how I use it and I have like some text that I pre prepared for the show. So we'll go through it and then let's just see how it goes. Cause I get. I haven't recorded this exact flow before, but this is the flow that.
Kip
Yeah, yeah, no, look, we can work on it in real time. That's fine.
Sully Omar
Okay, awesome. Okay, so here's an example of how I've been using O1. I know a lot of people like to use it for coding and engineering, which I do, but I think that's the obvious use case. Like, sure, just, you know, you give it a bunch of code. So what I've been doing is I've been deferring anything revolving my business around KPI's metrics, anything I need to do around my business where I'm like, hey, I need another decision maker. I go to 01 and I go to 01 Pro. But the way that I do it is I start with GPT4O. So I have got this scenario here that I have created and I'm going to paste it into chatgpt here. So I have a bunch of these metrics so we can look it through. So for example, what I would do is I would maybe grab these from Stripe or maybe you have like a analytics dashboard or something like that. And I'll come in here and I have all of these metrics. So here, this is obviously all made up. These are my monthly active users. My daily active users. And you can see there's a ton of metrics that I will pull from Some sort of reporting dashboard. So what I want to do, maybe for example, is let's say I have all of these sort of things I need to reduce churn. I need to increase all of these various metrics here. So now here's this giant metrics field, right? So here I have this. Now normally what you would do is maybe you would go into ChatGPT. You say like, hey, how can I optimize this? But the workflow I have is like a little bit different. So here are my metrics. And then I'll say, hey, I'm crafting a prompt to a smarter model than you. I want to optimize two to three key metrics for this quarter. Now craft a prompt I can send to this model that will give me a detailed report on how to do this. And I'll say, note, don't do this yourself.
Kieran
Yeah, 4o is getting slammed here.
Kip
You always have to.
Sully Omar
Yeah, I know, but you have to.
Kip
You actually have to do that for all AI assistants. Like if you ask it for something, you say, please don't do this part because it will just try to do the next part itself. Right?
Sully Omar
Yes, exactly. So here we'll go ahead. And this is sort of how I start all my problems.
Kip
Just to explain this to the listener, Sully, what you're doing is basically using chat functional 4.0 to basically craft a prompt for 01.
Sully Omar
Exactly. So I'm taking this. And for this example, I have maybe even slightly cleaner than usual numbers. What I'll do, actually, a lot of the times is I'll just literally copy the text on the page of my reporting dashboard and I'll throw it into 01 so it's even more unstructured. It's really messy. And I'll come into here and I'll be like, here's all of this just disgustingly raw, ugly data that are some metrics that I care about in my business.
Kieran
And you can even screenshot that too, right?
Sully Omar
Yes.
Kieran
And upload the image. There's a bunch for people watching, there's a bunch of different ways you could get that data in here. That's really easy for you.
Sully Omar
Exactly. And the key note here is that you're not throwing this into 01 right off the bat because you might have a dashboard and there's like three screens. So that would be three screenshots. The whole idea is to make this as easy as possible. So I would just dump all of the data that I care about into ChatGPT. So I was like, here, just take all of this you can upload files. But the main key thing here is we want to make sure that we have the right prompt with all the information to send to A1. So here, let's read what it says. Okay.
Kip
Prompt.
Sully Omar
I'm working on a SaaS business that's. These are the key metrics. Cool. Additional context revenue metrics. So right off the bat says a breakdown of the strategies, potential experiments, initiatives to implement. Like, I'll be honest with you, to craft this prompt really nicely and this articulate probably would have taken me like 20 minutes. Sometimes I don't even know what to ask. And that's the nice part is if I'm here and I'm just chatting with GPT, I can say, hey, can you improve the prompt like this? Can you improve it like that? And it gives you this ability to go back and forth without waiting.
Kip
Yeah.
Sully Omar
So I use O1 Pro mode.
Kip
So O1 Pro is the $200 a month?
Sully Omar
Yes. And some people ask me, what's the difference between 0101 Pro? To be honest, I don't know the very specific technical details, but from my personal experience, I've compared 01 and 01 Pro, and I found that the nuance and the ability to sort of decipher what it is that you're asking, especially if you have a large piece of text or documents, Zero1Pro is just a cut above the rest. Like, it genuinely does feel like a person at some point. I was just like, is this someone behind the scenes, like, doing this? And I would compare them, like side by side, and I would be the same prompt, same everything. And the nuance, I think, was what got it for me, where it would say, like, okay, you could do this, but potentially this could cause this, which means you should do this. Whereas O1 would just be like, just do this. Personally, I don't know about you guys. Whenever I'm trying to get an answer to a difficult question, I want pushback. I want to know what are the caveats. Pro Mode was the only model I've ever tried that did that.
Kip
It proactively pushed back without you having to say, find flaws in this and debate with me and why I'm wrong.
Sully Omar
Yes, it would give the nuance into what are the different things. And now I'll show you how I would do that. So you could prompt it, which is more likely to happen, or you can just let 01 think. So here's I have this prompt, right? It says if I were to break down to the structure, it's like, I'M working on a SaaS business. I got my metrics and I'm just going to paste this right off the bat just because to show you guys how easy it is. Okay, so goal is to support iq. Now it's going to go out and think and this is actually the boring part where it's going to think for a undisclosed amount of time. So we'll let it go ahead and do that.
Kip
We can talk about a few things here. So it's continuing to think. When you think about your usage here, in that you are using 01 Pro, I suspect you are similar to us. You are a power user of AI. You're using all of the different models for different things. Now, have you collapsed all of your usage into this single model? I find Claude just better for a range of things. Actually, it's interesting. We had Scott who runs the product team for Claude and I was trying to describe to him outside of it's a better writer, it's a better communicator. Why I love Claude. It's a little bit like why I like one of my friends versus another friend. I just like it better. Like I just like hanging out with Claude. Like, what's your usage look like? Is this your only model? Or how do you think about the different models?
Sully Omar
I think I'm on the same boat as you. So Claude is by far a better writer. She sounds a lot more human. I like to converse.
Kip
Yeah, me too.
Sully Omar
I just like to chat more with Claude. Yeah, so I'm on the same boat as you. My usage is very interesting actually. So in this example, I just wanted to keep it simple. I actually would have gone to Claude. I actually wouldn't have used 40 to craft that prompt for me. I would have used Claude. I'll use ChatGPT when I want to use web search just because it's just like, okay, cool, it's fast. And then I'll use it for like voice just because it's a little bit easier. But basically I've split it into like three categories. So the thinking and stuff goes to ChatGPT. All these like really difficult, analytical things. Right. So this is a pretty analytical thing. Here's all these metrics and all these like data points. Please analyze it for things like writing, for things like emails, for things like, you know, I want to chat with an AI about something. I'll go into Claude and I use cloud a lot as well because I like to still code a ton. I'll use it for coding and that's a whole separate workflow So I would say that it's still very much so like 5050 in how I use Claude and how I use ChatGPT. But I would say ChatGPT01Pro. Analytical things. I want to talk to someone with a conversation and just maybe enjoy my time. Going to go to Quad.
Kip
Nothing for Google.
Sully Omar
So for Google, the way that we use Google, I use it for multimodal things, but I don't really use multimodal that much. And at our company I will use the model for a lot of the tasks. I think when I was mentioning the three tiered system. So we use Gemini on it more than actually all the other models in terms of pure tokens because it's the cheapest and it's the smartest.
Kip
Biggest content winner. Yeah, yeah. Kip and I have been using Gemini Enterprise Seats. They are dope and they are dope because they're plugged into your email and your G Suite.
Kieran
Google Gemini with my G Suite wrote my performance review in like 30 seconds and it was great.
Kip
Yeah, it's dope.
Sully Omar
Yeah. Maybe that's the one part I haven't used it in, which is the G Suite integration, which I guess is really powerful now that I think about it. One time I was getting it to write an email for me. I like wrote an email and I was like trying to edit it and it like sent it and the email was like. I was like, it was not like the best written email. And I was like, oh no, it was like to someone important. I was like, no, it's flaky.
Kip
I will say it has some incredible G Suite, like one of the incredible use cases and we're probably going to do a little video on some of the use cases is you can basically say like aggregate all of the Google Docs that I've been tagged in, put all of the comments in different topics and then craft replies. Because I lose track of all the kind of comments that I'm getting tagged in G Suite and I can see the comments in docs and then it has access to your email. So you can say, well show me all of the emails this week that mentioned the same topics so I can collapse all my comms and aggregate them together. Okay, I see some pretty interesting use cases.
Sully Omar
That's good. That's actually an interesting one that I hadn't thought about because I've been using it on sort of like the Write and Help me, but you're using it sort of more on like I have like a bajillion different things that are coming in and I need that agent in front to like sort of be my first layer of defense.
Kip
A comms assistant. I think most knowledge workers need a comms assistant.
Sully Omar
I'm pretty sure I might just go after this and just go set it up on my end because I need that.
Kip
All right, your prompt is back.
Sully Omar
Okay. Yes. Let's see. It's back. Let's see. Oh man. See?
Kieran
Oh, geez. It's like a McKinsey report over here.
Sully Omar
I know. And I've had it think for like seven minutes once, which was like actually blew me away. But when that was like a record and it's funny, it thought for seven minutes and the answer was like four lines. But okay, I'm gonna read this and I don't want to read too much of it because like this is literally like a report. So here it says it gives me a below as a structured prioritized action plan. It's broken down by metrics, initiatives, and sort of all the things that I asked for it. Initially I was like, okay, so reducing the turn from 6.5 to 4.5 onboarding initiatives. Develop a more personalized guided onboarding experience that highlights the core product value. Immediately implement walk tour tours. Okay, it's not a bad idea. And then it actually gives me the KPIs. Okay, so provide personalized emails and notifications triggered by user behavior. I'm sure you guys know this is actually a really good way to sort of, you know.
Kieran
Very good.
Sully Omar
You're good. Okay. Seven day retention, onboarding, percentage of complete. It even tells me the KPIs that it's going to optimize for. So it's going to optimize for the seven day retention and it's to kind of reduce that first month churn. All right. And even gives me like where I need to go.
Kip
This is so good.
Kieran
So good.
Kip
Remember like when you used to have a search query like this and you used to have to like try to aggregate information from a hundred blue links and people think search isn't dead. Like basically your prompt is a search query.
Sully Omar
Right.
Kip
But you can just do more advanced search queries and it just gives you back an entire guide. Your cross functional stuff is pretty great.
Kieran
It's wild.
Kip
I'll let you call it out. But if you could just call that part out. So let's go. What it's actually put down for cross functional, the one I'm looking at is the fact that it's figured out customer success for training materials.
Sully Omar
And this was like an example that I kind of crafted before this, where it was like, okay, here's some metrics. It gets even more powerful when you give it more context. So in the context of my company or in the context of your team, the more context you throw, the better these models do. I have yet to run into a scenario where it didn't blow me away very few times. And I would look back and I'd say, oh, where did this model mess up? And I actually realized I asked it the wrong thing. I gave it the invalid metric or I gave it something that was conflicting and I would be reading it and I'm like, why did it say it? Is this model not smart? And I was like, oh, wait a minute. It's because I said one thing at the top of my prompt and then I said something that completely contradicted it and this actually noticed that I contradicted myself and it would say like, hey, you said this and then you said this. So I'm going to go ahead and take this guess because I couldn't really figure it out. So very rarely has it made mistakes. But yeah, this is a pretty long report because it's give me the customer success. I don't want to go through all this because this is like a very, very large report. But like increasing strategy initiatives, upsells and cross sell campaigns. That's actually really good idea to increase our average order KPIs.
Kip
That's a really good idea actually. Upselling cross sell campaigns. One of the things it's called out is to segment customers by usage patterns because heavy users are more likely to need advanced features. That is quite nuanced.
Kieran
You got to really know a lot about SaaS to understand that actually.
Kip
Yeah, like that's not a typical like average. Read the average like guide and get that. There's some nuance to it actually that.
Sully Omar
What I want to do to show you guys the difference, I'm going to throw this into Claude as we're continuing to read and we can compare it. So I'm going to paste this here directly into Claude and while that's running, we'll just continue to read through. And I wanted to show because I think you'll read this, you'll be like, all right, this is cool. Can't other AI apps do it? And then you'll read the response from another model and you'll say, okay, wait a minute, this actually does make a difference. So it's here risk mitigation. I just want to read the headings because I think there's okay. Increasing the conversion rate from 3 to 5% A B testing continuously. Experiments using Google Optimizely KPIs. Okay, see here, Reducing the friction. All right. These are actually all like pretty good SaaS. Things to know, right?
Kieran
Yep.
Kip
And they're all right, by the way. Someone who's worked a lot of growth. Yeah.
Kieran
These are all the things you would be doing.
Kip
These are actually all correct.
Kieran
If you had no clue how to do any of this stuff.
Kip
This is all right.
Kieran
You could just go and build a plan to do all of them.
Kip
I haven't seen a thing that it's got wrong yet, if I'm being honest.
Sully Omar
Let's see here. Personalized lead nurturing and retargeting drip campaigns. Okay.
Kieran
Yeah, that's also right.
Kip
It's got retargeting for abandoned signups.
Kieran
Kieran is like, oh my gosh.
Kip
I actually. This is pretty mind blowing.
Kieran
You're ready to move to Spain and retire.
Kip
I would say if I set this up as a use case. I've hired a lot of growth people. If I set this up as a use case, like a practical like working case as part of the interview process and I asked that person to come back with a plan, I suspect this is like way above average of what I would get back for sure. And I know like the additional thinking time is. And I might make myself look stupid because I was like doing a ton of digging, but I've read so much AI stuff it might just fall out my brain. But like the additional thinking time allows it to bring back a bunch of answers and then have additional time to like filter through for the correct one. Which is why it's getting you above average and better nuance. But it's the nuance actually when I look at the headlines like a landing page conversion rate, interactive product search and tutorials, accelerated onboarded and quick win checklists, like they're all things I think you could get from like best in class blog post. But it's the actual tactics in there, the how that it's got some real nuance and things that you would not normally get from just like the average. Which usually it will just return the average.
Sully Omar
Yeah. So let's just compare this. So I'm just going to compare the length of what I've got. Okay. I don't know if you guys can see how much I'm scrolling.
Kieran
It's a lot of scrolling.
Sully Omar
This is a lot of scrolling. We go to Claude and by the way, Claude is a great model.
Kieran
Great model.
Sully Omar
And it's done.
Kip
I love Claude so much.
Kieran
Wow.
Sully Omar
Like, and this is the same prompt. Like, this is not me changing the prompt. And I think it's when you compare the two, you'll start to like, say, like, okay, wait a minute, let's see. Churn reduction 6 point to 4%, whatever. With your strong enhanced onboarding. Okay, so let's say, which makes sense. Like, I think this is also what ChatGPT said, but let's actually just read the difference. It's just the granular detail that it gives me. It just gives me like very specific direction and like more nuance on what to do.
Kip
Plus it's giving you a structure that you hadn't asked for, which is basically giving you the tactic KPI and then cross team collaboration. Yeah, it's able to actually build upon what you asked for with things that you probably want that you didn't ask for. That's what I think is surprising.
Sully Omar
And like, just to showcase. So here, Claude gives me one bullet point. Like, okay, set up early warning systems right here. It just like so much more in depth. And I think it stops at like the second point.
Kieran
It's night and day.
Sully Omar
Here's the second point. And then it gives me the third point. It's night and day. And then like, I think a lot of people say, like, oh, you know, how would you use O1 versus, like, I hope to anyone that's watching this you can expense 01 Pro because it is expensive. I would say that if you're doing anything that revolves, like just heavy thinking and reasoning, and I know that the term gets thrown around, I like to define it as something that like, you can give to a person. I would defer everything to O1 Pro because like, again, Claude is still really good to chat about things and like, have that quick back and forth. Because here it thought for, you know, whatever, two minutes and something versus here is like right away. But the difference is just night and day. For me personally, and I notice this in anything that I use, whether it's coding, whether it's sort of like even writing different things, I will say the one thing that I've noticed is that it's not the best writer.
Kip
Exactly. I was going to say that it's not good for creative tasks. I find the same thing. Yeah, but thinking models, they're actually worse, I think.
Kieran
So you would take this plan and put it into something like Claude and.
Kip
Have it write the memo, do all.
Kieran
The writing of the memo and how you're going to communicate it and everything. Is that right?
Sully Omar
Yeah. Maybe this is a good way to show like how I would do that, because there is a little bit of a trick to sort of prompting it the right way. Because I wouldn't take this whole output and paste it into Claude. I would basically, like, segment it, if that makes sense. Like, let me just walk you through what I would do if I was like, okay, I want to maybe write like an internal memo to send to my team for each of these sections. Right. I could go here to ChatGPT and say, okay, you know, for each section, starting by the first one, write an internal memo to each team so I can send it to them. Let's just say it for now. We'll ask ChatGPT to do this and then I'll just take this full output, which is a lot paste. Oh, my God. I'll copy it and then I'll take this. I'll have to do a little bit of sort of prompting here. And this is a good live example of how I prompt. I sort of transfer between the two models. So it's like I have these metrics and I'll paste those metrics, you know, metrics like this and I'll paste that metrics. And I got a response from super smart model here. Smart thinking analysis. And now for each section, help me write internal memos I can send to my team. We'll leave this here, by the way.
Kieran
This is days worth of work in like 20 minutes, by the way, while.
Sully Omar
I'm on, like, the podcast and I'm trying to, like, you know, make sure the screen is recorded. Like, I'm not even paying attention to this, right?
Kip
So, oh, my God, it's great.
Sully Omar
And so here, right off the bat.
Kip
Picking dates and everything, we did a quick video, is using cloud projects to build yourself an executive assistant. It's actually a great task manager. You see the way it's starting to plot out deliverable dates?
Sully Omar
Yes.
Kip
It's like, that's another thing you didn't ask for. Like, I'm sure they're not correct, but it's pretty good at that stuff if you start to use it as like a task manager.
Sully Omar
Yeah. And one thing too, I will actually say here for how I prompted it. What I noticed is if you might say, like, okay, hey, write me an internal memo for all of these sections. What you want to do with any model is you want to break it down into as many sort of specific tasks as possible. So here I said, only do it for the CHURN reduction initiative. It looks good, I can iterate. And I'd be like, okay, looks good to me. Move on to the next. And you can obviously iterate it here. It's already. While this is running, this thing is still thinking. Just shows that there's the nuance in how you can use one and how you can take the output of one and throw it into the other. And like you say, I don't actually know how many hours of work this would quantify as and how many dollar values, but I'm going to go out on a limit. Say it's probably more than 200amonth.
Kip
Yeah, right.
Sully Omar
Like this.
Kip
Way more than 200amonth.
Sully Omar
Right? Yeah.
Kip
The thing I was slacking. Kip, I'm now really curious how far you can push this. Like if we actually give it a bunch of hotspots, data and like strategic docs, how good it is at creating a first draft of an annual plan which is like millions and dollars of human.
Sully Omar
So to answer your question, I've done that before. I've taken the entirety of our company docs and I translated it to a single, I think 50,000 token document. And I took that document and I said, here's what you need to know about the company. Every little detail is in this giant document. And I didn't actually take this to 4.0, I took it straight to 01 Pro and I said, just have at it. And I said, give me like, you know, I had all these things I wanted to know, these initiatives, I wanted to know like landing page copy. I wanted to know just everything. And it just thought for like eight minutes and it just gave me this giant output. And again, it is one of those models where I think a lot of people will use like ChatGPT to upload files. I think that's a really common use case.
Kip
Yeah.
Sully Omar
But what I've noticed in my personal experience is that a lot of the models, they can't reason over large text files. So if you throw for example, a million tokens or like, you know, let's just say 20 pages into Gemini, it'll be able to answer questions, but it will never ever be able to air quote, reason or think. It'll never be able to look at all of the data and really internalize what it means. The only model that I've found that has been able to genuinely like, here is everything you need to know about my company. Please do this for me is 0101 Pro specifically. And you'll see a lot of companies say, hey, 2 million contacts, 4 million contacts. It has this thing called needle in the haystack. For those of you don't know, it's like, if you give it 10,000 pages or whatever, how good can it pinpoint, like the thing? And you'll, and you'll see these benchmarks and these reports and they'll say, hey, it's the best. But you actually use it and you're like, well, yeah, I guess it could like control F in like two sentences, which is somewhat useful. But if you really genuinely want it to do analysis, like, Nothing comes close to 01 Pro.
Kip
Yeah. Your pitch would be like, if you have data analysts in your company, like $200 is a tiny amount to pay for the additional power you get here.
Kieran
Every analyst should have this, right?
Kip
Any serious company that's doing any type of research analysis.
Sully Omar
Yeah, I think every single analyst C suite, like, they should have this and they should find someone to grab all of the data and put it into a giant text file for them. Like, that should be a role in of itself. I don't know if there's an agent that maybe can do that.
Kip
I fundamentally believe, like, chief of staff is a great role, but a chief of staff really today should just be someone who actually has these tools and are like, the chief of staff should be completely powered by AI. That's what the chief of staff can do, is like do all the market analysis, make sure all of these things are getting done and all of that should be done by AI.
Sully Omar
And then there's no reason not to because it's not like it's a difficult thing and you're able to just get so much more leverage. Like the one thing I think that not a lot of people have been maybe speaking too much yet on sort of maybe the marketing side and sales side is because there's still a little bit of a learning curve. But what I've seen, because, you know, again, we're engineering focused company, is that on the engineering side, using the same model. Right. 01 Claude we've probably been able to, I'm not exaggerating here, 5x our output as a team. Like we're a small team and it probably feels like we have 5x the headcount and we're moving it 2x the speed because we have smaller team members. Everyone is able to know what they need to work on. And they're just sitting here with Claude or Cursor or whatever AI tool. And I think this year will be the year where I think outside of the engineering side of things, people will start to see the value because they'll be like, wait, this literally took me like a week. Would have cost tens of thousands of dollars.
Kip
Yeah.
Sully Omar
And I'm just sitting here and I'm like, just talking to this, like, computer interface and it's doing all this work. So I think this will be the year for where outside of engineering, a lot of teams will see like immense value from the analytical side because we've seen a lot of people using it for like, copywriting. But I think the analysis side is where, you know, these thinking models are able just to like, do things that. Just looking at the comparison between Claude and 01, like.
Kip
No, it's a huge difference.
Kieran
Huge difference.
Kip
Yeah. I think that was one of the best, like, comparables. If anyone wants to see why I would want to pay for this. I think just looking at that side by side is pretty great.
Kieran
I love the framing to start too. Super clear on the difference between the different models.
Kip
This has been a super great episode to like, look at the power of 01 Pro. I think our listeners are going to love this.
Sully Omar
Awesome.
Kieran
Kieran and I are both going to go buy O1 Pro right now and we got some work to do. This was awesome.
Kip
Yeah, we've been sleeping back in when you'd be showing us.
Kieran
Sully, thank you so much for joining us.
Sully Omar
Yeah, you can follow me on X sullyomar. That's most of the places that I post on and I. I like to talk about all my prompting engineering and all those kind of stuff. So that's where the place to follow me.
Kieran
This is going to be great for everyone. If you have any questions, pop those in the comments below. Give us a like and a subscribe and we'll see you next time on marketing against the Crane.
Podcast Summary: Marketing Against The Grain
Episode: "Which AI Model Should You Use? (Claude vs GPT & O1 Pro Live Prompt Guide)"
Release Date: January 28, 2025
Host: HubSpot Podcast Network (Kipp Bodnar and Kieran Flanagan)
Guest: Sully Omar, Founder of Autogrid
In this enlightening episode of Marketing Against The Grain, hosts Kipp Bodnar and Kieran Flanagan delve into the evolving landscape of AI models, particularly focusing on Claude, GPT, and the emerging O1 Pro. Joined by Sully Omar, the founder of Autogrid—an AI-first spreadsheet startup—the trio explores advanced strategies for leveraging these AI models to enhance business operations, streamline workflows, and optimize marketing efforts.
The episode kicks off with Kieran highlighting common misconceptions about using AI models like ChatGPT. He emphasizes that many users are not fully utilizing the capabilities of these sophisticated tools. Sully Omar introduces Autogrid, describing it as a "smart spreadsheet" with integrated AI agents that autonomously handle tasks such as data enrichment and lead management. He notes, “[...] every single cell is its own little agent and it gives you this nice little output” (01:25).
Kipp expands on the concept of micro agents within Autogrid, clarifying that these agents handle specific data tasks rather than overarching roles like customer support or sales. He states, “I think you're going to have agents with human in the loop doing that” (02:53). Sully agrees, explaining that past attempts at building fully autonomous agents resulted in unreliable performance, akin to a Slack channel failing 25% of the time. This unreliability leads to user frustration and abandonment of the tool (03:15).
Sully introduces his "three-tiered system" for categorizing AI models based on their capabilities and use cases:
Working Models:
Day-to-Day Usage Models:
Reasoning Models:
He highlights that reasoning models like O1 Pro require more sophisticated prompting techniques and are akin to having a "super intelligent coworker" (08:55).
Sully emphasizes the importance of crafting well-structured prompts to maximize the potential of reasoning models. He shares his personal workflow:
Initial Data Compilation: Collecting all relevant metrics from sources like Stripe or analytics dashboards.
Crafting the Prompt in ChatGPT: Using ChatGPT to help formulate a detailed prompt for O1 Pro by engaging in a conversational process to refine the request.
Using O1 Pro for Advanced Analysis: Sending the refined prompt to O1 Pro to generate comprehensive reports and strategic plans.
Sully introduces a compelling concept called "meta prompting," where an intermediary layer refines user inputs before passing them to the AI model. He envisions future integrations automating this process to enhance efficiency (11:37).
A significant portion of the discussion centers on comparing O1 Pro with Claude:
O1 Pro:
Claude:
Sully demonstrates real-time how O1 Pro generates a comprehensive report from raw metrics, highlighting its ability to produce actionable insights that surpass standard AI responses. In contrast, Claude offers more refined and human-like outputs but lacks the deep analytical prowess of O1 Pro (27:28).
Notable Comparison Points:
The conversation shifts to practical applications of these AI models in business settings:
Sully shares how Autogrid's team achieved a fivefold increase in productivity by integrating O1 Pro into their workflow, emphasizing the transformative impact of advanced AI models on team efficiency (36:14).
Notable Quotes:
Despite the impressive capabilities, Sully acknowledges the learning curve associated with effectively utilizing advanced AI models. He anticipates that as users become more adept at prompting and integrating these tools into their workflows, the broader business community will increasingly recognize their value outside of engineering departments.
He posits that 2025 will be a pivotal year where non-engineering teams, such as marketing and sales, will adopt reasoning models to significantly enhance their analytical and operational tasks (36:36).
Wrapping up the episode, Sully encourages listeners to explore O1 Pro to harness its full potential for business analytics and strategic planning. Kipp and Kieran express their enthusiasm for integrating these insights into their own workflows, underscoring the episode's practical value.
Final Notable Quote: “Every analyst should have this, right?” emphasizing the indispensable nature of advanced AI tools in modern business analysis (34:45).
Actionable Insights:
This episode of Marketing Against The Grain offers a deep dive into the strategic use of AI models in business, providing listeners with actionable strategies to optimize their operations and drive growth through intelligent AI integration.