Loading summary
A
Hey, it's Mike Stelzner. What if you had an AI system that wrote your ad copy, generated the images and video, reviewed its own work and made revisions, all customized to your brand without you ever touching it? That's exactly what Caleb Cruise is teaching AI Business Society members to build on July 16th. No coding required. Join today at socialmediaexaminer.com AI and if you're listening to this in the future and you missed it, the full recording and resource kit are waiting for you. Visit socialmediaexaminer.com AI
B
welcome to the AI Explored podcast, helping you put AI to work. And now, here's your host, Michael Stelzner.
A
Hello, hello, hello. Thank you so much for joining me for the AI Explored podcast podcast brought to you by Social Media Examiner. I'm your host, Michael Stelzer, and this is the podcast for marketers, creators and business owners who want to know how to put AI to work. Today we'll explore OpenAI models and how they can save you thousands of dollars. My special Guest is an AI consultant and data scientist who helps mid market B2B companies make AI useful. He's co founder of Trust Insights and author of Almost Timeless 48 Foundational Principles of Generative AI. His courses include Geo 101 and Geo 201, which focuses on becoming recommended with AI. Christopher Penn, welcome back to the show. How are you doing today?
B
Thank you for having me back. You know, I can't complain because it does no good.
A
Well, I'm excited. We're going to talk about some really exciting stuff today here, which is these open models. Before we really get into this, I want to ask you a question. What do you see as one of the biggest misconceptions for when it comes to generally using AI these days?
B
It's a misconception now and people are having a harsh reality check that it is not, and that is that AI is cheap. If you look at what you pay, for example, on the Claude Max plan, which is up to $200 a month, and then you look at the actual usage and at what you would pay if you were paying rack rate, it is a 97.5% discount. A Cloud Max plan gets you $8,000 worth of usage. Anthropic is heavily subsidizing this, and this is true across all the major providers. What they're all trying to do is habit formation, get people used to using AI. And we see this in the news when you look at, you know, OpenAI talking about their earnings report, saying, hey, we made $3 billion but we spent 9 billion. It's like, okay, clearly that's not a successful long term business model. And it is again because companies are trying to establish market share, they're trying to build habit formation with users and are hiding the real cost of using these tools. And that's part of what we wanted to talk about today. Because if you look at what these tools actually cost and companies at the enterprise level are now seeing this, if you use Microsoft Copilot, if you use Anthrop, if you use OpenAI at the Enterprise level, those are all pay as you go. And companies are getting extremely large bills from the AI companies. You know, they're saying, oh, AI was supposed to save us all this money. Yet Meta, for example, very recently had to put up a big pause on a lot of their AI development because their developers spent the equivalent of $2.7 billion in token usage because they did some really stupid things internally. But they really just kind of let people to spend as much as they want on an enterprise plan, which is a terrible idea.
A
Okay, so what I'm really hearing you say is it's almost like the early days of Facebook, right? In the early days of Facebook, everybody got pretty much unlimited reach and it didn't cost them anything. And then all of a sudden the algorithms changed because they had to pay their bills, right? And all of a sudden reach went down and people had to pay for advertising. Well, what effectively I'm hearing you say is that these large businesses like anthropic Google, Gemini, OpenAI with ChatGPT, they're effectively giving away the farm at a really, really cheap rate to have a lock in effect so that you're going to be with them over a lifetime. Is that kind of what I'm hearing
B
you say it is? They are trying to do that habit formation and make that the place to go. Now what has thrown a bit of a wrinkle into this is what are called open weights models. And I want to do a bit of definition here.
A
Before you go there, I want to answer this one question. Open models, when done well, what's. Let's talk about the benefits and let's define what the heck they are. Like, let's just start with like, hey, there's this alternative, this open model thing and if you all pay attention, this is the upside. So let's talk about the benefits and let's get into defining what it is, if you don't mind.
B
Sure. There's three major benefits of open models, open weights models. Number one, they are substantially cheaper to operate because they typically run either on a dedicated provider that you select or they run internally on your own infrastructure. Number two, they are the only guaranteed provider, private AI, because if you set it up properly, your data never leaves your infrastructure. So for anybody in a regulated industry, that matters a whole lot. You cannot, you should not be doing things like processing people's health information in ChatGPT. Terrible idea. And number three, so we've talked about cost, we've talked about privacy. The number three is that in terms of capability, today's open weights models are three to six months behind closed weights models in terms of smartness. So for example, the newly released Zhipu AI GLM 5.2 model is roughly equivalent to Claude Opus 4.8 in terms of benchmarked capabilities. So that's one generation behind Claude Fable 5, which is now available again. And it's available for literally even on their architecture, their structure is available for one twentieth of the cost. So if you are worried about getting, you know, huge bills from anthropic, you might say, say, you know what, let's go get some of our own hardware to run this sucker. Let's download it for free from this company, let's install it, let's run it, and we'll have an OPUS class model for the cost of electricity within our company.
A
Well, and I want to really dial in on this a little bit because right now the tokens that are being used up by these different models are billed at a very extraordinary rate. Right. But what you're saying is you can run these a lot cheaper if you happen to have a capable computer and you follow some of the stuff we're talking about today. And for even those people that are on like a hundred dollar a month plan, like I am with Opus Max, I could potentially save money in the long run by not hitting limits if I were to pay attention to what we're really talking about here today. And it, and it sounds like this is really a trend, right? There's a lot of talk right now where a lot of people are talking about this is ridiculous. These costs are completely out of control, especially in the big enterprise. And if they pay attention to what we're going to talk about today, big, medium and small businesses could all benefit from this. Is that what I'm hearing you say?
B
Exactly right. I'll give you a very simple example. I do a lot of software development. A small model, a small open model like Qin 3.6, which I run on my MacBook and I've done it on a plane, right? No Internet can write code very capably if it has a good plan. So the way that we generally tell people, and this is actually in the book Plan Big, act small, use a big foundation model like Opus to come up with a detailed plan and then you hand the plan to a small model. That's an usually open weights model. Say implement the plan. You don't have to think about anything. The plan is made, you execute the plan, and the tools are very, very good at doing that. So the recommendation for everybody, particularly if you care about things like sustainability, which is another reason to use open weights models, is you can have the grunt work, which is a lot of the typing stuff done by a small model that uses very little electricity that you can run on your own laptop, doesn't consume any fresh water, doesn't use a data center, doesn't cost you money other than electricity, and get great results.
A
Love it. Okay, so earlier you were going to kind of describe what the heck an open weights model is. Why don't you just help everybody understand conceptually what it is? Because most people listening to this podcast are familiar with the large language models, but help everybody kind of wrap their head around what's the difference between these two different things.
B
Sure. So we'll use the analogy of a car. In a car, there's an engine. There's also other stuff like seats, air conditioning, it's steering wheel, et cetera. A model, when we talk about AI, a model is effectively the engine. So these are things like GPT 5.5, Claude Opus 4.8, et cetera, et cetera. In the Western world, most of the big named models are closed weights, which means that like Claude Opus, GPT 5.5, Google Gemini, you can't go and download it, you can't see what the engine is. You know it's there, you know it works. But you are never allowed to get a version of the engine for yourself. It's always protected by some corporate entity. Open weights models, and there are a number of companies that make. These are models where the engine is downloadable. You can download it and run it for free. As long as you have the hardware to do it, you can basically get your own engine. So Alibaba's Quinn Deepseek, which is one that a lot of people have heard about, Google's Gemma 4 Francis Mistral all these companies make engines and then you put that engine into your car and you can go places. So your car would be things like a web chat interface or a coding tool or what have you. And the nice thing about the way these, a lot of the current engines and cars work like a tool, like open code, for example, you can change the engine. You could say, well, I want to do some planning, so I'm going to use Claude Opus and yeah, it's going to cost me some money to do the planning, but then I've got the plan. Now I'm going to just change models, I change the engine, I'm going to use coin 3.6 and now I'm going to implement the plan with that and it will be the cost of electricity. So instead of going into Claude and code and just running and just letting it do its thing and watch the bill go up, you can selectively choose which engine your car is running on.
A
Is open source and open weight models the same thing or are they slightly different?
B
They're slightly different. Open source means that the code used to make the engine is available. And there are very, very, very few models that give you all of the training data. Allen AI does that and Apple has done that with one of their models. But that is usually, you know, petabytes worth of data. So A, it's somewhat impractical and B a lot of that training data for everyone is confidential. A because if you hand somebody the training day and they find copyrighted works in there, now you've got a lawsuit. And B that is the training data is the secret sauce, a part of the secret sauce for how model makers make their engines. And so you can give away the engine, no one can unpack the engine and reverse engineer back to the original training data. So everybody wants to keep these things under lock and key. The difference that we're talking about today is open weights models where you can get the engine or closed weights models like Gemini where you can't get the engine, you will never get the engine.
A
You know, this kind of reminds me of the early days of Microsoft Word. And you tell me if this is a good analogy. You could download Word and use it forever until the operating system would no longer run it. But then they changed to software where you had to log in on the cloud and you're on a monthly subscription. Is it analogous to that in some kind of way?
B
It's more like Microsoft Office versus say OpenOffice or LibreOffice. Those are open source packages where you can not only download it, but you can modify itself with the open awaits model. You can download the model and you can make some, some level of modification to it. You can do what's called fine tuning on it. For example, there are very specialist models out there. There's one called Skyfall that is an excellent writer. It's not very good at coding, but it has been tuned to be a good writer. And so the open weights community has taken some of the models made by these, these open weights companies and tuned them. That's why there's 3 million of them out over on Hugging Face. Because there's a lot of specialist models and in things like marketing and business and strategy, you might want to use a tuned model for if there's a very specific task. Like for example, if you are a Python coder, you might want a model that has been tuned to be really good at Python even though it's dumb at everything else, because that is the thing you want to use that particular engine for.
A
Love it. Okay, so let's say we're interested in actually starting to use some open weight models. Let's talk about kind of the stuff we're going to need. Right, we're going to need a computer. You kind of already mentioned that a little bit. But let's talk about like what everybody is going to need to be able to pull something like this off.
B
Sure. So you need a computer that has a good graphics card. If you can play Call of Duty on like, you know, at 4K resolution, at top speed, you know, 60 frames per second and stuff, and had no trouble doing it, your computer probably has enough video RAM on its graphics card to do that. Because all AI inference runs on graphics cards, GPUs. Many computers now have specialist chips in them that can help do that. So Apple, for example, in all of Apple's M Series Max, they have what are called neural processing units that are specialized for AI. And Apple's architecture actually allows you to share all the memory on a system. It's literally called shared memory.
A
You're using a MacBook Pro, it sounds like, and I'm using a Mac Studio, so I could definitely run it on here. But you need more than just good hardware. You need a lot of memory too, don't you?
B
You need a lot of memory on a PC. You need a lot of video memory. So you will need like a really nice graphics card. Or you have a Mac or if you got some money, you can buy an AI specific device. So this is things like the Nvidia DGX Spark, which is about a $4,000 little. You know, it's like a Mac Studio size box. It sits there on your desk. Asus has the GX10 God. AMD just released a new one on their ROCM architecture, but they're all these little boxes, four to six thousand dollars. They sit on your desk, they chew up about 160 watts, 200 watts of power. And they are dedicated AI machines. They don't do very much else, but they can run AI models, open weights models on them. They can run image generation models, they can run video generation models, and they can run text generation models. So those are the three kinds of hardware you might have. You'll have either an integrated machine like a Mac, you'll have dedicated video hardware like a PC, or you'll have a dedicated AI computer.
A
And when you talk about watts and stuff, help people translate, is that really a lot of power or is that not necessarily a lot of electricity?
B
So your average MacBook is between 80 and 140 watts. Your average microwave is 1200 watts. So yeah, it's not a lot of power.
A
And then what about the software? Like, are you going to need special software to be able to pull something like this off?
B
So you need two kinds of software plus the model. You need to have the model, of course, which is the weights. You need to have a server that the model fits in, and then you need to have a client that talks to it. So this, this should sound very similar because it's kind of like a website, right? We have web content, you have a web server, and they have a web browser. Very similar. You have the AI weights, you have the AI server, and then you have the AI client of some kind. So on the Mac you would have, for example, a server like OMLX or LM Studio. On Windows you might have something like anything LLM or LLAMA cpp. Those are all servers that you would have that can serve up the model.
A
And that's just an app that you're installing on your computer? Pretty much, right?
B
Yeah, it's just an app.
A
Because sometimes when you say server, people think hardware, right? And it's not necessarily hardware in this case, it's a software server. Okay, that's right.
B
And then you have some kind of app. So apps like Open Code or Claude Code or openwork, LLM Studio has an app portion to it, so does anything LLM. So some of the things are kind of all in one. So you have those three pieces, the model, the server that makes the model available because the model is really just a big database. And then the client, the thing that you're going to be productive with, what I generally recommend people get started with when they're starting with open weights, is OMLX on a Mac, LLAMA CPP on Windows. Or Linux as the servers. And then for the app, Open Code or openwork, both of which are open source tools, completely free and very, very capable.
A
You mentioned Claude code. Is Claude code also an app that you could potentially install on the server or is that not necessarily the case?
B
Well, you can install cloud code on pretty much anything, so. And then you have to, you know, repoint it away from anthropic to your server. But you can. Yes.
A
Oh, okay. Okay. So those that are familiar with Claude code, it's possible to be able to use Claude code already on a server. You just have to like tell it where to go when it's processing its model effectively, is what you're saying.
B
Exactly. And so one of the things that's interesting is that it's all can be on one machine. Like if you have a really nice MacBook, you can do it all on your Mac. You can also, if you're a small to mid sized business and you've got an office, you've got, you know, computers in your office. There are projects, the XO Project X O is one of them that allows you to essentially network all your computers together in the office and function as an AI supercomputer. So you might not have to buy any new hardware if you can glue together, you know, the 20 or 30 Macs that are already in your office.
A
Love it. Okay. You know what most marketers are doing manually? Ad creative, writing the copy, coordinating with the designer for the images, waiting on video edits, managing revisions, every campaign, every cycle. I wanted to fix that for our members. So on July 16th, I'm bringing Caleb Cruz to teach inside the AI Business Society. He's going to show members how to build personalized AI pipelines using Claude code that runs the entire creative process on its own. Copy, images, video, revisions, all customized to your brand. And you don't need to be a coder to do it. I'll be in the room for the entire session. I'm in every session. And I can tell you members are going to walk away with a clear buildable system that they can put to work that very same week. If you can't make it live, the full recording and resource kit will be waiting for you. Transcript, slides, key takeaways, the whole shebang. This is what we do every month inside the AI Business Society. Come and see for yourself. And don't miss this training. Join@SocialMediaExaminer.com AI Again SocialMediaExaminer.com AI let's talk about scenarios where it would make sense for people to Consider something like this because I'm sure that's going to be really helpful for people to kind of like, okay, this is really interesting, but I have no idea what I would do with something like this. So let's talk about like how do we know when we're actually ready to do something like this? Like when does it make logical sense to consider something like this? And then we can explore some of the things that you've done with it.
B
The first time you get a much larger build than you expected from an AI company is when you, you go, oh, this is not working out well in terms of things that you would do. Any tasks that you would, that is I would call an execution task, write this code, write this page, summarize this document. Are tasks that open weights models do very, very well. Research this thing, you know, search the web for this. Any, a lot of agentic tasks. For example, if you use, if you've heard of agentic software like Claude Cowork or OpenClaw or Hermes agent, these are all things that do very well with open weights models because they don't need a lot of a super lot of intelligence. They just need to have good instruction following and be able to work with different tools like web search, et cetera. So any of those tasks. So that's a lot of research. That is a lot. Searching the web, there's a lot of data aggregation. I'll give you a real simple use case. Imagine you're a small to mid sized business and you want to do some prospecting for new customers and you might say, I know the kinds of companies that I care about, but I don't know who to talk to. You would have, you would set up an agent of some kind along with an open weights model and say you're going to take this company and you're just going to use your web search tools to search for the LinkedIn profiles of all the people who work at this company and you're going to write them down in a spreadsheet that I can review later. And it would go off and for 12, 24, 36, 48 hours it's going to go off and be out there searching, search and copy paste, search and search and copy paste. And then at the end it says, hey, here's your spreadsheet that I made for you. And you'd be like this is terrific. And then you would go and check its work of course, and probably end up with a really nice, some research that you could use for prospecting for your business of stuff that Honestly, an intern could do, but this means you don't have to pay an intern. And also the intern does need to eat and sleep and an agent does not.
A
I think you said earlier, plan big, act small. And if you didn't, I know we talked about it when we were prepping for this, but you kind of alluded earlier that hey, you can kind of start with some of these models that you're already familiar with like Claude or Gemini or ChatGPT and basically have it kind of build the bones for this thing and then take it over an open model. Is that, can you talk about that a little bit?
B
Exactly right. So in a big model like GPT 5.5 inside of ChatGPT or Gemini 3.1 or Claude Opus 4.8, you might say I want to build a plan to implement a piece of software or to build a prospecting list or whatever. And you would go back and forth with it. Maybe you're using something like one of the brainstorming plugins and say here's what I want. Ask me questions about this until you have enough information to write up a detailed plan that I can hand off to an AI tool to execute. And you'll have a 25 or 30 minute conversation chatting back and forth forth. You know, maybe it's taking notes in its canvas for you along the way. And when you're done, you should probably have like a 10 page plan of all the stuff that you want a machine to do. We call it the 5P framework by trust Insights. So purpose people, process, platform performance. The purpose of what I want to do is I want to do some prospecting so I can grow my business by finding new people I haven't reached. Right. That's a pretty clear purpose people. Who are the people that you, you want to look for? Process? Hey, you're going to use your web search tools, you know, platform, maybe you use DuckDuckGo. It's got a, it's very friendly for AI agents and the performance you succeeded when you've got 200 new prospects by LinkedIn URL and their LinkedIn URLs all work so that I know that you didn't make it up. Following the 5P framework by trust Insights, you would work with the big models, build the plan and they download it as a PDF or as a markdown file or what have you. Then you go to your agent running on an open weights model and say is the plan, implement it and it just goes off and does it.
A
Okay, you casually said a brainstorming plugin so somebody listens, like, ask them about the brainstorming plugin. Is there one that you use that you really love and is it some sort of skill that you would just install?
B
It is, yes. Jesse Vincent's Superpowers. It's over on GitHub. It's completely free. It works in Claude Code Codex, Anti Gravity, and a bunch of the agentic software. And so it's terrific.
A
And it basically just asks you a bunch of questions or ideates with you. Is that effectively what it does?
B
Yes, exactly. It's a. It's actually an entire coding system, but the brainstorming part is applicable to everybody.
A
All right, so let's talk about some of the applications that you've done with this thing. Right? Because you've talked about how, hey, using open weight models are really good for executing repeatable regular tasks. And I know there's some repeatable regular tasks that Chris Penn does because you create a lot of video. So maybe you could talk to us a little bit about some of the cool things you've been able to do for yourself.
B
One of my favorite daily agents that runs is called the local paper. If you noticed, the quality of local news has gone down pretty precipitously over the last 20 years as companies have tried to fight for clicks and stuff like that. And there's a bunch of stuff that I care about locally that no newspaper covers anymore, like, oh, the local cornhole tournament on the town common is next Saturday. I worked with Claude first and then with Hermes Agent to build a local paper. And there's a series of Python scripts that run every morning to go check the city website, they'd go check the park website, the Chuck Public Library website, check the police department's website, grab all the data, and then an open weights model, which is Quinn 3.6, that runs on my laptop, reads all the data and turns it into a little newspaper. And then it drops as a PDF inside my Discord server. And every morning I've got a paper to read that is actually the stuff that I care about. Even though commercially this paper would be a failure, like, no one would read it as so boring. But it's all the little stuff that, like, oh, there's a French cooking class, the public library this week. Okay, cool. I would want to know that doesn't sell papers in the commercial world, you
A
know, and as I'm thinking, there are SaaS products that we pay for that have not kept up with the modern times. We could get very creative, couldn't we? And create our own little local version of it that crushes saves up those, those monthly rates, couldn't we?
B
Absolutely. So I used to have like 10 different WordPress plugins on my, my website that I all cost, you know, 9.95, $19.95, whatever. And one by one I have handed them off first to Claude Code and then to an openoids model to essentially just rewrite them. Say, you know, here's this piece of software, here's what I use. There's a whole bunch of stuff I don't use and there's some things that really annoy me about this. So let me go ahead and have AI first plan and then build with an open weights model a version of the software. And because the open weights model built it and documented it, I can also maintain it. So I am my own tech support, AI is my tech support now. And because it's again because the open waste models are now so smart they can effectively triage and fix things. They have, they know how to build tests. And so my bill monthly of software subscription keeps going down and down and down because a lot of these tools don't have a value add above and beyond the software. And software is effectively free today. What is valuable is either human stuff that goes on top of the software or data that nobody else can get. But software is a commodity.
A
Love it. Okay, so that's just giving people some possibilities. Now let's talk about what you've referred to as the engine, which is the open weight models. You've mentioned Quinn, you mentioned Gemma. Four, you mentioned Deepseek. Let's talk about a little bit of the models and why you might want to use different ones and just help people understand what their pros and cons are. For these different models, there's two things
B
to get to pay attention. Number one is how big is it? Because how big it is means how much memory you need. So for example, if you look at Quin 3.6, the 31 billion parameter model, size wise, that thinks about 35 gigabytes, right? Which means that you need that plus you know, an additional 25% in memory free memory available on your machine. So if you have a 32 gigabyte Mac Studio, that's how much RAM you have and say 10 is used up by the operating system. You only have 22 left. Quin 3.6, that 35, that 31 billion parameter model is not going to fit on your machine. However, the 12 billion parameter model probably will. So a big part is to look at and if you go to a site like Hugging face, you can actually see how big is this model on disk?
A
Because it runs in memory. Is that what I'm hearing you say? It doesn't run on a solid state drive, it runs in memory effectively all the time, right?
B
Yeah, that's right. It has to run in memory. For now it has to run in memory. So that's the first part is how much model can your hardware support? Like if you went all out, you spent six grand on the nicest MacBook and you got 128 gig of memory, you can run a good number of models very very well. Second is whether it is. It's what's called a dense or a mixture of experts model. Mixture of experts models you'll see in the name have two sets of numbers in them. So Quinn 3.6 35b3ab is a mix of experts and what that means is that for any anytime you're using it, only a small part of it is active. But internally has these little experts that kind of do its own routing to make it very fast. It's less accurate, but it's super fast. So if you're doing stuff like article summarization, those are the, you know, those are exactly ones to use. So if you're doing a say social media sentiment scoring, you want one of these because they're going to, they can chew through a lot of data very quickly.
A
Is it generally disclosed that it's a mixture of experts model? Do you know that going into it?
B
Okay, you know that bit by the name because the name has three pieces to has the name and it has total number of parameters AKA how much knowledge it has. And then it has what's called active parameters, AKA how much knowledge is active any at any given time when you're using it.
A
So if it's a lower number than the other number, that means it's a mixture of experts. Is that effectively what you're saying?
B
Yeah, effectively. So you'll have like Quin 3.6 which is the model name. 35B is the total amount of knowledge it has. 3AB is how much is is active at any given time. If a model does not have that last set of numbers is what's called a dense model, which means that it's all its knowledge is active all the time, which can be kind of a waste. Right, because you don't need its knowledge about French cooking when you're writing Python code. So those are the two types of all models. But open waste models that you're going to see the most.
A
Okay. And let's talk about kind of the pros and cons between some of the bigger options that are out there. You've mentioned Quinn a lot. You've mentioned some of these others as well. Let's talk about like Gemma for who it's from, Quinn who it's. Let's just kind of do a quick overview flyby of all the options.
B
Sure. So there's hundreds of options, but the ones that I find the most useful are Alibaba, which is the Chinese e commerce company, makes the Quinn family of models. And this is really important, this is an important distinction. If you use their web based service on their website, you are using their models that are based in the People's Republic of China and you do not have data safety. Right. So if you're using their website, it's not safe for your data. However, if you download the model and run it on your hardware, it is completely safe to use. So that's the important distinction because a lot of people get very concerned about Chinese AI.
A
Explain what Alibaba is so people understand why it's so darn powerful. Because that's kind of like the Amazon of America, isn't it?
B
It's bigger. Yeah. Alibaba is the largest e commerce company on the planet.
A
So they've got all this data that they've used to train Quen, which is why it's so darn good. Is that right?
B
And they also have really, really, really good scientists. One of the perpetual jokes I have is that generally shouldn't get into a math war with China.
A
Okay, so we got, we got Quinn by Alibaba. That's, that's, you know, a Chinese model.
B
Yep, they have a great selection of models. The second is Google. Google has a model family called Gemma. This is out to version four and Gemma comes in a variety of flavors as well, all sorts of different sizes. Very smart model. Also very capable. If you think about it. It's kind of like what Gemini Flash
A
is like the latest Flash model basically, right?
B
Essentially, yes. And then depending on the size, as the size of the model goes down, its intelligence goes down. So the small version is kind of like flashlight, but that's an excellent family that can run on a lot of different devices.
A
Okay, and what about any other options
B
we should consider if you have either a lot of money for a lot of hardware or you're working with what's called an inference provider, companies like Cerebr, Grok with a Q or Deep Infra, they've spent, you know, hundreds of millions of dollars on hardware. So there's two other families that are very smart. One is the Minimax family, also by a Chinese company called minimax. Their M3 model, you need like $50,000 worth of hardware to run it, but it is a super, super smart, very powerful model. And then finally is the Deep Seq family from the Chinese company Deepseek. Again, you're going to need 50 grand of hardware or use a third party company that's, that is a cloud hosted provider and Deepseek version 4 flash are generally considered sort of the best open weights models. There's also GPU AIs, GLM 5.2, but that's kind of a lesser known model at this point. Still very good. The company that I use in the US is called Deep Infra. They're based in Sunnyvale. They have a zero data retention API, which means they record nothing when you use their APIs. And their pricing is pretty good.
A
Okay, but Deep Infra is. What is that? What is, what are we talking about? Is that a model or is that something totally different you just brought up?
B
That's a hosting company. So just like, you know, you can host a web server on your computer, there are also hosting companies. So there are AI hosting companies that host these models, these open weights models.
A
Ah, okay. So if you don't want to have a computer, you could put any one of these in a hosting situation. Is it pretty expensive or is it pretty reasonable to host?
B
It's most of the time it's pay as you go, so you're paying by token the same as you are with Claude or, you know, OpenAI, but at like one tenth of the cost.
A
I see. Okay, so it sounds like most people are going to start probably with, with either Quinn or Gemma to begin with. And if you're concerned, like how much better is Quinn over Gemma for those that are considering, you know, Google versus
B
Alibaba here, in my experience, Quen is a smarter model, particularly at tool handling. So if you're going to be using it for AI agents, it is a smarter model for that. So if you, if you're setting up a Hermes agent or an open claw or something that's going to do a lot of autonomous work, Quinn is definitely smarter at that.
A
Got it. And if you're just doing basic data processing and stuff, Gemma is probably going to be just fine at it. Is that what I'm hearing you say?
B
Exactly.
A
Anything people need to know about how to set these things up and run
B
them, it's not complicated to get set up. It is complicated to get it Perfected for your specific use case.
A
Okay, so if people want to learn about this, is this where a tool like Claude Code or Cowork could kind of layer over the top and help you set it up and kind of see what's going.
B
Yep, you could do that. You could also just use Google AI
A
Mode, tell everybody what that means.
B
If you go to google.com, the default search button now says AI mode and it literally is a chat with Gemini
A
and Google knows code pretty well and it can help you set this stuff up is what I've discovered. All right, we picked a model and now we got to talk about creating some of these apps and skills, right? So you've already mentioned a couple different tools that you're going to be using, which you refer to as app. We've got the server, right? We got the LLM and then we got the client side app. I don't know if I got that right or not. So talk to me a little bit more about what do we do now that we've got the model picked and we want to start creating some apps and skills, Right?
B
So you've got the model, you've got the server running that's serving up the model. Now you need the client. So I recommended either Open Code or openwork as starting points. Both are very, very good. Open Code also has a desktop app which is very nice if you've not used a tool like Claude Code or Claude Cowork. These are essentially chat style apps. You connect to your serv, you say, hey, I've got the server here, you're going to use it. It will say, great, I'll go look it up and looks it up and it says, okay, I'm ready to go now. It's literally just like working inside of Claude Cowork or ChatGPT or what have you. You're ready to go, you're ready to do anything that you can do in like a chat GPT you can do in these tools. You can tell them to do web searches, you can upload documents, you can drop in images. And as long as the models that you know that you've picked support those things, they can very capably be your personal assistant.
A
Okay, so let's talk about Open Code versus openwork, right? Because those are the two that you recommended, like what are the pros and cons and your professional opinion on each one of them. So people might be able to listen to you and kind of discern which one maybe they they want to work with first.
B
Open Code is tuned for Coding. It's very good at writing code and a lot of its built in system prompts are meant for coding. It can do everything, but it will do coding best. Openwork is a general productivity tool, so its prompts are tuned more for general work like making PowerPoints and opening spreadsheets and things like that and not coding as much. So if you're a small business owner, say, and you want to close out the books at the end of the month, you download your spreadsheets from QuickBooks or whatever, openwork is the best choice. You drop your spreadsheets in there, you turn it on, say, hey, help me close the books. Here's the, you know, here's what I'm trying to do. And it will go off, it will do those things. You could do it in open code, but it would take a bit more prompting.
A
Is this by the same company?
B
No, they're separate projects.
A
Okay. And they're open source, is that right?
B
That's correct.
A
So, okay, one of the logical questions that I have in my brain is that, okay, these models keep changing, right? And let's start with how often should we upgrade our models? Let's just start there. Because one thing we know is, well, I can speak for myself. I'm a writer and Opus 4.8 is not as good of a writer as Opus 4.6. So you don't want to just automatically upgrade the models. Or do you like, let's talk about that a little bit. Like, because they keep coming out with more capable models, but sometimes their capabilities are in very narrow areas. And I'm not sure what you would recommend for people. Do you recommend they experiment or wait for others to do experiments? Like, let's talk about that a little bit.
B
I absolutely recommend you experiment. And the nice thing with open Weight's models is because it's literally just a big, big file on your computer in your server. You could say, okay, hey, Quinn, 3.7 just came out, I downloaded it. Switch. And you switch that model and you, and you run. You have a benchmark test of some kind, like, hey, I'm going to give you some inputs, write a short story about this or what have you, and you look at the results. Maybe you run five or six times and you compare and say, okay, clearly this newer model didn't do as good a job. So I'm going to keep my old model around. And that is one of the nicest things about open weights models. When Claude, for example, when Anthropic says, hey, you know what, we're retiring Sonnet 4.5. You're like, yeah, but I like that one. Like, too bad you can't have it anymore. With open weights models, as long as you got the disk space for it, you can have. If you. There's a model that you really like that does a great job at a specific tactic, you just keep it around and you switch it around when you need it.
A
Does that mean when you're using a tool like Open Code or openwork, it just has a little selector, kind of like what we're used to seeing inside of Claude and these other models where you just select which model you want to use and then you just apply the task to it and it'll just select the options that are downloaded on the device. Is that correct?
B
Yep, exactly. Right.
A
Interesting. Okay, each time we add one of these things, how big are the file sizes typically? I mean, if these are running in memory all at the same time, I mean, help me understand how this could be a real problem, right? Or is there only one running at a time typically?
B
It would typically run one or you can run it. Well, it depends on how much memory you have. So on my machine, I have a server called OMLX which after a model's been idle for 10 minutes, it unloads it out of memory so that it frees up memory. But I have scripts that will say, hey, load up Gemma and do this thing. And it will do that. And then after 10 minutes, it just puts Gemma away, puts it back on the shelf, if you will, of another script that uses Quinn. You know, when that script runs the servers, it asks the server, hey, I need Quinn. So pulls out of the library, loads it up. Now I have a 2 terabyte external hard drive that has. Is filled with models. So yeah, you do want to have some prudence because remember how I was talking about, you know, the. The billion parameters in the model name that basically tells you how large it is on disk. So if it's a 31 billion parameter model, it's probably a 31 gigabyte file of text. So yeah, you can run out of disk space real fast.
A
Well, and that does beg another question, which is can you get these solid state drives and just plug them into your computer or do you need certain kinds of. I mean, I guess it doesn't really matter because it loads it into memory effectively, is what we're saying. It's not currently referencing the models currently on your drive, but that's probably going to come, is what I heard you say. What about the apps like Open Code and openwork? I would imagine they constantly have updates coming out. Do you recommend to just keep it up to date on the latest one or not necessarily do that? What's your thoughts on that?
B
Yeah, you can definitely keep those up to date because those are basically the rest of the car. They get maintained frequently. Where you put your car customization are in things like you reference skills. For example, if you've used skills in Claude cowork or Claude Code or OpenAI Codex, everybody follows the skills format because it's an open standard in the industry. So you can take your skills right out of Claude code and put them into open code and they will work one to one.
A
And can you update your skills with these local models and bring them back the opposite direction as well or.
B
Yeah.
A
Okay. And does that mean. Do you recommend there being a central repository for your skills in general so that when you update them in one place, somehow you've got a central space on that?
B
Yep. If you're an organization, something like SharePoint, because the skill is just a text file, it's a CS text file. So you can put on SharePoint. More advanced organizations will use a private GitHub repository that everybody checks skills out of in one place. You can put them in your Google Drive or your Dropbox or whatever. Any place you would centralize a pile of files is where is where you could store your skills.
A
You're clearly very technical. Chris, one of the more technical people I know. How technical does one need to be to be able to begin with something like this, in your professional opinion?
B
It helps to have a technical friend to set it up. Once it is up and running, it's pretty easy.
A
Okay, what's something people might want to consider as an experiment? Just like if they're going to download and experiment with this stuff, what's the simple first thing that they might want to consider? Maybe trying as a proof of concept.
B
If you will take something simple like, hey, here's a giant research report summarized just so you can see it happen. Because one of the really cool things about these tools that I like a lot is that you can see under the hood. You can watch it working, you can watch it building the tokens, like literally assembling them. Which does to me is a super useful way of dispelling the idea that these tools are somehow magic or whatever. Like you would watch it just through predictions going, oh, it's literally just guessing what the next logical sequence is. And it takes away that magic sense, which means that it helps you remember and keep some distance. Like, these are not Living machines, there's no sentience, there's no sense of self. They're not your friend. Right. It's just a prediction engine at work. And I find that's a very good thing for people to see because we forget that, you know, we, you know, you start chatgpt, your clothes says hey Michael, how what do you want to work on today? You're like oh that's so nice. And you realize, oh, it's just a machine.
A
So if someone's putting this on a laptop and they close the lid on their laptop, effectively this thing is going to go to sleep, right? Or is it not going to go to sleep? So is there something to having a computer that's on all the time, a separate machine from your normal workstation for this kind of thing? And also, is this where cloud hosting might come in handy because it's never going to sleep?
B
Yes, absolutely. And again, there are three different kinds of hardware. That dedicated AI computer might just be the thing with the acknowledgment. You got to remember if it's on all the time, it's using electricity all the time. So if sustainability is your concern, obviously you might want to not have it on all the time.
A
Time. Does your company provide services for this kind of thing? If people want to work with you on other kinds of projects like where would they go?
B
We do in the mid market and enterprise, particularly for people who get those really large co pilot bills. But yeah, if you, if you want to check out who we are, what we do, go to Trust Insights AI
A
and then Chris, tell everybody if they want to like follow along with all the great stuff that you're doing because you're constantly creating content. Where would they go for that?
B
TrustInsights AI is one place and the second place is ChristopherSpenn.com those are the two places that you can find everything else all over there. But those are the places that I put my best stuff.
A
Yeah. And folks, I would suggest you subscribe to his newsletter as well because every week you come out with a ridiculously long newsletter. And you're also creating a lot of great content on YouTube as well. Where can they find you there?
B
S. Penn is pretty much my social handle everywhere, including on YouTube. Yeah, my YouTube videos I make every Sunday morning while I'm making, you know, Sunday meal and stuff because I just, I literally just stick my phone to the side of my refrigerator and cook.
A
That's super cool. Christopher Penn, thank you for coming on the show and answering all my questions about open wave models. I can assure you some people listening are going to give it a shot. So thank you so much for sharing your insights today.
B
Thank you for having me.
A
Hey, if you missed anything, we took all the notes for you over@social mediaexaminer.com a116 if you are new to the show, be sure to follow us on whatever app you're listening to us on. And if you've been a listener for a while, we would love a review. And do share this with your friends. You can tag me on Facebook, LinkedIn and or X. And do check out my other show, the Social Media Marketing Podcast. This brings us to the end of the AI Explored podcast. I'm your host, Michael Stelzner. I'll be back with you next week. I hope you make the best out of your day and may AI help you become more successful.
B
The AI Explorer Podcast is a production of Social Media Examiner.
A
Before you go, if you want to stop grinding through the ad creative process manually, the AI Business Society can help. Caleb Cruise is teaching members how to automate the entire process on July 16th. Full recordings and all the resources are included with your membership. Join@SocialMediaExaminer.com AI again SocialMediaExaminer.com AI.
Host: Michael Stelzner
Guest: Christopher Penn, AI consultant & data scientist (Co-founder, Trust Insights)
Release Date: July 28, 2026
In this episode, Michael Stelzner dives deep into open-weight AI models with expert Christopher Penn. The discussion centers on how these open models offer substantial cost savings, privacy benefits, and operational flexibility for businesses of all sizes—especially as the true costs of closed, subscription-based AI become clear. Penn offers advice on getting started, hardware requirements, practical use-cases, and the best available models. If you’re a marketer, creator, or business owner seeking practical, actionable guidance to harness AI without runaway costs, this episode is a goldmine.
(01:56–04:33)
“It is not, and that is that AI is cheap…. Companies are getting extremely large bills from the AI companies.”
—Christopher Penn [02:14]
(04:18–06:19, 08:27–11:15)
Definition & Analogy:
Three Key Benefits:
“You can basically get your own engine…. For literally one-twentieth the cost.”
—Christopher Penn [05:27]
(18:58–22:43)
“Plan big, act small. Use a big foundation model to come up with a detailed plan, and then you hand the plan to a small model.”
—Christopher Penn [07:31]
(12:48–17:24, 41:52–42:09)
(26:04–33:09)
“If you use their [Alibaba] web-based service… you do not have data safety. If you download and run the models, it is completely safe.”
—Christopher Penn [29:24]
(23:36–26:04, 24:42–26:04, 28:15–29:06)
“I have handed them off first to Claude Code and then to an open-weights model to essentially just rewrite them.”
—Christopher Penn [24:54]
(33:10–39:51)
(40:13–41:33)
Does it Require Heavy Technical Knowledge?
Starter Experiment:
“One of the really cool things about these tools is… you can see [them] under the hood… a useful way of dispelling the idea that these tools are somehow magic…”
—Christopher Penn [40:41]
On Habit Formation and Subsidized AI:
“What they're all trying to do is habit formation, get people used to using AI…. And are hiding the real cost of using these tools.”
—Christopher Penn [02:14]
On Cost and Privacy Benefits:
“Open weights models… are the only guarantee of private AI… your data never leaves your infrastructure.”
—Christopher Penn [05:11]
On Real-World Automation (Use Case):
“You would set up an agent… to search for the LinkedIn profiles… and at the end it says, ‘Hey, here's your spreadsheet that I made for you.’”
—Christopher Penn [19:52]
On the Value of Experimentation:
“I absolutely recommend you experiment… you look at the results… and you switch it around when you need it.”
—Christopher Penn [36:35]
On the Non-Sentience of AI:
“It takes away that magic sense, which means that it helps you remember… [these] are not living machines… It’s just a prediction engine at work.”
—Christopher Penn [41:10]
Open-weight AI models offer a timely, cost-effective, and privacy-focused alternative to cloud-based, closed models from major vendors—without significant sacrifice in intelligence for most business tasks. Christopher Penn guides listeners through the what, why, and how, providing concrete examples and tool recommendations for anyone ready to regain control of their AI costs and workflows.
For more details and show notes, visit socialmediaexaminer.com/aipod