
Loading summary
George Sivulka
You'll actually have hybrid AI and human employees working alongside each other. People that are really good at prompting be the best managers. Intelligence will become too cheap to meter.
Matt Turk
Welcome back to the MAD Podcast. In this episode, I'm thrilled to sit down once again with George Sivulka, a Stanford prodigy who walked away from a fully funded PhD when the launch of GPT3 convinced him that the real action.
Mat
Was outside the lab.
Matt Turk
Instead, he decided to launch Hebia and became one of the very few founders to get a pre seed investment from Peter Thiel. Fast forward to today. Hebia has become an AI platform that chews through billions of pages for top Wall street asset managers and tier one law firms and has raised 160 million in venture capital. In this episode we explore why generalist AI beats vertical AI.
George Sivulka
The best investors reading really crazy data sources that have nothing to do with finance. The best lawyers have backgrounds that are way outside of law. You don't want specialization, you want generalization.
Matt Turk
The power of inference, time, super scaling.
George Sivulka
And actually the whole scaling ed inference paradigm was pioneered at Hebia. We're one of the first people to say, hey, you get way better accuracy from using more large language model calls.
Matt Turk
At runtime and how AI is already redesigning org charts.
George Sivulka
And when you think of that from an organizational design perspective, you can start to define the agent employee as another node in your org chart. It will probably have to have an email, it'll probably have to have a slack, it'll probably be doing things the wrong way and you'll have to manage it in the right direction.
Matt Turk
This episode was recorded live at a recent Data Driven nyc, the in person monthly event we run in New York in partnership with our friends at Foursquare. Please enjoy this energizing chat with George Sivulka.
Mat
Good to see you.
George Sivulka
It was, I think almost two years ago to this day where we started talking about Matrix in a very similar setting.
Mat
Actually Matrix was, if I remember correctly, not even a name then it was stealth.
George Sivulka
We it was pre marketing lingo, but it was, it was still in stealth, but a lot of the early ideas and you were my first podcast ever. So very excited to be back.
Mat
Wonderful. Well, welcome back. So maybe for anybody that did not attend that event or does not yet know about hebia, what is 45 second elevator pitch for what you do and.
George Sivulka
Why it matters if we're in an elevator? Hebia is we actually build AI agents for finance, for investors, bankers and lawyers. So anyone that does white collar work, whether you're doing or you're discovering as Part of what you work on, we build an AI agent platform so that you can leverage some of the latest and the greatest.
Mat
And you know, rewinding back to that chat from a couple of years ago, I remember that you had a very powerful kind of mission statement which was to keep smart people from doing stupid tasks. And if you watch that back and if you sort of fast forward to today, are you still doing exactly that? In what way has Mission evolved?
George Sivulka
I think when we started Hebia, I think that one of the fundamental insights was that you had a lot of the smartest people in the world doing the stupidest tasks. And it seemed like there was an arbitrage opportunity there to save them some time or some effort or some sweat. And the goal of Hebia was never to just stop at really highly paid knowledge professionals, the investor or the finance guy, the lawyer. It was actually always to go much more broad. And our vision and mission have solidified really over the last couple of years into building capable AI platform for a billion people. And the idea is that there's lots of consumer AI products, products where you can go and, you know, talk about your sushi restaurant and wine in San Francisco, but there's actually not very horizontal general purpose workplace AI products. And what we mean when we talk about capable AI and AI that can do things is actually something much more horizontal. It's something that's much richer in its interaction. It's actually not as verticalized as you might think, but it still has the depth of what an expert platform would be. So if you think about how that vision and mission have evolved, we started on Wall street, we started with lawyers, we started with investors and bankers. But now we've built an AI platform for experts that lets anyone, whether you can code or you don't know how to code, or you can only use cursor to code to build whatever kind of knowledge, work, application or agent you'd like.
Mat
Could you rewind back to the initial sort of light bulb moment when I guess you were a student at Stanford, if I remember correctly. How did that all come about?
George Sivulka
It was 2020, and at the time, around June or May of 2020, no one was talking about AI, no one was talking about large language models. And I remember, I think the topic of what I was studying at the time was meta learning, or the idea that you could figure out how to build a machine that could learn how to learn or generalize to any task. And I thought that meta learning would become the most important software of all time. And at the same time, I was working on in my research. And then all of a sudden, June of 2020 comes around and OpenAI actually released GPT3. So before ChatGPT, I remember using it, and I think, you know, after using it a few times, I sat back in my chair. I was like, this thing is a meta learner. It has beaten me to the research punch that I was working on. And if I couldn't actually play a part in creating that fundamental technology or that technological revolution, I knew I wanted to be a part of applying that in a really interesting and meaningful way. And I made the bet that, hey, this would be worth basically throwing my whole life, my whole research career away on and starting a company from scratch.
Mat
And not to make you blush, but. So you were doing like a PhD. You were like 21 or 22 when you were starting a PhD.
George Sivulka
I'm blushing. Stop Mat. Not really.
Mat
Stop telling me. Stop forcing me to tell people how smart I am. No, but there was something like that. You graduated super early from Stanford. When did your PhD just tell that story in two sentences?
George Sivulka
I was a young hustler, is the reality of the story. But I was one of the youngest people to graduate Stanford and was the youngest person in my PhD program. I was fully funded, so I was throwing away at the time, millions of dollars of research grants and not having to be a TA for the extent of my graduate school, which was actually quite the luxury. Didn't know I could raise. Didn't know we'd be able to build a company that we built and the team that we built.
Mat
Great. So, speaking of that, so you know, what's, what's. What's happened over the last couple of years, like any, you know, metric, including vanity metrics that you can share, fundraising history, number of customer, number of documents, whatever it is that you want to share, to give a sense to people for the reality of the company.
George Sivulka
As of today, I can't forget that you're a vc. So I'm not allowed to share any metrics over here, but there's some public ones. And I think things that I'm really proud of, are really lasting, are first and foremost, our team. So we built a team of 100amazing, incredibly smart folks, all five days a week, sometimes six, in office in New York City. And we are now just starting to become a multinational corporation. And so we're opening sf and we've already opened a London office, which is incredibly exciting, with goals to end the year at 300 to 400 employees. Lots of exciting growth, A lot of it Here in Silicon Alley and unfortunately Silicon Valley, we have to move out there a little bit. But I think on AI metrics or a little bit more about the product and how it's used, one of our favorite things to track is the amount of unstructured data, the amount of pages that are processed by the platform. And a really interesting thing is, hey, last year Hebia and probably all of the other major consumer model providers processed around 100 million pages, probably around whatever, hundreds of years, maybe thousands of years of reading. This year we're already on track to process around 4 to 5 billion pages. So somewhere around 50,000 years of reading for a human that's, that's taking the right amount of breaks. And that exponential is just one of the most phenomenal curves that you've ever seen. It's like in a big board in our office. And I think the reason we're so proud of it is whereas other AI platforms are having a bit of a transactional relationship with the AI, you ask a question, you get a response. With Hebbia, you can give it these complex tasks and it really turns through vast quantities of data and does work the way you work. It's much more of an agent than a chatbot. When you're actually looking at the work that it's doing, or that it would take in people, hours, a team of really highly paid professionals to do, it's actually having a massive impact with the organizations we've rolled out now. We're deployed at, I think between 40 and 50% of the world's largest asset managers, some of the tier one investment banks in the world. We have an incredibly fast growing legal segment. So I think even just sort of the share of our revenue that's in law is increasing at an exponential clip. And so it's just an incredibly exciting time to be at Hebbia and in AI.
Mat
So we're going to go into all things Hebia, the product, the tech and all the things. But before we do that, just maybe six or seven days ago, you had a really interesting blog post that you published where you talked about the concept of agent employee.
George Sivulka
Yes.
Mat
So going to that, what is an agent employee in your world?
George Sivulka
I think that the idea behind the blog post and talking about an agent employee actually stemmed from a way that I think organizational design is starting to change our customers and actually likened it to an old organizational trichotomy of remote work where the Internet actually came out and took. It basically decoupled the output of labor from where you were so you could be all in the same room in New York, or you could be across the world and the Internet, basically democratized access to talent. With AI. And these agent employees, you're actually also starting to see, hey, I'm going to have full time things, whether they're employees or not employees or whatever you want to call them, where instead of actually having output or work or labor coupled with personhood, it will now be completely decoupled not only from location, but from whether or not you can pay salaries or have humans doing the jobs. And so the entire essence of the piece was, hey, just as we have remote employees, we have hybrid organizations, fully remote organizations, and in person organizations, you're going to start to have fully human organizations, actually fully AI organizations like the one person billion dollar startup, or, you know, these things that are effectively just APIs. And you'll actually have, and this will be the most common thing, hybrid AI and human employees, agent and human employees working alongside each other. And when you think of that from an organizational design perspective, you can start to define the agent employee as another node in your org chart. And it becomes very clear that if that's a node in your org chart, it will probably have to have an email, it'll probably have to have a slack, it'll probably be doing things the wrong way and you'll have to manage it in the right direction. And the makeup, whether you're 50% agents or 10% agents, or 90% agents or 0% agents, will actually define the output of the organization. Amazon is a perfect example of how this happened, how org design shaped what they built. If you look at AWS's offerings, every single offering in that big menu is a different startup altogether with its own gm. And that intentional organizational decision impacts their product, it impacts how it works, it impacts whether or not those products work together really elegantly. The same will be true as you start to think through these agent employees as nodes and how you organize them and whether or not you're, you know, where you are on the 0 to 100% AI agent dialectic.
Mat
So fast forward a few years, what does that all look like? Whether you call it, you know, 2030, let's say not 20 years out. 2030. So we all, what AI managers? We spend half hour days interacting with agents. Do we still interact with humans? What does that look like?
George Sivulka
I think that right now we're already AI managers. The only difference is the AI takes a single step. So if you are a AI modern organization, you probably are using some sort of chat or rag application and how good you are at prompting that AI is how good you are at managing that AI. And instead of, you know, kind of letting the AI go and prosecute a task over and over and over and self correct, it can only take a single step. As agents are rolled out, you'll actually start to see, you know, people that are really good at prompting, really good at defining a process, be the best managers and actually be the best at extending whatever their agenda is in the organization or making their function the best. And so I think that everyone will be prompting and prompting is managing and it will all blur pretty soon.
Mat
As a quick reality check, I think you mentioned somewhere that AI agents will contribute more to GDP than human workers within a decade. How backload is that? In other words, what's your sense as somebody who's in the proverbial trenches every day of the reality of agents in the enterprise in terms of what they can do and what's overhyped?
George Sivulka
We will still be deploying AI in its current form as a chatbot in 10 years. I think the way to think about that is there's actually still plenty of businesses in New York that don't take credit cards. They only take cash. And organizational change and technological change will take a lot of time and a lot of effort at the same time. Most people use credit cards. And when things work, and they're not just experiments, they actually happen very fast. You're starting to see that with certain AI companies that are massively penetrating markets that have gone from zero to some double digit percentage of their sam. And we're fortunate that Hebia is one of them. But at the end of the day, the stragglers, a long tail of adopters will still take time. And I think that to answer your question on the nose, everything is going to be backloaded, right? I think it'll happen in the decade. Like the change to credit cards happened over five to 10 years and now this change to just point and click and credit cards happen even faster than that. But there will still be stragglers. It'll solve the time.
Mat
One last question then, before we get into Hebia in detail. What do you find particularly interesting in AI today? AI research, open source reasoning models that, you know, you may or may not sort of import into Hebia directly. But what do you find interesting in terms of like what's happening in research? You know, there seems to be a new model every other day, a new thing every other day. What catches your attention?
George Sivulka
That's a good question. There's the felt sense, I think in communities like this one, and then maybe the larger AI community that the scaling laws for training have slowed down a little bit and we really haven't had a massive paradigm shift that has been released recently. At the same time, there's a lot of really interesting research direction in scaling laws for scaling during inference. And what that means is, hey, maybe we have a fundamental unit of compute and that is a single inference, a single forward pass on a large language model. How do we actually now use that? Or run lots of inference like an agent, like OpenAI0304 and all of these kind of more agentic, more reasoning models to actually get better at doing these difficult tasks. And actually the whole scaling at inference paradigm was pioneered at Hebia. So our early matrix product two years ago were one of the first people to say, hey, you get way better accuracy from using more large language model calls at runtime. And so we built lots of infrastructure to scale that up. We actually built an agents team before it was even called agents to actually go out and run these larger jobs. And I think that's probably the most interesting research direction moving forward is like, hey, let's go and say this current scaling law of training has slowed down the scaling law of inference or test time compute is very interesting. Let's double click there. There's lots going on with multimodal research. Obviously context windows there, how you apply them, how you tie them in to the larger ecosystem of foundation models that are used for text applications is very interesting. And I think the number one problem in all of AI is elongating the context window. And so we have a few different metrics. It's now a focus. You've got these needle in the haystack Twitter threads that measure how good they are. One of the areas that Hebia has actually been doing a lot of research has been in artificially extending the context window or effectively extending it. And that's been a massive actually boon leveraging inference time compute for accuracy for research and diligence workflows.
Mat
Since we're on the topic, I cannot resist asking for a double click on this. What does that mean? How do you artificially extend a context window?
George Sivulka
Well, the, the idea of a context window is that you can think about all of the things in the context window. So as, as humans we, we currently can remember whatever the last 10 minutes of conversation, maybe not this entire conversation, but we also have bits and pieces over our life and our training and our early careers that we've collected that make us really good employees. AI, you've got a system Prompts and then whatever, up to a million tokens to jam as much context as you can in there. When you think about what you actually want AI to do, you want it to reason over all of your data. Applications like Rag can jam the context window with as many search results as you want, but it's not going to reason over that data. It'll just search for stuff that exists by elongating the context window. Ideally you'd be able to connect the dots, find stuff that's there, but also stuff that's not there, stuff that's missing. And what that means actually tangibly from a product perspective is how do we very elegantly use the current context window, scale that up, or jam as much of a current maximum context window runs into a single question to get the right answer. And so it's actually more of an infrastructure problem, it's more of a kind of patchwork or convolution problem of how you actually apply that with the current length of context windows. And then it's actually kind of an information theory thing where you have to think through how what I get out of each of these runs can be maximized every time, but also maximally compressive so that every single future iteration it gets just the right information. Nothing that's not relevant.
Mat
As you mentioned a year or two ago you released Matrix, which is the flagship product for Hebia. So let's get into it. What does that actually do? I think it was billed as the interface to AGI. What's almost my reality as a user of the product, what do I have in front of me? What does it do for me if.
George Sivulka
The most important job in the future will be how you actually manage AI agents or how you prompt these things at scale. You can think of Matrix as actually running a bunch of sub agents or an interface like even a trello board where you can assign a lot of tasks and then a bunch of agents will do these things. And so it could be a network or a multi agent platform where you can basically say, hey, I want all of these things done. They all are related to each other in some specific way and every cell in what looks like a grid output ends up becoming a task that's completed by.
Mat
And you mentioned you started in finance and then you expanded to law and consulting. How do you customize something like this? And relatedly, how do you even decide which next vertical you need to go into?
George Sivulka
One of the things that is a massive fallacy in AI applications today is verticalization as paramount. All of the VCs a lot of entrepreneurs as well believe, because it's been true for the last 20 years, that building a very verticalized piece of software is the only answer. And so you've got plenty of startups that are like, okay, I'm only going to be AI for blank, AI for compliance, or AI for law, or AI for X, Y or Z. When you have a capabilities, like AI capabilities approach to the problem, you actually start to realize that generalization will beat specialization every single time. And what I mean by that is if you, let's say you want an investing agent, your investing agent will become a researching agent and the researching agent should become a learning agent. And then you get to something closer to AGI. If you want a law, a legal agent, right, your legal agent will become maybe like a diligence agent or like something that's like very hyper logical. And if it's a really logical, it'll be like basically the smartest possible agent and end up at AGI. Or if you're a banker, your banking agent will become a marketing agent because banking is effectively dressing up a company to be sold. And then the marketing agent will become, okay, the persuasion agent. And the persuasion agent will become something closer to AGI. And what I'm trying to highlight here is that it's a way to hack, like to prompt one or two things for a specific set of customers at this point in time today. But as AGI comes closer or is actually here, as these models improve and as we as people get better at using them, we're going to get way more interested in actually extending the context window or pulling these expert knobs or taking something that's general and writing our own prompts and building our own agents. Because the reason why you as an investor beat the market isn't because you're all using the same ChatGPT wrapper prompt. And the reason why you as a lawyer are getting paid $2,000 an hour isn't because, oh, I have the same shortcuts from AI for law company. It's because you are an expert and you can customize the greatest and the best tools to the way that your firm works, to the way that you actually can find an edge. And that's not specific. The best investors aren't only looking at SEC filings, they're like reading really crazy data sources that have nothing to do with finance. The best lawyers, whether they're persuasive or logical, they have backgrounds that are way outside of law. You don't want specialization, you want generalization. And a lot of our customers are realizing that that's actually the superpower of Hebbia and that's why so many people are buying us.
Mat
Don't seem to be a huge fan of chatbot as an interface.
George Sivulka
Nothing against chatbots.
Mat
Talk about how different the hebbier interface is and what you're trying to achieve there.
George Sivulka
I think chatbots are in my eyes like the TI84 or like the HP12C, like, like a calculator. It's a one off, you know, I'm going to go and put in my equation and then get a response. Nobody does their taxes in a calculator. Nobody computes a DCF or like, does serious knowledge work in a TI84? You like, stop using them. In high school you might have a simulator on your laptop to like because it's novel. The way that we work is we have these really broad, flexible interfaces called spreadsheets. And the way that everyone interfaces with a computer or the fundamental paradigm of computing is through a spreadsheet. It's flexible, it's highly modular. You could call it an expert system. And that's actually the way that humans, billions of people use capable computing today. The spreadsheet in Excel is the most important software to ever been invented. Hebia takes an approach that's not dissimilar. Just like a single cell in Excel is a calculator, a single cell in Hebia could be a chat, or it could be a variety of other AI operations. But the Matrix platform, how people use it is so much broader. It's so much more flexible. It's an expert platform. And we're building that expert platform for how humans will interface with AI agents.
Mat
And so the way they interact with the product and realizing it's difficult to talk about an interface verbally, but it's more.
George Sivulka
You could have invited me for a demo, I would have done a demo.
Mat
For you guys next time. But it's a spreadsheet, right? I mean it's talk about it one of the mods.
George Sivulka
So we think we're building effectively the new Microsoft Office suite. And just as Microsoft Office had Word and PowerPoint and Excel, we have Matrix and then actually a variety of other applications, some of which I can talk about, some of which I can't. And we have a Chatbot and we think that there's going to be an entirely new productivity suite for how humans, how experts want to interface with AI without having to get in the weeds in code. And it's one of those applications. Yeah, great.
Mat
All right. If that's okay. With you a little bit of a technical deep dive into how that whole works behind the scenes. So presumably to start with, you need to connect with a bunch of sources of enterprise data and process the data. So how does that stage of ingestion work?
George Sivulka
Ingestion and indexing is like one of the fundamental pieces of having any knowledge work application. There's different steps to it. So first is just like collecting the data sources and hooking into as many providers as possible. And that's not fun or not technical. The more interesting thing is actually once you have the data, how you index it, we believe that you can do a lot of pre processing. And I'm not talking about like keyword search and building a BM25 index or like having like you know, some sort of semantic search index with embeddings, even your super long embeddings that are coming out. That's not actually that interesting to our users. Again, that's good at doing searches, it's good at finding things in the data, but you want to start to process information before the user even asks a question. You want these agents to be doing work ahead of time. What we do and instead is we actually ingest documents and pre process, pre populate, depending on the doctype, depending on the context of the document, a really rich schema and understanding of each document that ideally would be, hey, we've pre indexed and pre done 90% of the work at the same time. That's not enough for asking and running user questions on the fly. And for the 10% of the work that is, okay, this war just broke out here or this crisis is happening here, or this event that you couldn't have predicted is happening. We use the same indexing engine and the same infrastructure to run things on the fly. And so you can consider any query that goes through Hebbier, that goes through matrix, as a combination of a lot of pre index work and then some stuff that's on the fly and custom and the pipeline to do that from parsing documents, sometimes using multimodal documents to figure out formatting and structure and images, actually how we then go and run that index and what we're retrieving over and how do we save the work of the pre indexing agents. All of those are interesting search problems that I could talk about for a while. That's a bit of the thesis behind how we index.
Mat
And then perhaps that's part of what you just described, but I wrote down something about what you call the ISD architecture. So that's the rag Part after that. And you were saying that somewhere that RAG sort of doesn't cut it for very hard problems. So what is your approach to rag? And I assume a lot of people here know, but maybe use the opportunity to define what RAG is in the first place.
George Sivulka
Sure. RAG is a architecture for using AI that's retrieval, augmented generation. It was first coined in a paper in March of 2020 by a bunch of Facebook researchers. Hebia were actually the first to turn that into a product. So it's like a very close thing to my heart. So back in 2020, we were the first people to actually productionize it, roll it out. And the idea is that you could hook a search engine up to an LLM. It's basically what Perplexity uses. It's like a search engine to LLM, except over the web you always have an answer. Over offline and unstructured and private documents, you don't always have an answer. And what we say, when we say RAG doesn't cut it or that heavy. Actually after, you know, making that as our baby, we had to kill the baby and we turned away to, from, from RAG to what we call ISD is we said, well, we need an architecture that's closer to extending the context window versus just searching or calling an external tool. And ISD leverages that index. So a lot of work that has been done ahead of time. But then also it leverages a way of recursively kind of reading sub documents. So still leveraging tools, sometimes leveraging that infinite effective context window to kind of bubble up an answer. And it's really a decomposition agent at its core. So you ask a question instead of it just pinging a search tool, it can ping a variety of different tools, does rich decomposition, shows its work in the matrix, and then synthesizes a complete response towards the end.
Mat
And do you have pieces still of I guess almost classical RAG architecture like RE rankers and that type of stuff.
George Sivulka
Somewhere or you completely very funny thing. And my first five employees will get a laugh of this and laugh out of this and no one else. But we have to this day the most accurate RE ranker that has ever been released. And we do not use it. We spent a lot like the first year, year and a half just training, embedding algorithms, training different versions of Colbert with like multi embedding architectures for a single passage, and then training RE rankers. And we came up with a novel RE ranker architecture which four years later, academia and the industry have not beat and we do not use it. And the Reason we don't use it is because a lot of the time search isn't what's important. Right. If we were building a public web API, ll like re rankers would be important. But we scrapped all of that for a really heavy infrastructure play that uses tons of large language model calls. It is really expensive from a latency and just latent dollar costs perspective, but achieves really accurate answers for deep research, for deep diligence tasks that people can only do in the heavy platform.
Mat
Great, let's talk about the model part. So you guys are very close to OpenAI. So do you use multiple OpenAI models? Do you use all this stuff that you can talk about, open source, what have you?
George Sivulka
Yeah, we've got partnerships with OpenAI but also anthropic and also Amazon. So we're playing the field. Don't tell Sam Altman, but the idea behind Hebbia is we believe the model layer will become commoditized. It's no longer a hot take. Uh, so whatever models you want to use and hopefully eventually computing on the edge will be the way that you run these things. Just like you don't ping the cloud to run your Excel model unless you're in Google sheets.
Mat
And then you have, you've built some very smart stuff. What I can tell around sort of scaling all of this. So you have this concept of maximizer which sounds like it's a router but smarter. What is it?
George Sivulka
Yeah, I, I think that some of the most interesting stuff at Hebia isn't actually only the agents research, but actually the cure systems problems where like running like I think we run around like 250 billion large language model calls a month and like our workflows. I think before we had maximizer, with all the rate limits that we had, we could run a million tokens a minute and now we can do 500 or 450 million tokens a minute. All of that actually wasn't an AI problem, it was a systems problem. And like you mentioned, it is a router. You can think of it as a two sided marketplace, but one of our systems there is called Maximizer. We liken it from an analogous perspective to an air traffic controller where just like an air traffic controller basically tells you who can fly where and when and apparently you're not supposed to fly out of Newark right now. I hope it sticks. Maximizer basically has a handshake between license grants and license requests. So if a user or a job has a really high or low priority and wants to go and use GPT4O. It can make a license request and license grant will say, okay, you have to use Azure or you have to use X, Y or Z models. And it actually ends up with the theoretic, the information theoretic, maximum utilization of any rate limits. So not only are we running the most amounts of pages, the most amount of tokens through our platform, but we're doing that in the most efficient, perfect way. And that's all our infrastructure, which is cool. I could get into more, but yeah, I don't want to bore everyone.
Mat
Inevitable question about hallucinations, you know, high stakes professional context where people are paid a lot of money to provide very reliable results to the customers. Is that something that, you know, in 2025 is as much of a problem as it was?
George Sivulka
I think it's old news and I think the only reason we still talk about, it's like everyone talks about hallucinations and, you know, no one knows if it's happening or what's going on. It's like fugazi. Fugazi. Like it was a problem back when the models were stupid, but they're obviously not stupid now. And I'd even say that they're way better than any human. So, you know, when we start to think about like hallucinations, it's like, well, where does that come up? Like, have we actually seen a lot of that happen? Probably not recently, but the place where it does come up is actually a limitation of rac. It's when you've got the wrong documents, where you're searching for something and it finds something that's broken. And what you're realizing is that the models are actually pretty good at doing reasoning. They're really good meta learners. Like the hallucination. Like whether it's a problem or not, like no one really cares. What people care about is when these things fail because they don't have the right context or they're not reading the documents in the right way. And that's why we've done a lot of research on, you know, how to feed the right stuff at the right time. And if that's really, really expensive, we don't care. Because models are going to become. Intelligence will become too cheap to meter. That's, that's what our industry is predicated upon. So yeah, we'll spend $10,000 in model costs per user a year. We don't care. We don't. But indicative.
Mat
You just mentioned research. How does that work at adapia? Do you guys have a concept of an AI lab within the company or is it distributed across the team. And how important is research, especially in a context where, again, there seems to be something new every other day. How do you think about doing your own research versus just free wr, whatever comes.
George Sivulka
Yeah, I. It feels like there's something new every day, but also like, if you've been tracking the field, it also doesn't like it. It kind of doesn't feel that different. I don't know if I'm allowed to say that, but it doesn't for me. Like, there's new models, but they don't feel that different. Like, you know, some of them are more sycophantic. Like, you know, that's the news. I, I think not a lot of people are saying it, but I think like, SF has a really big group think where it's like, okay, everyone's going to go out and work on the same things. And like, there's like this sort of like mimetic attraction rivalry between OpenAI and Anthropic and like, you know, what are they going to do next? And like, oh, they have a new interface or like the chat bar got a little bit jumpier with the animation. And, you know, we kind of know, like, okay, they can make videos and it's too slow and they can make images and it's better. Now the issue is to come up with new interfaces and AI that can really do new things. I think you kind of have to avoid the group think. And so we currently don't have any AI researchers hired out in San Francisco. We don't even listen a lot to what people are requesting and really try to think about, okay, what is the tool that AGI would use? Like, if you actually had AI that could choose any other AI tool to use, what would that look like? And we have that guide what the interfaces look like. That's kind of how we came up with Matrix and a lot of the other interfaces that we're talking about. That's how we came up with ISD versus RAG and what the rest of the industry uses. It's kind of by avoiding what is going on on Twitter and just keeping our heads down. So we do have a research lab. It is not a standard research lab. It does spend a lot of time on information retrieval and context renewals.
Mat
So do you think that, to take the extreme of what you said, that research is actually slowing down? So there's test time, compute that we were talking about, that sort of felt like the major innovation, you know, of the last six months. But after we're done with test time Compute. Are we sort of running out of tricks?
George Sivulka
I think there, there will be more tricks. I do believe that AI will start to work on itself. I think that, I don't mean to be relatively pessimistic, but I wouldn't start a company in AI right now. I mean that from the perspective of like the alpha is gone. Like you know, when I was starting the company there was, everyone was like working on crypto and it was like, you know, the alpha was gone. I was like, okay, we're going to work on AI. And so I don't know what it is, is like the, the next thing. But I, I do think that the models will continue to get better, they'll get way cheaper, they'll get way faster and the user experience will improve. I think longer term, just as I'm talking about how generalization beats specialization. You saw over the last 20 years of enterprise SaaS Excel get unbundled into a million specialization things. It only made Excel more powerful. But there will still be specialization as a second wave to create a lot of value. But I don't see lots of very large changes in the near future. But maybe I've also gotten spoiled from how much stuff has changed. So I'll have to think about that one. Great.
Mat
So everyone pivot to crypto right now. Announcing the rebranding of the event. Crypto driven Crypto.
George Sivulka
Crypto New York City.
Mat
Yeah, exactly.
George Sivulka
There we go.
Mat
I wouldn't come switching text a little bit and talking about go to market. You've innovated there as well. I think among other things you have a very different take on who salespeople should be for a company like Hebia in terms of backgrounds. Do you want to talk to that?
George Sivulka
Yeah, I'm happy to. I don't know if this audience is interested in like sales playbook and lineages and all that stuff, but yeah, I.
Mat
Think there's a fair amount of people building companies.
George Sivulka
So one of the things that I commonly think about is the best enterprise SaaS selling organizations over the last whatever, two, three decades all come from the same lineage or the same DNA. So they all come from like BMC and AppDynamics and now they went to MongoDB and they're all kind of like the same thing. They have the same playbook, they have a framework called Medic or Medpic which is all around how you sell value. AI is very different from those organizations products. If you actually try to apply like this standard, okay, like we're going to do this type of discovery and we're going to do this exact sort of medic sale. There's some things that we'll take and that you can extend to AI, but right now it's less around pain, it's less around, like standard enterprise SaaS, cycles, and probably more around FOMO, around missed upside, around value cases that are really hard to define. And so the profile of a traditional seller that would excel at these best in class quote unquote playbook organizations, you know, we'll see if that, if that takes for Hebbia and if that takes for our peers in B2B AI application spaces, the jury's still out. At the same time, one of the things that we've seen really work are people with domain expertise, people that can speak to the customer's lingo, vernacular, processes and work. And honestly, people that are very similar to consultants. And kind of having that playbook consulting muscle together is, I believe, what will work when there's a true paradigm shift over when, you know, the last two years of enterprise assets. It's really just unbundling of things that people already know exist.
Mat
So interesting, right? To that concept of like FOMO buying. Do people still ask you about roi, how to justify your price, and if so, how do you do it?
George Sivulka
Part of the beauty of Hebbia is that we, I think we're one of the only AI companies that builds value cases and that is really significantly upselling our customers. We treat it like a traditional playbook company. We actually build a really strong understanding of if you're spending this much, this is what you're getting in return, or these are the expenses that you can actually remove. And we commonly partner with the world's best financial firms, not to just give them an AI tool like our competitors, but to actually go out and understand and say, okay, it's a board level. You're trying to impact your PNL to the tune of $100 million. These are the ways that we've seen it done. And it's like, you know, a bit of technology and a bit of consulting. Yeah.
Mat
And is your ROI based on cost saving, meaning hours not spent doing a task, or is it based on increased revenue?
George Sivulka
There's cost saving. It's not only hours spent. Sometimes it's removing third party legal expenses or third party consulting expenses, or third party expert network spend. But a lot of what really is driving value and the amount capturing the imagination of customers is the idea of, hey, you can actually make way more money with AI. If I had an infinite number of employees that were experts at a task Maybe I could end up like really reading every single SEC filing and finding a bunch of red flags and finding a place where the markets really are inefficient. Or if I had a bunch of expert employees with an infinite amount of time, maybe my law firm could actually litigate a case better than any other law firm in the world. And that's much harder to capture. We can't defend it with an ROI calculation, but it's part of the reason people buy.
Mat
Do you come across actual cases where people will say, well, we're not going to hire X many junior consultants, analysts, bankers, because now we have this tool or similar kind of AI technology?
George Sivulka
I'll say one final story. I think we're running out of time, but I think it's a really interesting story and it kind of talks about the future and whether or not we'll have juniors. I think we will have juniors. I've seen some people that talk about it. I haven't seen a lot of people that do it. I've seen a lot of third party expenses like legal fees, I'm very short Accenture and consultancies, but. And the big four. But in terms of juniors, Morgan Stanley and some of the folks there always claim that they invented the analyst. And the story behind that is, you know, one day they got a bunch of computers at a computer room and they said, hey, you know, we don't know how to use these computers. All the bankers and they hired a bunch of kids that were like nerds from Columbia and they brought them down into the computer room and they said, hey, we're going to go and have you use the computers. And the person that made that decision came back to Morgan Stanley, whatever, 20, 30, 40 years later and said, hey, you still haven't figured out how to use computers. And I think the intuition or like the underlying sentiment there is when technology is created like the computer or like AI, and it's a true revolution, you end up actually having lots of people come in and do those jobs, like do jobs related to the technology. So I firmly believe that being an investment banking junior won't look the same as it looked five years ago. I definitely believe the same for investors, for lawyers, for everyone else. But I actually don't really think that that will decrease the amount of jobs. I think that someone like Morgan Stanley will claim to invent the prompt engineer and we can laugh about it 50 years later.
Mat
Actually, a couple of last questions from me and then I'll open it up to folks. Can I ask a personal question? Always so you are, in the grand scheme of things, incredibly young, and you're building an incredibly impressive company. Perhaps that's inspirational for anyone in the room that's thinking about building their own company. How do you do it? How do you. Maybe not in front of investors, because investors like very young founders, but you're going to go to the top people at some huge asset manager and whatever, and you're going to say, hey, I'm going to revolutionize your business. Like, how do you lead as a young founder?
George Sivulka
I think naivet is a superpower because it's incredibly hard and a terrible existence to be a founder. Everyone always says, like, oh, I wouldn't do it if I knew what it was like. And I think that's actually true. And so I think that young people end up changing the world because you just don't know how hard it will be. And then you're like, oh, I'm here. I don't have any other option. But I think it definitely demands a lot. I'm not the person that does everything at Hebia. I do a lot of things at Heavy, but we have a lot of talented later in their career folks that I learn from every day that I'm incredibly fortunate to work with and that I'm incredibly grateful that they believe in me, but, you know, still learning about sales playbooks and information retrieval and all the rest from. From people that I love working with. So I think that's. I think that's the superpower. Is the team always great and maybe.
Mat
Just to zoom out to finish next two to three years, whether that's product roadmap or what have you, what does success look like in three years from now?
George Sivulka
I think the only thing that I care about is again, like putting capable, truly capable AI in the hands of as many people as possible. I don't think that's a chatbot. And in three years, if the same amount of people that use chatbots today are using much more adept agentic applications and driving them to do things, to create value, to start businesses, to discover new information, that's what gets me going. That's like, that's my dream. And I really want to at least play a small part of that. So.
Mat
All right, on this note, George, this was fantastic. Thank you so much.
George Sivulka
Thank you guys for coming. Appreciate it.
Matt Turk
Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
Guest: George Sivulka, Founder & CEO of Hebbia
Date: May 29, 2025
Host: Matt Turck
This episode features an in-depth conversation with George Sivulka, founder and CEO of Hebbia, an AI platform designed to automate complex knowledge work for high-value professionals. The discussion revolves around the proliferation of AI “agent employees,” the future of workplace automation, and the evolution of AI-driven organizational design. Sivulka shares Hebbia’s journey from a niche financial tool to a generalist platform, delivering insights on product innovation, technical architecture, organizational change, and what it takes to build and scale in the competitive AI landscape.
| Timestamp | Quote | Speaker | |---|---|---| | [00:00] | “You’ll actually have hybrid AI and human employees working alongside each other. People that are really good at prompting be the best managers. Intelligence will become too cheap to meter.” | George Sivulka | | [03:05] | “...The goal of Hebbia was never to just stop at really highly paid knowledge professionals... our vision and mission have solidified... into building capable AI platform for a billion people.” | George Sivulka | | [04:52] | “This thing is a meta learner. It has beaten me to the research punch that I was working on.” | George Sivulka | | [07:03] | “With Hebbia, you can give it these complex tasks and it really turns through vast quantities of data and does work the way you work. It’s much more of an agent than a chatbot.” | George Sivulka | | [09:58] | “You’re going to start to have fully human organizations, actually fully AI organizations like the one person billion dollar startup… and this will be the most common thing: hybrid AI and human employees, agent and human employees working alongside each other.” | George Sivulka | | [12:54] | “Everyone will be prompting and prompting is managing and it will all blur pretty soon.” | George Sivulka | | [15:43] | “There's a lot of really interesting research direction in scaling laws for scaling during inference...” | George Sivulka | | [21:01] | “One of the things that is a massive fallacy in AI applications today is verticalization as paramount... generalization will beat specialization every single time.” | George Sivulka | | [24:08] | “Chatbots are in my eyes like the TI84… like a calculator. Nobody does their taxes in a calculator.” | George Sivulka | | [31:48] | “We believe the model layer will become commoditized… whatever models you want to use…” | George Sivulka | | [32:28] | “We liken [Maximizer] to an air traffic controller... information theoretic, maximum utilization of any rate limits.” | George Sivulka | | [34:26] | “I think [hallucinations are] old news... they’re way better than any human... intelligence will become too cheap to meter.” | George Sivulka | | [38:20] | “I wouldn’t start a company in AI right now... the alpha is gone.” | George Sivulka | | [42:56] | “A lot of what really is driving value... is the idea of, hey, you can actually make way more money with AI... it’s part of the reason people buy.” | George Sivulka | | [46:23] | “I think naiveté is a superpower because it’s incredibly hard and a terrible existence to be a founder. Everyone always says, like, oh, I wouldn’t do it if I knew what it was like. And I think that’s actually true.” | George Sivulka |
The episode offers a comprehensive look at how AI “agent employees” and generalist AI platforms are set to reshape not just repetitive work, but entire organizational paradigms. Through Hebbia’s story, Sivulka argues for horizontal, agentic AI—rejecting the chatbot status quo and verticalized solutions—in favor of infrastructure that enables expertise, customization, and massive scalability. For founders, operators, and anyone interested in AI’s workplace impact, this conversation is rich with technical insight, strategic perspective, and pragmatic predictions about where the industry is headed next.