
Loading summary
Hernan Lardiaz
He looked at me in the eye and said, ernald, I'm not going to implement AI because I'm afraid of the outcome. We are in a regulated market. If something happens, we cannot control risk.
Chris Staigle
So who is the person in the organization on the client side or the internal stakeholder that is most likely the person that will understand the concept, understand the requirements of it and be able to fulfill on the client side.
Hernan Lardiaz
There are two audiences, the business owner and the group that develops and implements AI solutions.
Chris Staigle
Well, anand any like, I guess closing thoughts or considerations for everybody as they go and chew on this.
Hernan Lardiaz
Don't feel shy about AI. Don't feel shy about the output of AI because it's good for everyone. It helps in a lot of ways, but ensure that you have the correct control points and the correct boundaries to understand what's going on with the AI output.
Chris Staigle
Hernan Lardiaz is a seasoned tech leader with 25 years in global sales and operations. A COO of RAG Metrics, he helps companies reduce risk and boost trust in generative AI systems. Welcome to Using AI at Work. I'm your host Chris Staigle. Each week we'll be learning how today's business owners, entrepreneurs and ambitious professionals are getting more done with smart use of tomorrow's tech. Let's get started. Right now, every business leader is asking the same question, what are we going to do about AI? If this is you, Chief, AIOfficer.com has the answer. We give you a simple path forward where we provide executive and team training so your people know exactly how to safely use generative AI in their day to day. We also manage the deployment and implementation to make sure tools actually get adopted and deliver results. And we'll also guide company wide transformation so AI becomes part of your operating system, not just another shiny object. The companies that act now will increase productivity, cut costs and grow faster than their competitors. Those that wait will get left behind. So if you want to make AI work in your business, visit chiefaiofficer.com and see how we're helping companies of all sizes finally get results from AI. All right, Greetings esteemed listeners. This is Chris. I'm the host of using AI at Work where we explore generative AIs application and participation in the work day through a non technical lens. Now the topic that we'll be covering today, we could get very technical in the weeds, but I'm going to make sure that the conversation sticks to how can I as an executive, first off, understand the application of and get benefit from model evaluations The. The testing of the results that we're getting from AI that we're putting into production. And our guest today is a fantastic thought leader on that topic. Hernan is a. He's been working with AI since 1991, was his first exposure as a software development engineer. Has worked with a lot of the known brands, multinational brands that you know of. But we were recently introduced through a post that he and his partner had put in Rachel Wood's group called the AI Exchange. If you're not familiar with that, definitely worth checking out. And I know Rachel's a friend, but she comes from kind of a geekier background when it comes to AI. And in all of the training that she does, she always stresses evals. Evals. Ev. Evals. Well, as somebody who didn't come from a technical background or a development background where that was important, that was kind of a. A step maybe that I. I skipped or didn't pay a lot of attention to until recently. Now we find ourselves at chief AI officer doing a lot of work with clients now where it's not. If there's a mistake, it's not internal to our company. If there's a mistake, it has, it has impact, it has gravitas on a client company, which we definitely don't want to do. So when I had the opportunity to connect with Hernan and his business partner on this topic so that I could understand it, I was excited to share that information with you. And this interview will be a continuation of a basic conversation that Hernan and I had maybe a couple of weeks ago. So with that, Hernan, welcome to the show. And anything that I left out, maybe that you want, you think people should.
Hernan Lardiaz
Understand now, first of all, thank you for having me, Chris. A pleasure talking to you today. You define a little bit what I've been doing. So my first interaction with AI was in 1991. At the time was purely academic because there was no processing power. The main difference between now and then is the processing power that we have today. So for sure, in the middle of my career, I've been doing other things, but in the last few years came back to the AI world. Nice.
Chris Staigle
And it's a new habit that I've started at the beginning of the shows. I'm borrowing from Greg Eisenberg and his show, but I think it's a good question to ask by the end of this episode, considering our audience and the preamble that I gave, what would you want the listeners to walk away with?
Hernan Lardiaz
Why? Having a measure to control or to look at the output of AI, it's important and we for sure will get into the details of that. But why do we need to look at what we are getting from AI, analyze it and have some kind of statistic behind it?
Chris Staigle
I think it's a fantastic target. Okay, so let's jump into it. We had a brief conversation where I was. What started our initial dialog was that I was interested in exploring your services to bring that into what we're doing for clients so that our chief AI officers had some tools and frameworks to be able to evaluate. Hey, is the solution that we're building for this client, is it tenable? Is it, is it stable? Is it going to give them a high degree of accuracy when it comes to the output of that intelligence? And that really, I think that's a great place for us to start. So can you kind of explain to the audience what, what does evals mean in the context of using AI?
Hernan Lardiaz
AI is based on a pro, a probabilistic feature. So in order to have an outcome, there's some probabilistic stochastic process that is embedded so that drives to probabilistic outcomes. So as I always say, if you play the lottery enough times, at some point in time you'll win it.
Chris Staigle
I think this is good. Before we go on, explain the difference.
Hernan Lardiaz
In your.
Chris Staigle
Position on deterministic versus probabilistic. Probabilistic, yes.
Hernan Lardiaz
Yeah. And in the old days of. In the old days, actually up to a few years ago, that's the old days of computers, computers program were deterministic. So there are a series of algorithms that produce a result. So if you put 1 plus 2, you will always get 3 in AI, it's probabilistic. So the outcome is not necessarily always the same one. And that's basically the fun of AI. So that also drives that sometimes the outcome is not correct or is what we call. It's hallucinating. And the nature of AI hallucination cannot be avoided. Errors cannot be avoided. And that's basically the core of what we want to resolve with what we do at Ragmetrics. Try to understand what is that output of AI and how to measure and evaluate it.
Chris Staigle
Okay, so as a listener, I'm sure that you have used, you've got maybe a few prompts that you use regularly. You pick them up at maybe a meeting or you develop them yourselves. And you've noticed that I use the same prompt, but the output is different. Maybe not wildly different every time, but it's not the exact same every Time. And that is a key concept as we continue this conversation. If I have an activity or a workflow or a process in my business that I want to accelerate the result by using AI, I need to be prepared for the fact that if it's a very precise activity, if I don't have frameworks in place, I can't always count on the result being meeting that, that, that tolerance of precision required for that workflow output or whatever. Hope this is making sense, but it's very important because outside of if you just want to use AI for thought leadership and ideation and working on strategy, that's okay. But if you want to operationalize AI and make sure that it's something that becomes part of your workflows and processes that you can count on and you can say, hey, you know what? We know we can now scale this activity in our business by leveraging AI and we're going to get those consistent results, then this effort on evaluation is something that's extremely important. So I just, I want you guys to know that this is not just some theoretical topic. This is a key piece of being able to operationalize generative AI at scale. Okay, so Hernan, thank you for that definition. What is the experience that you're having with clients who are the type of clients that are looking for, that are even savvy enough to know that they need this as part of their process and that are looking for help solutions like what you guys provide?
Hernan Lardiaz
Yeah. The market is continuously evolving. So we are on the phase where more and more companies and enterpr are starting to implement AI, from large enterprises to small SMBs starting to implement some kind of AI. So anyone that has any risk with customers or customers could be external or internal. Any risk with customers should and is using AI should have an option to control that from simple chatbots, implementations that give you directions for a product, how to set it up, to fintechs, companies that provide information about financials, to healthcare companies, to anything in regulated markets. So everyone that has an AI tool that requires, as you defined it before, very clear, precision requires control on that. Yep.
Chris Staigle
Okay, so for those of you listening some of the activities in your business as you're introducing AI, you may be fine just saying, hey, for this role, I'm going to give you a license to ChatGPT, some basic governance guardrails, a little bit of training, and that's all you may need in this particular role. For other roles, you know that boy, if we could have the bot, the agent, the app that could support this role, that Individual in that role could do 10x or 100x productivity. But when you are outsourcing or relying on the models to have that much of an impact on that processor workflow or, or job in your business, the reliance on consistency is paramount. There's, there's no two ways around it or else you will scale it, it'll break and then you'll go back to the drawing board and you'll say, hey, this, this AI stuff, it's too hard. Let's just keep with prompting. So Ernan, for businesses who are like this is a new concept to them. How do they, how would you suggest they start to identify those activities where. No question, we need to have an evaluation process in place before we put this into production.
Hernan Lardiaz
Yeah. So let's, let's, let's recap a little bit. So the evolution process goes into, into in two areas if you want. One is to ensure that the output is correct in terms of what general correctness means. It's not hallucinating. The answer is more or less correct. The other one is what is what you really want to measure. So different businesses have different needs and those needs rely on the quality specific of what they are trying to implement. So as an example, a retail entity, a retail chatbot, might be focusing or not on discounts and users might want to measure the quality of those discounts. So it's not only just providing a general answer that is correct in form and in context, but also looking at the details behind it. The process to generate a good AI product and to monitor your a good AR product starts on the development phase. The development phase should include testing and optimization. So without getting too technical, if you have a rag system, there are different ways to get information out of the rag system, different policies. So the optimization of those policies can help to reduce cost, can help to get better throughput, and can help to get more better accuracy on the answers. So it is important when the AI solution is being developed to have a good evaluation process or tooling to evaluate during there then the other portion comes is when the AI solution, it's alive, it's up and running. So we heard in the market concepts about hallucination, concepts about drift. The objective is how to detect those, how to detect if something is drifting. An AI solution is drifting. So one way to do it is to continuously monitor the output of that AI tool in time. And if you are measuring accuracy and you start to see that the accuracy slowly starts to degrade, then you can see that you're starting to have a problem. So accuracy, let's say in a scale from 1 to 5, might be, I don't know, 4.9, good, 4.9 next week, 4.9. And then started to come 4.8, 4.75, 4.73. Then you start to see that is degradation. So having the tools in place is important for that to start. It's basically having the commitment to understanding that there's something that needs to be done. The tools in itself that are out there like ours are easy to implement, are easy to monitor, are easy to handle. So that's not the issue. I think it's more important to have the concept, the idea that monitoring is required and evaluations are required in order to get a better, better output of AI.
Chris Staigle
And you know, spoiler alert for the listeners. When I was interested in this for our, you know, our own business and when I inquired about the pricing, I was overwhelmed with how generous their pricing is. So just so you guys know, this is something that you need to be doing in your business for any, obviously some activities are more important than others. It's doubtful that many of you have. And because, listen, the reason I'm saying this is I'm inside of a lot of businesses as an advisor, as a, you know, a leader of bringing AI into their business. I doubt that many of you have any type of evaluation, whether manual or, you know, computer assisted, for the AI activities that are occurring in your business. And the main thing I want you to be thinking about right now is where are we at risk? Because the business owners that we talk to who are interested in, hey, we've got to do something about AI, we're ready to go, but we don't understand the risks enough to go all in. If that's you right, then I've just exposed you to a whole nother vector of risk that you probably hadn't even thought about yet. But the good news is that now that you understand this and now that you're becoming aware of solutions like ragmetrics, it's not going to be a big deal for you to confidently operationalize AI into a lot more areas, even areas that are, you know, there's compliance involved, there's, you know, high, high penalties for any type of mistakes and things like that. So the good news is the risk is being mitigated on this call. So hang on, what we've been doing is manually, like human map mapping this on spreadsheet. And it's better than nothing, right? Version one is better than version none, however, takes a lot of time from both the chief AI officer and the pilot teams. It is open to human error. And all the things that, you know are the promise of why you'd want AI in the first place. How does a company transition from not doing this at all and being like, oh, wow, we didn't even think about that. To maybe those that are randomly checking outputs and grading them on a, on a, whatever their, their scale is and tracking it in a spreadsheet to going into an environment where is it automated? I mean, what does that progression look like with the final result being your solution?
Hernan Lardiaz
Yeah, so yeah, for sure, the objective is to be automated. So there are three types of projects or companies working on AI and how they address this issue. There are three different. The ones that do nothing and basically say, no AI is good enough, let's continue with that. And then the problems comes and there are several stories in the market about that. The ones that test manually, that I think from the ones that are testing, the few that test something, they test manually, and then the automated process like ours. So the process is very, is very easy in concept because there are two ways, as we said it before, before the product is in production and after the product is in production, before the product is in production, the process is very easy. It's getting the knowledge base, getting the information that should be used to train the AI tool to understand how the AI tool should behave, getting that information and creating test samples, we call them data sets, test samples, and inject those test samples through the pipeline, the AI pipeline, and see what comes on the other side and basically create that. And that process is basically just without getting too technical, just injecting information to an API and getting the response from the API. And then we look and we evaluate. And as I said before, there are hundreds of metrics or criterias that we can use to evaluate. So the process itself, from the technical point of view, is extremely easy, similar to what we call live AI evaluations or monitoring. That is, we get the information through an API of the context, the answer and the question, or the input and output that that AI agent AI component had. And we look at that through the same metrics or different metrics or the metrics that are required or need in that specific point. So the process on the technical point of view is very simp, as I said before, and I think I'm becoming repetitive. That comes with age and becoming repetitive. The importance here is to define what needs to be tested and how it's going to be tested. And Chris, there are hundreds of examples and probably we should have mentioned those at the beginnings from Air Canada being sued because a chatbot didn't provide right information about a flight and tickets to a fast food company getting an order wrong and putting hundreds of thousands of french fries in the same ticket to New York City government providing information to someone to break the law. So they're hundreds and thousands of examples that can happen any day. The most important thing is, are we ready to start testing? Are we ready to start the evaluation? And it's in the leaders of the organization to drive that forward.
Chris Staigle
So, okay, I'm going to run through this so that I can extract information from you and apply it to the client experience that we're giving for our people.
Hernan Lardiaz
Right.
Chris Staigle
Okay. So who is the person in the organization that I need to. On the client side or the internal stakeholder that is most likely the person that will understand the concept, understand the requirements of it and be able to fulfill on the, on the client side on whatever those obligations are for this.
Hernan Lardiaz
Yes. So there are two. Two audiences. So one of the reasons why I started with, with this project is I had a conversation with the CIO of a bank in the Midwest customer for a long time in other business. So I went to see him to see if he wanted to implement some chatbots and voice bots to handle customer service. He looked at me in the eye and said, hernan, I'm not going to implement it. I'm not going to implement AI because I'm afraid of the outcome. We are in a regulated market. If something happens, we cannot control risk. We cannot manage that. So that's how I started here. So that is one audience. One audience is the business owner. And I mean business owner in the broad sense could be division manager, ownership mentality. Yeah, yeah. The business owner that wants to have to reduce risk and the risk profile of the services that he's given, either internal or external. So that's one group. The second group is on the IT side or on the development side. That is people that are every day working with AI tools and need to develop and need to accelerate the process. And in order to bring a product with quality, they need to test and evaluate. So those are the two audiences. So the business owner and the group that develops and implements AI solutions, AI agents in general.
Chris Staigle
So now I've got my internal representatives for this effort on any AI deployment. Next my question would be, who's going to develop the tests?
Hernan Lardiaz
So the tests are very easy to develop and we can, we can help our customer or we help our customers to develop that. The Test is very easy to develop, and I sort of gave an indication before. Yeah. And basically it has three components. One is what we call the data sets. And basically our tool creates data sets based on the knowledge base. Let's give an example. Assume that you want to test a chatbot for retail that deals with retail policies for refunds, exchange or whatever. So there's a policy that is written probably in a PDF document. So, and probably that same policy is the one that has been loaded into the chatbot to provide the answers from. So we get that policy and we create the test scenarios based on that policy. It's as simple as uploading the policy into our tool, our system, and asking the system to generate hundreds of questions based on that policy. So that is phase one. Phase two is to define the criteria that you're going to use. What do you want to measure? Accuracy. Do you want to measure the thickness? Do you want to look at how many tokens it's used? Do you want to measure the quality of the output, the tone? There's several. So that's the second phase. What is what you want? Which is the criteria. What is what you want to measure? And then the first, the third phase is running the tests and looking at the results. And as I said before, in general, the tests are run through an API or the information can be uploaded to our system, our system can run it offline. For that, we also need the prompts, we also need which are the models that you're going to be using and the basic things. But yeah, that's basically how it runs. Once you have the results of the test, you can look and define, refine your prompts or change the model if you want to try another model to see what it's better and things like that. So you have all the tools there to optimize your platform.
Chris Staigle
So as a listener, I want you to think about this. You could you do all those steps manually? Could you go to the models and say, hey, here's our situation. Can you create, you know, these test data?
Hernan Lardiaz
Sure.
Chris Staigle
Could you, as a human, go in there and copy and paste and hit Enter, and again and again and again and again? You sure could. However, before you got exhausted, maybe you've run through X number of scenarios how. But that's not probably going to touch all of the edge cases that'll come up in the wild with humans interacting with your AI system. But when you have a technology like this, and I would assume, Ernan, that one of the main benefits of the API integration is that I can do high speed testing very quickly. A lot of scenarios.
Hernan Lardiaz
That's the objective for sure. You can always do it manually. Yeah, but basically instead of taking a minute per question, two minutes per questions, we can do hundreds of questions almost simultaneously. And you put them there and you look at the answers and you get the statistics on the back end so you can look first. Overall statistics. The overall of this is as we said, 4.9, so it's good enough. The other set that we tested is 4.92. So let's look why this one is better. So you work on the real things, that is optimization and getting the things up and running faster and better, not just spending time in creating. The other thing is whatever the model is, you are ensuring that you are having consistency in how it's evaluated. So if you continue evaluations on, if you have people, people may look at them differently. If I analyze all the questions, I might, yes, have the same rationale. But if I give some to you, you might say that something is a four when I thought it's a three or a five and there might be differences. So here there's consistency throughout the process. So if you have an update on the product and in two weeks you come and change a few things on the product, then you can test with the same data sets. The same data set and look at how that changed in your output. It got worse. So the update is not what we needed. It got better. There we go. So that's the idea behind it.
Chris Staigle
Yeah, that's very helpful. And I mean, I guess to give it an analogy, I've got to be in San Diego on Saturday. I could walk to San Diego, which would be the equivalent of me manually doing these evals. Or of course I can jump on a delta flight and which would be the equivalent of leveraging technology to get the result faster. Okay, so, okay, great, I've got this. We've, we've defined it with the stakeholders internally. We're starting the testing. So I would imagine that the, the evaluation of that though you still want some human in the loop to maybe do a sampling of comparison to what the system says is accurate output and the human saying, hey, well wait a minute, I do this every single day. I respond to these tickets, I answer these questions, whatever that is. I don't think that's the same. Where is human in the loop on further evaluating what the evaluation system evaluated?
Hernan Lardiaz
So we see human in the loop in a very minimal aspect. Not because we want to minimize humans, that's not the idea. But in a very Minimal aspect. And it's to retrofit information back into the system in the moments where it cannot be done automatically. And then we can talk a little bit more how to correct a few things automatically, but in a model that cannot be automatically. So you look at numbers, you decide that there's something needs to be improved, a prompt or things like that. In some cases need a human to do those things. Our system, we tested, we did some testings and our system is 95%, has 95% human agreement. That means that 95% of the time a person on the system will rate in the same way. So we believe it's pretty good and pretty consistent in that front. So we try to minimize the use of people only for those cases where it's extremely needed. The information that someone that does it every day and how it's done or things like that, that should be part of the information that's loaded into the AI tool, the AI agent, and work on that front.
Chris Staigle
How are you guys capturing that? Is that sitting with the process owner and documenting, you know, how they do a certain process and what they look for, or how are you capturing that information?
Hernan Lardiaz
Basically, we don't develop the agent, so that's part of the process that the agent development has to do. We look at what we call the knowledge base. So that information should be in the knowledge base? Yeah, a knowledge base in the broad sense, if you want, but that information should be in the knowledge base. So if a knowledge base, as an example could be an faq, a frequently asked question. So that means that if a customer says I want to return this, the first thing that the person asks is when did you buy it? Or things like that. So that is part of the knowledge base and that is uploaded into the evaluation process. Okay.
Chris Staigle
And again, through the lens of our practice of doing this for companies, I'm asking these questions, but I think they're applicable to a business owner so that they can understand when they need to plug this type of thing in. What, what scope of a, of a project would you say would be too small for really? Oh, you don't really need evals because the human's going to be involved anyway as compared to something where it's obvious. As we mentioned, there's no way that my team can test nearly as many scenarios as the tech.
Hernan Lardiaz
So there are two, as we spoke, there are two phases, one pre production and one post production. I would say that the post production area is any. So there's no meaning. As an example, we develop a node for N8N and non code tool for AI agents that that's basically for we'll say for any segment but starting with a low tier as soon as an AI approved it's going to be available. But basically we are targeting any segment and that's the objective of that is what we call the monitoring piece of the live AI evaluations. So that has no minimum or nothing is too small from there from the rest is what you as a developer or implement someone that implements an AI agent, what you define that your need is. So I can give you a very simple example. I develop internally Chatbot for a demo internal and before using it I run it through some evaluations to measure the accuracy and also to measure the consumption. Because we are startup we don't have too many funds. So yeah, we measure the consumption. So we realized that some of the things we were doing were consuming more than we should. And I can give you an exact example and explain how it went. But so and basically this was just for a demo internally. Then I ran a 20 minute evaluation and I decided to change the way that the rag was retrieving information. And it was at that basically the new way used half of the tokens that the previous way. So it's not only accuracy, it's also to optimize on the cost side. So yeah, it's, it's whatever you want. And basically this was with 25 questions, it's not hundreds. It was with 25 questions just to have a sample of the consumption. So there's no real limit in what you want to do.
Chris Staigle
Okay, the evaluation process, if I'm already in action with some pilot projects and at least with chief AI officer, we kind of have a three phase, we model the process then we kind of do like a build and assessment and then once we feel it's hit an accuracy level and then we roll it out into production but monitored rollout over 30, 60 days. So in this environment that would fall into step two for us where we're actually we built the test thing and we're evaluating. Let's say we run it through there and you mentioned a 95% accuracy target. That's an internal target that we also use when we're working with clients. But again we're doing it manually. How long? Provided that it hits that 95% accuracy in the, in our assess phase. How long is that that evaluation process going to take? Minutes, a few days?
Hernan Lardiaz
Minutes or okay short hours depending how many, how many samples you want, but couple of hours. So it's in the day. Range. So it's a few hours during the day. And that's, that's basically it. Not. It shouldn't take. It shouldn't take. It shouldn't take more. More than that in my mind, Chris, what it's, what it's more important is not only doing an evaluation once, but finding a way to continuously do evaluations either by monitoring in real time or every couple of days, inject a set of a data set to evaluate and see the output and see how that is changing in time or not. Because again, we go back to drift, we go back to hallucinations, we go back to errors and that is what we want to prevent. So ensure that there's something always there that is being, it's driving evaluations.
Chris Staigle
So for me, I'm seeing a constraint in our model because here's our option now if we're doing it manually, put it in a checklist and give it to the client and hope, fingers crossed that they're like, oh yeah, it's working fine, I don't need to test it. Right. Let's hope that they don't do that. Or a chief AI officer, highly compensated individual with, you know, specialized talent is going in there and doing some basic.
Hernan Lardiaz
Right.
Chris Staigle
Is this accurate?
Hernan Lardiaz
Right.
Chris Staigle
So it's not a high leverage activity for that stakeholder in the company or an external chief AI officer. Could this all be set up to where there was an automatic trigger to run a slight eval every 72 hours or whatever and then report back to. Like that's all automatable?
Hernan Lardiaz
Okay, yeah. Basically you have the data sets, you have the experiment that has the parameters of what you want to evaluate. Inject the data set through an API, get the results on the other side and look at the numbers. So at night, every other day, you can run it, you can test it. And I'm saying night. Assuming that that is, there's less when it's running. Sure, yeah, yeah. Or the other way, it's having the real time evaluation, evaluation process continuously running. So those are the two ways to do it.
Chris Staigle
Depending on the, I don't know, the process itself, some might be once a week, once every three days, as compared to some are doing it multiple times per day. Because there's such a, you know, like we have to make sure this is right every time. Okay, this is very helpful. And actually it's, this is one of the things I love about generative AI1, there's so many layers to it.
Hernan Lardiaz
Right.
Chris Staigle
So we've got listeners to this show that they're, they're like, you know what? I'm getting good at prompting. I'm starting, right? And they think that they have, like. And again, guys, if you're listening and this is where you are, there's no slight to this. This continues to happen to me, somebody who's immersed in this for, you know, past several years, every single day. But you, you, you do something and then you're like, oh, wait a minute. That's just like the basics. There's this next level, and then you kind of get in there, and then, oh, my gosh, there's another level. And that's kind of what I'm experiencing right now with this understanding of. And now that I'm seeing it, it's like, you know, oh, of course. But until you get exposed to an idea, you don't. It's. It's not like, oh, yeah, of course that makes sense. You just, you don't even know that that exists. And that's what's happening here. And I hope for the listeners that you're seeing the wow. Oh, this eval concept. Hadn't even thought about that before. And yeah, I mean, I'm a smart guy. I work with a lot of businesses. It's not something that I. Until I started talking to her, not in his business partner a couple weeks ago, I didn't realize the impact of this. We were doing it manually. We were trying to hit that 95% accuracy before we. We were using the clients pilot teams to evaluate it and all those sorts of things. And now I'm realizing another layer to this thing. This is so much smarter, better, more accurate, cheaper, faster, everything. So, yeah, this is good.
Hernan Lardiaz
The example that you gave on the guy that is good at prompting, and he knows. So my question to those guys in general is, okay, you change five words on that prompt, right? How do you know it's better? How do you know it's better than before? Because you run 10 tests and you say, yeah, it's better because these 10. Because that is what. Why you change it for. Because you wanted that outcome. But what happens with the other 400,000 scenarios that you haven't test? So the idea of having automated testing is to resolve that in 10 minutes. You know, yeah, this is better because these, that and the other. And you have confidence on. On. On that. So, yeah, so, yeah, so.
Chris Staigle
And just again, through the lens of the practice that we have, what does spinning this up in an environment look like? Like, what's the onboarding for me as a client who would be interested in.
Hernan Lardiaz
Testing the tools As I said, if it is pre production testing, the onboarding is extremely simple because it's basically using an API and understanding how it works. So in general, we work with customers on that front. It's very simple, it's no big deal. The live hallucination has two components, one that we haven't mentioned yet, but I'm going to mention it now because that adds a little bit more tricky things, let's put it that way. One component is just getting the monitoring piece. So basically every interaction that goes or every output that goes from an AI agent is being evaluated and you get an answer. The other component on that is how to retrofit that evaluation. Then what? I mean, evaluation is a score and a reason why the score is what it is or why that evaluation failed or whatever. So how to retrofit that into the AI agent to force the AI agent to correct itself. That has a little bit more components because that requires a modification of the prompt and other things, but still it's extremely easy and friendly. It's all a couple of APIs that go from one side to the other. So if someone wants to try the tool now and has a little bit of knowledge on the technology side, I'm pretty sure that in an hour he can be using it without any problem.
Chris Staigle
So my takeaway from that answer is that it's way more important that we introduce this early in the process as compared to, oh, we'll figure it out once it goes live. Like again, seems obvious when I say it, but.
Hernan Lardiaz
Let me give you a real example that I was mentioning about that chatbot. So we're working in a specific demo for a retail customer, so got the retail policy uploaded into a rack, a pine cone rack, nothing too far fancy or something very, very simple. So I said, okay, let's get five chunks. So the top, the top K is five chunks and let's start to work from there. Nothing too crazy. Did the chatbot in Python streamlit nothing too fancy, something very simple, very normal. And I said this is work, it's fine. What happens if between. If instead of five chunks we get two chunks, so top K for the ones that now it's two instead of five, that means that basically we are getting less data from the rack to evaluate. So the first thing is, if we look only at the output, probably the accuracy, it's better in the two because it's more precise. As you know, the more data that you get from the racing, the more accuracy you lose because basically you get what is more accurate first and then you come down the line, fine, good. But it was interesting when we tested, the generation piece also was more accurate when we talk two chunks instead of five. And that's because simple rag, simple questions, simple chatbot. In two chunks, it has enough data to provide the correct answer. So the other three chunks were adding noise. So it was more accurate with top K2. So basically two chunks. And additionally when you were using all. When we are using two chunks, we have less data that is transferred to the LLM to generate the response, so less tokens. So it ended to be around half the consumption of tokens. So those are the things that if you are way down the line, you don't necessarily are coming back to test. So the earlier that you can start testing and evaluating, the more that you can drive accuracy to your system.
Chris Staigle
Yeah, that makes perfect sense. Like, I get it. This is helpful. This is now like my mind is spinning because I'm thinking about all this training. I need to go not necessarily rework, but enhance with the concept of this new, this new approach which would be leveraging technology to do high velocity, high volume evals as compared to what I've been doing, which is horse and buggy. Right. Like the old way of got a spreadsheet, we're doing some tests. This is fantastic. So it's a huge win for not just my, my business, but the efficiency that will, you know, any client will be able to experience from this eval and a confidence on our side. Or, and if like folks, if you're listening to this and you're the person leading this internally, or you're working with an external or an internal person who's leading this, it's important for you to know this. And I would actually take this back to them and say, hey, what's, what's our, our protocol for evals on our processes right now that you have a good understanding of what it is and why it's important, I would be curious, are your people doing this right? And if not. And Hernan, now what I'd like to do is kind of shift to where, if people want to find out more, where they can go. But let me ask you first, you guys don't run the evals. You have the technology that supports the running of the evals.
Hernan Lardiaz
Yes, for sure. If someone needs help, we will provide it. There's no, no doubt. But we don't run the evals. The idea is that the people implementing, the people monitoring have their own implementation of the evaluation process and they run it and they run it whenever they want. However they want. They can change metrics, they can look if there's some issue, I don't know, they detect that these counts are coming too often. So they can create an evolution process or monitoring process to look at the discounts in that chatbot, in that agent. So yes, we allow people to do it. So, okay.
Chris Staigle
One of the things that I would be thinking about if I was listening to this was that's great, but my people don't know how to do that. Where can teams like pilot teams or implementation teams learn how to set up the evaluation?
Hernan Lardiaz
We have a very. Yeah, we believe we have very, very good information in our website, Ragmetrics AI. There's a resources area and there's the documentation area. The resources area more tells the story on how to run evals, what is important, different type of evals. So I think there's a lot of information there. Great. For sure. For sure. They can contact us through the website. And we are more than willing to help people. So we believe that this is new in the market. It's a new category that is going to be defined. We are more than willing to help anyone that has a question, anyone that has a comment, doesn't know how to do it, more than willing to jump on the phone and give examples, even provide code of things that we've done internally in order to help them.
Chris Staigle
Okay, now I'm preparing to roll out AI or I'm talking to my board next week and we've going to come together with thoughts. I've just listened to this podcast episode. I know they're going to ask about budget, let's say middle market company. We're going to start with maybe one or two pilots across finance, hr, operations, sales and marketing. What's my budget? Need to go ahead.
Hernan Lardiaz
So basically the unit of measure is the evaluations. So our package that we are launching now, it's around $250 starting from there that provides around 3,000 evaluations. So it's a good number to start in the process. From there it goes up and for sure, if someone has a specific need, we are more than willing to address it. But it's not that you need to sell the house in order to get an evaluation process. That's not the idea. And the more that this runs, the cost per evaluation will go down and. Yeah, and it's going to make it easier for everyone.
Chris Staigle
Yeah. I'll tell you, when we first had our conversation and once I realized a. The impact, but also the technical nature of it, I was expecting a pretty big. I was expecting that this to be inaccessible price wise, except for companies that were like working with Deloitte or whatever.
Hernan Lardiaz
Right.
Chris Staigle
I didn't realize how affordable that was. And when you work out the math, 3,000 evaluations, let's say we said, you know what, we're going to save the 250. We're going to have our people do it for your people to run 3,000 evaluations. The payroll on that would be way beyond the 250. Pony up the 250, do it right, remove the human error from it and let the tech handle it for you in a day as compared to a couple of weeks. So that's my position.
Hernan Lardiaz
And again, more than willing to. So for sure we have models that if you're a large company and want to have something in house due to privacy, we can host it in house on prem. So we have different models. But the idea here is to help the AI developers understand what they need to do. And also to your question about the board, I thought you were going to go to the phase. Okay. And which is the risk of implementing AI. Why do we want to implement AI? And I think the big part of the risk of implementing AI is that AI hallucinates AI brings error that jeopardize cost, that jeopardize customer retention, loyalty. So by having a tool that evaluates and monitors in real time, you are mitigating all those risks. I think that in my mind that's the most important thing. Yeah.
Chris Staigle
And it's a low cost insurance to make sure that all this energy you're putting into introducing AI into your business, that it's working and that it's something.
Hernan Lardiaz
You can count on.
Chris Staigle
So like this has been fantastic for me because now I feel like I've got a whole different level of comprehension so that I can now explain this to other businesses. And obviously the stuff that we're doing internally at chief AI Officer on why it's important and why that like you, you don't want to not do this, right?
Hernan Lardiaz
Yeah.
Chris Staigle
You don't want to put AI into production and just hope that, you know, your employees are, you know, spot checking it on gut and getting it right.
Hernan Lardiaz
No.
Chris Staigle
You got a business at stake here.
Hernan Lardiaz
So.
Chris Staigle
Well, any, any like, I guess closing thoughts or considerations for everybody as they go and chew on this?
Hernan Lardiaz
No. Just don't feel shy about AI. Don't feel shy about the output of AI because it's good for everyone. It helps in a lot of ways. But ensure that you have the correct control points and the correct boundaries. To understand what's going on with the AI output. That's basically it.
Chris Staigle
I think that's fantastic advice. So we're going to have links to Ragmetrics AI. We're going to have links to your social connections and all that sort of thing. And I want to encourage you, listener, dear listener, regardless of where you are in the phase, if you're leading, if you're investigating, if you're somebody who is at the top of the organization, if you're somebody who's kind of peripheral to what's happening, if this isn't being discussed, I would say, hey, like, send them the episode or whatever. But, but I would, I would say this is certainly a consideration that we need to do. And I think I've got a resource that we can start with. Right. So. Awesome. Well, I look forward to staying in touch about this subject as you guys continue to make breakthroughs in the startup environment. And I look forward to actually engaging with Ragmetrics as a client for our clients to make sure that we're, you know, we're doing evals in a way that I feel very comfortable and confident with.
Hernan Lardiaz
Yeah, Chris, thank you very much for the opportunity, sharing this with you on your audience and yeah, looking forward to continue conversations in all fronts. Thank you very much.
Chris Staigle
Thanks everybody. We'll catch you on the next episode. In the meantime, just keep using AI. Thanks for tuning in to Using AI at Work. Don't forget to subscribe for more conversations about how to use AI at work and a special thank you to our sponsor, Chief AI Officer for Empowering Businesses with AI Education and Training. Visit their website for a free AI Readiness Assessment and AI Strategy Guide to help you get started using AI at work. That's www.chiefai office. Follow us on Twitter at the handle usingaiatwork and visit www.usingaiatwork.com for free resources to help you harness AI in your role.
Episode 90: Evals and AI Output Evaluation with Hernan Lardiaz (PodEdit - D1)
Date: February 9, 2026
Host: Chris Daigle
Guest: Hernan Lardiaz, COO of Ragmetrics
This episode explores the essential role of evaluation (“evals”) in deploying AI systems in business environments. The discussion focuses on how business leaders and technical teams can ensure AI reliability, accuracy, and risk mitigation—especially as AI becomes integral to operations. Hernan Lardiaz, with 25 years in tech and AI risk-reduction, shares strategies for measuring AI output to bolster trust and operationalize generative AI confidently at scale.
AI Is Probabilistic, Not Deterministic
"In AI, it's probabilistic. So the outcome is not necessarily always the same one. And that's basically the fun of AI. That also drives that sometimes the outcome is not correct or is what we call hallucinating."
— Hernan Lardiaz, [06:20]
Risks in Regulated Industries
"He looked at me in the eye and said, Hernan, I'm not going to implement AI because I'm afraid of the outcome. We are in a regulated market. If something happens, we cannot control risk."
— Hernan Lardiaz, [20:59]
Operationalizing AI Requires Trust
Definition and Purpose
"...this is not just some theoretical topic. This is a key piece of being able to operationalize generative AI at scale."
— Chris Daigle, [08:06]
Manual vs. Automated Evaluation
Who Needs Evals?
"So there are two audiences, the business owner and the group that develops and implements AI solutions."
— Hernan Lardiaz, [20:55]
How to Build and Run Evals
Three steps:
"It's as simple as uploading the policy into our tool, our system, and asking the system to generate hundreds of questions based on that policy."
— Hernan Lardiaz, [22:41]
Continuous Monitoring and Drift Detection
"If you are measuring accuracy and you start to see that the accuracy slowly starts to degrade, then you can see that you're starting to have a problem."
— Hernan Lardiaz, [13:33]
Automation Brings Speed and Consistency
"We see human in the loop in a very minimal aspect...to retrofit information back into the system in the moments where it cannot be done automatically."
— Hernan Lardiaz, [27:55]
"...for your people to run 3,000 evaluations, the payroll on that would be way beyond the 250. Pony up the 250, do it right, remove the human error..."
— Chris Daigle, [48:01]
Mitigating Risk and Building Trust
Efficiency and Scalability
"...if someone wants to try the tool now and has a little bit of knowledge on the technology side, I'm pretty sure that in an hour he can be using it without any problem."
— Hernan Lardiaz, [39:01]
“If you play the lottery enough times, at some point in time you'll win it.”
— Hernan Lardiaz, highlighting AI’s probabilistic nature, [05:50]
“Errors cannot be avoided. That’s basically the core of what we want to resolve with what we do at Ragmetrics: try to understand what is that output of AI and how to measure and evaluate it.”
— Hernan Lardiaz, [06:20]
“Outside of if you just want to use AI for thought leadership and ideation... But if you want to operationalize AI...then this effort on evaluation is something that's extremely important.”
— Chris Daigle, [08:06]
“There are hundreds of examples...from Air Canada being sued because a chatbot didn't provide right information, to a fast food company putting hundreds of thousands of french fries on the same ticket, to New York City government providing information to someone to break the law.”
— Hernan Lardiaz, [18:46]
“My question to those guys [prompt engineers] in general is, OK, you change five words on that prompt, right? How do you know it's better?... what happens with the other 400,000 scenarios you haven't tested?”
— Hernan Lardiaz, [38:07]
“You don’t want to put AI into production and just hope that, you know, your employees are spot checking it on gut and getting it right. You got a business at stake here.”
— Chris Daigle, [50:07]
| Time | Topic / Quote | |-----------|-----------------------------------------------------------------------------------------------| | 00:00–01:00 | Introduction, real-world business fear of AI implementation in regulated sectors | | 04:07–04:49 | Hernan’s background and evolution of AI from academia to business | | 05:50–07:10 | Definition and importance of “evals”; probabilistic vs deterministic outputs | | 09:07–10:47 | Which businesses need AI evaluations; risk factors and regulated industries | | 11:27–14:48 | How to identify processes that must be evaluated, and steps to start | | 17:10–18:46 | Three approaches to testing: nothing, manual, automated | | 20:35–22:29 | Two audiences for evals: business owner and developers | | 22:41–24:45 | How to develop and run tests: policy documents, data sets, and criteria | | 29:32–30:47 | Where process-knowledge fits: capturing expertise in knowledge bases | | 33:09–35:48 | Continuous evaluation; automation and scheduling | | 38:07–39:01 | How to know when prompt changes are truly better—need for broad testing | | 40:49–43:15 | Real example: optimizing AI system for accuracy and token consumption | | 47:00–48:01 | Pricing models and cost/benefit of automated evaluations | | 49:33–50:19 | Evaluations as risk mitigation; “low-cost insurance” | | 50:27–50:50 | Final advice: don’t fear AI, but control and monitor its outputs |
For business owners, executives, or transformation leaders: Ensure you ask, “What’s our protocol for evals on our AI processes?” If you’re not sure, now’s the time to start.