
Loading summary
A
Welcome to Practical AI in Healthcare, the podcast that cuts through the noise to spotlight real world solutions delivering real world value. From patient care to clinical research, from life sciences to patient engagement, we focus on what truly matters in healthcare today. No hype, no theory, just practical insights where AI is making a true impact. Dr. Steven Lapkoff and Dr. Leanne Rosenblitt are your hosts as we explore what's real and moving the needle in this exciting new domain. Welcome aboard and let's get to it. Hello, welcome to this week's edition of Practical AI in Healthcare. My name is Dr. Stephen Lapkoff and I'm here as I am every week with my colleague, Dr. Leon Rosenblatt. How's it going, Leon?
B
Great. Really happy to have our guests here today. Going to be a really fun conversation.
A
Yeah.
C
So let's go right into it.
A
So interoperability has been kind of the broken promise of healthcare IT for more than 20 years. Pretty much at every conference I've ever been in, the same slide has always come up the same data silos. There's always a lament, someone saying that it's genuinely changed. Now you can, now you can do something different. Well, they've been saying that for a really long time, but I'm here to tell you that we've met somebody who's actually doing it, someone who's actually changed things, who is really actually now able to pull 85 to 90% of patient records, actual patient records, from almost anywhere in the country, thanks to a confluence of events, activities and laws. The evolution of healthcare information exchange, the evolution and dissemination of standards from HL7 like the CDA and FHIR, and the passing of the 21st Century Cure act, all these things have come together and allowed for a major shift in what's possible around the aggregation and collection of healthcare data. Mika Newton is our guest today. He's the CEO of XCures, a real world data and evidence company that spawned out from a cancer company. Mika, welcome to Practical AI and Healthcare. Why don't you help us unpack how these changes in healthcare data have allowed X cures to come into being.
C
Yeah, well, first of all, that was a great introduction of just what are the problems out there? And like you, I've been to many, many conferences for many years where we just lamented the fact that nobody had any access to data. All the data was all siloed in one place or another. So maybe the easiest way to kind of start unpacking is why don't I give you like a super quick briefing on our journey as a company through this, because we actually learned it by doing it. And I always think that's the best way to learn anything.
A
Perfect.
C
So, as you'd mentioned, we started out as a company focused on ecology. And we actually started in clinical decision support for advanced metastatic cancer patients. Where we were trying to build software is a medical device for which the comparator would be a tumor board. So basically, could we make software that would produce the same outputs as a tumor board? Well, we actually did make the software and it did work. And we published on it. Turns out that it wasn't easy to do. But the problem we ran into with every single patient who came in was how did we get our hands on the input data for that software? And the way we used to get the input data was to have the patient or their clinicians or providers send it to us in a FedEx box. That was pretty standard method, or fax it to us. And so we would get thousand page faxes of medical records and we worked on processing and bringing them all through, but it would still take us. And I think this was the biggest frustration for us, something like an average of 10 days. And we were dealing with very sick patients, and sometimes that 10 day delay was enough that they had already gone to the oncologist, they'd already made a decision on what their treatment pathway was going to be like. We couldn't help people because the delay was taking too long. And so we were trying to find a solution to that. When we really came into this world of interoperability that you started talking about, Steve, and that would have been right around got the company going in 2018. So right around 2022, we got super fascinated with the idea that there would actually be digital health data exchange across the country. And, and we started digging into it deeper and we realized that we had kind of stumbled into this moment in time where we thought there was going to be a massive change. And the thing that was happening is the stuff from the 21st Century Cures act was moving from kind of, you may want to do it to you have to do it, meaning that there were penalties being introduced in the legal regulatory space. And honestly, data exchange, the US went from 2022, call it 30% of providers, to 2025 over a two year period to the numbers you threw out of 85 to 90% data exchange.
A
So as all of this evolved and you moved from, you know, the cancer world into this new world, your entire world kind of had a shift. You didn't even start like this I mean, your, your, your world started in a whole different space, even before executors. And you know, what brought you, you know, we often ask our guests, in addition to their company's history, what about your personal history? What was your, your journey, your hero's journey to get you from where you started into this whole space? Because it's not an obvious, an obvious course.
C
So it's it. So where I started actually, if I go way back when, was in computational chemistry, proteomics and genomics software, so scientific software for early stage biological discovery. That's where I had very, my very first parts of my career. And I think the common theme across all of that part of the career was really focus on the modeling and simulation at the time. I mean, the software I sold back then ran on Silicon Graphics machines. We used to drive around with one in the trunk and go to pharma companies and demo it. I think if you had 16 processors, it was considered a massive amount of computing power at the time. And then as my career evolved, I kind of went through typical management stuff and I decided I wanted to do startup companies. And my very first startup opportunity was here in the San Francisco Bay area. I worked for a company called Archimedes Inc. That was owned by Kaiser Permanente and founded by a gentleman by the name of Dr. David M. Eddy. David arguably was the first person to ever write evidence based in the literature, even though a group in Canada then ended up kind of during Gordon Guyet's group ended up making so evidence based medicine and guidelines and policy kind of came from David and David's original thinking. I think he played a large part in the NCQA and HEDIS measures and other things. But we were building models of human physiology and healthcare systems. So we started trying to simulate what would happen. Could we create a SIM city or a SIM medical environment where people get old and get high blood pressure and have heart attacks and get diabetes and could we model the impact of various interventions, et cetera, on it? So that was in some ways the foundation for what brought me here. We got fascinated with the idea that could you use technology to predict basically health outcomes down to the individual level? It's like personalized medicine, but in a very different state than it was today. So we did that for a while. It got sold through private equity and ended up doing a lot of work in health economics and market access for biopharma, again using different models, a lot of Markov models who are built for that stuff along the way, and then ended up at a company called Dr. Evidence directly before exters there I was the chief commercial officer and we essentially were doing what I think of as very sounds a lot like what open evidence is today. We just didn't have the LLMs or the technology that they do today to do that. So when I think about exteriors and it actually always felt like a very natural evolution for me, which is how do we take real world data? In this case the real real world data not de identified but the actual data of individuals and make it work for them in a way. Originally it was meant to be essentially translational in the sense of like, hey, we're going to take your data and then we're going to analyze it and tell you what you should do. Right. And what that ended up then is the long journey into like how do we extract accurate, longitudinal, comprehensive evidence for any given individual and actually empower them and their healthcare providers with that data? Because it's just a mess at its fundamental level. It's very disorganized data.
B
So Micah, I think the story you're telling is just super interesting and it raises so many new fundamental insights. One, you're describing how the CDA document exchange now returns near complete records. Not just fire summary data. Right. And this is recent and real, not part of, not something that's going to happen in the future that we're all trying to hype. You know, I loved your background on the human. Why? I mean, you start out matching desperate cancer patients to therapies for vitro tumor boards and then you pivoted with the insight that there's machinery now to assemble one patient's record. But now we could turn into horizontal infrastructure for everybody. And I, you know, so I think that's incredibly exciting. There's a line you shared with us that I want to probe in our earlier conversations. Records travel, but they don't translate. The industry spent a decade building the pipes, so you know, fire tefca document exchange of various kinds. And you're saying that solved the transport problem but not the meaning problem. Unpack that for me. When a researcher, patient or an AI agent tries to actually use a full record that they got through this new infrastructure, they can get it more easily. But where does it fall apart?
C
Yeah, let me give you a couple concrete examples of how this data gets really, really messy fast. The first is in progress notes in electronic medical records. What is often done with the progress note is it's cut and pasted into the next progress node. So you end up with duplication of the progress note and basically the note Gets longer and longer and longer. And then somewhere along the lines someone writes over some part of it or et cetera, because it's just a free text field. And so when you actually get the documents, right, you might get 300 documents. It has kind of the same snippets of text that contain complex clinical information which are sometimes copied over correctly and sometimes an error. And so drawing the line through that and deciding what the actual ground truth is is not trivial. I'll give you another example. So, and I'll just. This is a made up example, but it's one I use pretty often. So I, I used to get care at Kaiser Permanente here. I now get care at John Muir as my wife happens to be a nurse who works at John Muir Hospital. So I, I like to go there. When I left Kaiser, I was a good guy and I got all my medical records and I brought them with me because it seemed like the right thing to do. Well, when I got to John Muir, they scanned them all in, right? And now exist as a, as a PDF, right? Or JPEG embedded in my medical records file in, wrapped in the CDA infrastructure. So if I now go to one of the national networks, call it care quality, and I run a query for myself, I'm going to get a response from Kaiser with my historical record and I'm going to get a response from John Muir. And the records that were in Kaiser are now duplicated. They exist once as the structured CDA output from Kaiser's EPIC emr and they're going to exist again as a set of images embedded inside the EPIC output from John Muir. So like imagine the stuff is all moving around and it just keeps stacking up and stacking up and stacking up. So if I think about just even the simple ingest problem, by the way, everybody has this in the transport problem. It's a common issue for EHR companies. The very first thing you want to do when you connect to these national level pipelines is normalize, deduplicate and convert to fire, right? Those are like very basic steps. And I will tell you, most people are trying to scratch that problem at this time. If I just look across the edge and the impact to patients and to the systems themselves of having this kind of disorganized data at the bank is actually enormous. There is a massive administrative burden that is being pushed throughout this system. Because if you don't have clean data, what do you do? You start again, Patient walks in, you started ground zero. You say, tell me about you. And I think We've all. I've been through this as a patient. I'm like, how come you don't know about me? I thought you had this medical record system, like, what, what's going on?
B
I was just recently a patient for a neck injury and oh my God, I've answered the same set, the same questionnaire, literally the same questions. I described it to Steve at one point, at least seven times within a of two days. And it gets old fast. So it's a real problem. And you know, so I think you're helping us differentiate a couple of things, right? So transport, moving the data is different from the meaning, making the data usable. Usable. That's sort of an organizing framework for the way you think about it. And then most of the record is unstructured. So a CDA doc, although it's organized and transportable, it may be comprehensive, but it's a bundle of pros. And you need to do a bunch of work to deduplicate it, pull out the important stuff and make it into something computable and comprehensible. And that process of converting unstructured data into clean, structured firesheet data is the hard problem, right? It's not the transport anymore. That's it. I mean, what I want to highlight for our audience, right, Just don't want people to miss this, is that in a sense, you're seeing something pretty radical that here has happened. The capabilities of EHR data aggregation have fundamentally changed just in the last two years and no one noticed like I was, I was trying to pay attention to it, but I think what you described was a bit. Was newer to me than it probably should have been. And I think you expanded that for us. And nevertheless, there's that problem of, of extracting meaning remains. Right. As and as we described in some of our conversations, new technologies often move problems downstream where the choke points come in later in the pipeline. Right. Because you removed an earlier checkpoint, it sounds like something similar is happening with meaning and structured data. I want to come back to two numbers you threw at us in our pre call that I thought were also pretty striking now. But you're getting 90% document retrieval success, right? So if I ask you for a document for a patient, 90% of the time you can get it, which is amazing and about 80% parsing accuracy on your platform, which is also amazing and frankly sounds like a lot, right? I mean, you know, having worked with EHR data, 80% would be very, very hard to achieve from a messy CDA file. I want to be honest with Our listeners about constra. So, and about that 80%, what's the 20% that's left? Right. What's wrong? That's that you just can't parse. That's. That'll be helpful to know. And where does a parse break in the. You know, and that what worries me, how do you know where. Where and when it's wrong as opposed to correct?
C
So let me kind of anchor on the numbers. So when we talk about getting access to data, we're talking about the level of participation of provider systems across the country, right? And that's actually the gating factor, because if you go to a provider who doesn't participate in health information exchange, right. Then that data is not going to be available. And there still are a fair number of smaller groups, in particular smaller practices that are not participating. Where it gets really good is around any major metro area that has large health systems at all. And so I just want people to imagine that it's a perfect world out there. It's definitely not. And your zip code matters a lot in this world as it does with everything else. Now, in terms of the types of data, if you think of the data structures, there's structured data in the CDA document, there's unstructured data, which is notes, and then there's embedded PDF, tiff, jpeg. So there's pictures of things. And I often say images. And I have to remind myself, I don't want people thinking DICOMs. We're not talking about DICOM transportation, RFI out from CMS right now, can't remember if it's been fully completed or not. But we're trying to understand how to do DICOM exchange on a national level. I think it would be a. I'm all for it. It's fantastic. But that problem has not been solved yet. All of that lives in proprietary PAC systems, right? In different parts. Those images are actually really important for lots of patients too, and physicians. But that hasn't been solved. So in terms of the actual structuring of the data, we look at kind of how do we take, you know, the structured data I would say is easy, but it's not. It's. There's a lot of actual normalization kind of term mapping that needs to be done with that. So you end up with normalized terms, number of companies operate in that space, you know, lots of tools. You could think of someone like an imo, for instance, I think is someone that people are very familiar with. Right. As to normalization, Steve used to work
B
for them Way back when.
C
Yeah, yeah, Great, Great product. When we look at the unstructured text, right. And extract from that, and then we think about also the unstructured images. Images have to get basically digitized. So we are really good at digitizing and OCR in that, because the PDF could be a digital, it could already be digitized, in which case it's really good. But when it's a pure picture, right. The biggest issue there is the quality of the actual image itself. So when something is. We've all done this, the photocopy of a photocopy, a photocopy. We have that problem in healthcare that eventually the quality of these images deteriorates. And when they deteriorate, the OCR NLP stack doesn't work as well. So just the term, you know, doing the digitization documents, that's probably the biggest area now, where there's still error and you have to go and check it now. There's some pretty cool stuff you can start doing with Bayesian kind of logic around that, which is, what is it, what do I think it is? And what would I expect to see in this document? And we've done some work around that sort of space and there's a lot you can do in that space when we take the final product, which in as we've thought about it, at least in my world, is not how much of the record have I actually parsed? But can I answer the question, meaning someone gives us a question, which is, I need to know these five things, these 10 things, I need to know it consistently for every single encounter, whatever that is. Right. And I want it, you know, just over and over again. We can then measure the accuracy of that. So in our world, we call those checklists, essentially, but they are the 10 or 15 things you need to do. The specific process, we are getting 95 to 98% accuracy, like really good F1 scores. And we published on this and so that's actually where we focused our effort is not really how complete is the record, but can we solve for the clinical use case on the back end? And then we like, we won't accept 80%. It needs to be 95% plus is basically our internal. Our internal metric and why we provide source verification with every document. I think it's really important, by the way, anyone who's doing AI at all in healthcare, you have to know where the source document is. If you lose that, I think that's one of the biggest issues we all have with trust in the AI world and healthcare, but you actually don't want people to have to verify. So if you say trust but verify and you go back and you verify something every single time, you really not move the problem down the street because you're still going and reading the medical record. Right. So you would have to get to a level of consistency and process where generally. Right. In most use cases, the end user, the consumer can just believe the data even though it's gone through this AI powered filter. And I think that that's really the critical part to unlocking this sort of technology. Otherwise we're just more tools on the same problem.
B
So, yeah, Mika, I just want to probe that a little bit because I think I agree with you that automation becomes valuable if we remove human attention. And that always is a crit, both. But during the learning process or during your QA phase, you, you know who's in the loop and where do you draw the do not trust it yet line? I mean, you understand it's, it's not, it shouldn't, I agree with you, it shouldn't be part of the production system at the end, you know, but, but it raises the two questions, right? You know, how do you do that checking during learning and refine and get it to whatever the acceptable line is and then where do you tell the customers not to trust it yet? Right? How do you decide to say, well, it's not good enough for this use case?
C
So the gold standard is human extraction, right. Of knowledge from medical workers. So we do kind of the dual extractor adjudication standard path, right? Two people look at the data, right? If they agree, it's good. If they don't agree, then either one of them or a third party comes in and decides what's right on the actual piece. And that's what we use for testing each one of our individual models. And by the way, I think that that again, has to be done for anything in AI. Now, we have a really interesting argument that we have with ourselves all the time, which is the humans also make mistakes like even, right. And by the way, dual, you know, extraction is actually not the common standard. The common standard is a person reads it and decides whatever they think they saw in the record. So the question is, is that actually better or not? And then we tune to the individual field level, meaning it's not that I get 95% of the whole thing, right? It's that each one of the individual fields is always tested right at that level. So you try to make it very, very specific to the model. Now what I'm excited about, right, and we're starting to move forward and this is really interesting is actually what we call auto validation frameworks. So if someone asks a new question, right, let's say that we've developed a model for a certain clinical answer. Well, what if the person who comes in next has a different clinical definition of it sounds like the same thing, but it's different. Which by the way happens in healthcare all the time. Like how do you define diabetes? Well, there's lots of different definitions. Right. And we see those typically in the literature. I would think of them as measures. There's all these validated measures that have been done in various studies come up, but even the measures kind of vary between themselves. So what, how do you accommodate the fact that the way two people think about something is different? Well, you want them to be able to explain what it is and then you want to be able to automatically validate. So think of taking a large set of either pre configured data or synthetic data that you can create and then essentially using a chain of AI agents to validate the model. So now we can build for instance, judge models and those type of things actually judge on how accurate the model is. And I think that that's going to be a big part of the future. That's going to give us a lot of flexibility about it. So sorry, I kind of went down, way down there where we're going to go, but
B
so.
A
Oh, I thought you froze there, Leon. Nope, you're still there.
B
No, that was, I was really fighting the urge to get into the multi agent, you know, agent councils and techniques that allow you to trust communities of agents. But we're going to save that for the end of the conversation if we have time because. Yeah, so Steve, take it away.
A
Mika, you've outlined a lot of the back end work that had to be put into position for this all to work. But let's talk about the mechanics of how it's actually working. You're processing millions and millions, like over 200 million records across 50 states in Puerto Rico. Can you talk a little bit about how that actually mechanically happens and what role AI plays in the final processing?
C
Yeah, so just to give you a number, this week we ran 600,000 people's records this week alone.
A
That's crazy. You know, I got to tell you something just to break in a second. When I, Leon and I worked on a big registry project at the MMRF and we had to pull medical records for that registry and we had a meager goal of 5,000 patients for that registry in the year after year and a half after we opened, we got 750 patients to register. But given how long it took us to get the records for those patients, it was taking between three and six months to collect these medical records. And they were coming in boxes, they were coming in envelopes, they were coming in faxes. Only a tiny, tiny percentage of them were coming as data. And, you know, we couldn't do what you're doing now in the, in the blink of an eye. Yeah, that was.
B
And the abstraction was. The abstraction was garbage. And honestly, and this was after everybody already agreed that patient mediated data extraction was the way to go.
A
Yeah.
B
There was just no way to actually force it to happen.
C
So you're describing the world that we started the company in, by the way
B
we lived it. Yeah. That's why you're telling, you're telling us you discovered a new continent. Right? Like, we're all guys going, like, that's not how the world works. This is hard. What do you mean? So it's, it's kind of cool.
C
So.
B
Yeah. Sorry to interrupt. But. But go. I mean to just. Yes, keep going. Tell us how, how are you doing it?
C
So 600,000. So basically what we have to do is so first of all, we work under baa. Our customers are covered entities who are provid care to the patients themselves. And so we're acting on their behalf. Think of us as the picks and shovels, the tools. Right. That help it get it done. We connect to a on ramp. The on ramp we we love. And there's many of them out there. But the one we use is called no2kano2.com great people. John and Trace over there are wonderful and their whole team and they provide us with access essentially to 2. Two national networks that we leverage. One is called Care Quality and the other one is called TEFCA or qhins Q H I N S. And I think your audience, you have an ability with. You put the podcast, I gave you guys a little primer with some links in it, maybe you can post it because there's a ton of information out there on those individual sites. So what we need is to get someone's data is basically their PII that comes over. So you need first name, last name, date of birth, sex at birth. And an address is ideal. Generally a zip code will get you started, but the better the address data is out there. And the first step is we have to find the data. I think a lot of people start immediately imagining some central warehouse or database that this all lives in. And it doesn't. It actually sits on all of the individual EHR instances, wherever they are all over the country. So if we think about that, we think of that as locations. And to date we've touched 550,000 locations. That's crazy. Yes. So how do you find the location? Well, the address starts to dial in, so the chances are that you got care around your address. So then we do a bunch of stuff with consumer databases. Try to figure out all the places you could ever have lived in your life. Right. There's a lot of other data out there. We try to figure out if you're close to other states. Right. Because some people will cross state lines. You gotta like kind of figure out where it is. And we've optimized for maximizing the return without overburdening the network. If you can imagine, these networks have certain throughput capacities themselves. And so you also want to be a good neighbor to all of the other network partners. So if you're constantly asking someone, hey, do you have Mika's data? Do you have Mika's data? And like a thousand times in, they're like, no. And by the way, I'm never answering your question again. We're just going to like turn it off. So there's, there's. And they're reciprocal networks. Right? Meaning if you pull data off the networks, you also have to give data back. So there's a whole bunch of rules of engagement around that. So step one is find out where all the data is. So that's a query that's run which identifies the location of the data. And then you send a data request which all happens electronically now, which says, ship me what you have. And you can put some date ranges around it, like kind of how far back do you want to go? And that sort of stuff. So there's just, these are features in the query logic of the networks themselves. There's a lot of configuration here that you can optimize around. Again, depending on what you want to do with it for different uses, you're going to have different components. And then what you generally get back are these raw CDA documents. You were talking about our XML files with all of the associated information in it. And then that starts the mechanical process I described of essentially moving through a data pipeline. You can turn on different parts of it. So part of it is the normalized dedupe. There's a convert to FHIR component. There's what I think of as metadata enrichment. So for instance, looking at all of the NPIs associated with the record and identifying what specialty wrote it. So you can actually tell which records were written by which type of specialist. Right. So that you actually understand who might be authoritative on a particular subject versus not so example, a nurse might refer to the oncologist's notes, but if you have an oncology question, I think you want the oncologist's notes, not the nurse's notes that refer to the oncologist's notes. And so a bunch of logic like that around who wrote what, where, when. There's semantic embedding and labeling, which is when you receive all this information, what exactly did you receive and what's in it? So think of. I always think of that as like making the master catalog of it. And that needs to go down to kind of the individual section level in each one of the documents. So you think of, you got the document, you clean it up, and then you start adding stuff to it. And you try to add as much useful stuff as possible, given that it costs money to add stuff. And you're burning compute on the back end. And then the last piece, right, is you now have these organized, normalized records. Looks like someone was in a single emr, right, for their entire lifetime with all the same naming through it. You then put models on top of it and actually ask questions and answers of it in logic. So that you have essentially prepared for what I call it a complex inference. So something like line of therapy and cancer disease progression, et cetera, not a single data point in the document. It's a clinical inference that you would get from reading many documents. And that's actually, I think, where the magic then lies. And by the way, the accuracy things we were talking about earlier, I think are really, really important because it's actually those inferences that are the crucial pieces of the information that has to flow forward.
A
You know, when I interacted with you guys, and I'm just for complete transparency, I was an advisor to you guys early on when you were kind of in a different instantiation of your company. And this year at himss when I ran into you, you guys had more or less changed your business model kind of dramatically. And I'd like you to dig in if you don't mind, if you feel comfortable with it. Can you explain how. Where you were when you started and what the pivot was about, and how that pivot, you know, either was facilitate or, you know, what caused that to come into practice? Because it. It's an interesting observation, and I saw that happen not just to you, but to others at HIMSS as Well, other companies I've seen.
C
So I always tell people I think it looks like from the outside we pivoted, but in many ways, in our own mind, we just went deeper and deeper down the rabbit hole on the technology, right? Because we couldn't do what we wanted to do. So we had to go solve something else first and then go solve something else first and go solve something else work. But we definitely pivoted on the deal and business model, right? And it was aligning the business model to that really allowed us to go and validate product market fit, which is the most important thing for small companies. And people ask me, what does that look like? I'm like, you know it when it happens. Because suddenly people start buying your stuff, right? And telling you how great it is versus telling you that that sounds really interesting and nice, but come back some other day. So when we started the company, the model was capture data from cancer patients who are going to give us the data in return for answers on how to treat their and we will go and monetize that data for research, right at various points. So we were going to be a true classic real world data company, like many of the real world data players out there. You would think of like the flatirons and codas of the world and so on and so forth who are out there. And our differentiator was going to be the direct to consumer, right? We would capture these advanced metastatic cancer patients who were falling outside of typical research networks and we would get more comprehensive data than other people. That was our thinking back then. So in doing that, we ended up starting to build the infrastructure to capture the data, right? So the tools we had today actually started back then. They started as the electronic data, the EDC Electronic data capture system for our software as a medical device. That was the original pieces of the software. And so we were building it that way and we continued to build it up until that was 2019, 2021. In 2022, right when we started connecting to these health information networks, a few of our clients had come to us and we were doing research work for them. At that time we were selling data and we were running intermediate expanded access studies. So we were providing access to cancer drugs and helping facilitate these treatment protocols, et cetera. So it's really a tech enabled services business. And one of our customers came to us and said, yeah, we want to do a lot of this. We've got a big load of patient, we really don't want to pay your fees. Like if we pay you to do this, it's going to be exorbitantly expensive. But we keep seeing the tools that your team is using, right? Because we see it when we're working with you, could we just use your tools? And we're like, we never designed them for anybody else to use, but sure, go ahead. And they came back and actually went also to a research institute who we also gave the tools to, because we thought, well, let's try it out. And they both came back and said, the tools are terrible. They're really hard to use. They don't really work very well. But one of them, who's our client to this day, actually came back and said, but we'll work with you on it. Right. And so we got a development partner in that model and we started. That was really our first SaaS license. But we kept the idea of selling data. So. And this was, I think, an interesting business decision for us. We would license people our software and then we would discount the software use in terms depending on the amount of data rights they would give us because we remain fixated with the idea that we could acquire and aggregate the data and that then we'd be able to do something with it. In 2024, when we started really expanding, so 2023 was a good year for us. We did all the test pilots. We got some clients up and running. In 2024, we went fully commercial on the software with the same kind of dual research, data and aggregation model in the result. And we were going live. And we were using a supplier at the time called Particle Health. And Particle Health and EPIC had a disagreement, which you all guys can go read about online. There's actually ongoing litigation related to it. And as part of that process, many of Particle Health's customers were disconnected. We were one of the groups that got disconnected. And in talking through the regulatory pieces with all the participants, we basically had to. We moved over to Kano 2. We had to bring all our clients from Particle over. We had to re qualify all of our clients on the network. The networks have a know your customer piece. And it got really serious at that point. Because of litigation. Yeah, we got really deep on what is or is not a permitted use case. And that discussion continues to this day on the platform. And what we decided at that point is that we did not want to take the risk any longer of trying to be in the data and insights business itself, other than we wanted to enable our customers to do that. And we went really pure SaaS at that point. We said, we are going to. And by the way, as soon as we did that, a whole ton of people who didn't want to do business with us were suddenly very open to doing business with us because they were very concerned about the commingling of data rights. It was a really difficult thing to talk about in contracting. It was difficult to contemplate giving another party access to rights. And then how would you control that other. It was just too complicated. And when we made it simple, it kind of worked for us. And by the way, it doesn't stop you from doing any of the interesting things about insights and all the rest of it. You just have to start thinking of the world a little bit differently. And many of our customers get insights from their own data. And seeing as the data that transfers with the patient is everybody's data, and we haven't talked about that yet, the fact that it isn't just the one institution, it's all the institutions that person has been to means that you end up with these really rich data sets even inside of an individual, call it secure tenant for someone that provide a huge amount of insight and have downstream purposes. So, and by the way, I don't think we're the only ones who do that. I think some of the, and I won't say this definitively because I don't work, but I think folks like Ontada and others not only provide you the data, but they provide you an interface to the data that allows you to track insights out of it, not just a raw dump of the information.
B
It's actually such an interesting business move and I could probably think of reasons why commingling of data rights is such a problem from like, you know, from a game theory point of view. Modern economists who follow like Kenneth Arrow's ideas would say there's an information asymmetry in their relationship and it causes real contract, real misalignment of interest. By the way, when back when I was CEO of Prometheus Research, we took a similar decision. We were focused on real world data and we just said we're just a software vendor or a services vendor. We are not here to get your rights. And it enabled us to work on a lot of very interesting projects where if we insisted on data rights like everybody else did, we just wouldn't have gotten them. In fact, we had an open source platform to make things even more interesting. But so I really appreciate you sharing that story because I think it's points away the way the market is going and where messy contractual arrangements are never going to get broad acceptance. They're just too Expensive to negotiate. And there are several other things that your earlier comments have pointed to as just real business model innovations. The business associates model that acting on behalf of covered entities and that being a covered entity is a really, really interesting twist in the way things used to work. I mean, you know, people used to fight tooth and nail not to be a covered entity because you did not want HIPAA applying to you. You're like, I'm not a covered entity. I'm just a little researcher doing some research over here. Covered by the common rule, but not by hipaa.
C
Right.
B
So I think that's a, that's a very novel and cool move that's worth people understanding. And then, you know, you shared an analogy with us in our earlier conversation. I want to kind of bubble up, which is that, you know, the, the TEFCA and care quality. These are like the, you know, this network model, this is like participating in the Visa card network.
C
Right.
B
And the hospitals are like the banks. Right. They've got the data, but there's the network. And then, you know, the, the question is, who is running the credit card and who, who owns the rail? Right. I, I, I think that's a very interesting analogy. And in generative way, it sort of got my, my juices going. Maybe not my creative juices, but something. But I, I want to test it with you just a little bit to see if it leads us anywhere. Right. So I want to sort of see where it doesn't bend far enough. So Visa moves one fungible thing. That's dollars you're moving meaning. Meaning doesn't reconcile cleanly the way that money does. Is the semantic translation problem. Exactly. The part of the card carrying, is that the part that the metaphor kind of obscures a little bit.
C
Right.
B
Because there are two parts. There's the transport and then there's the meaning.
C
Yeah. The Visa MasterCard is clearly the transport layer, just to be explicit.
B
Yeah, yeah. So I want to make sure that's like, I want to elevate that. So hear your reaction.
C
Yeah. And I think what you, what I think about is like, okay, so I need the payload, I need to get it. By the way, you're bringing up something I've been thinking about a lot, particularly as it relates to health information exchange. So if you think of the hies jumping a little off topic here, think of those as the regional banks or the aggregation of regional banks, particularly often community hospitals and other people who don't necessarily have all the infrastructure that they need to do patient data exchange. And they exist in Almost every state and have been around for decades now, for the most part. If you think about having. In California, you might have 1,045 million people and thousands of participants in the health information exchange. The transport layer actually gets really messy because if you send everything to everybody, because the way the query works now is, hey, send me data on this person. And then you say, here's everything I have on that person, right? You're moving a lot of information around, a lot of which is not required. I often compare this to, like, why was the facts actually kind of good for so long? Well, if I call you up and tell you to fax me the pathology report and you're another human being, you understand that I'm asking for the pathology report and you can go and find that pathology report and just send me the pathology report. If we're computers, you just send me everything. Now I have to find the pathology report. So there's actually more work for me, right, in having this exchange. But if we think about these local exchanges and everybody who's sitting on the edge of it, everybody who's sitting on the edge who's trying to provide some service to the patient doesn't actually need or want all of the data. And there's a huge amount of cost associated, right, with massive transport. Plus the duplication stuff we talked about and all the error, all that. So how do we start putting a surface on all of this data so you only get what you need? Meaning that question. Answer that. Semantic logic. I would like to see it work its way down into the Visa MasterCard transport layer. And it's not there right now. It's a rally.
B
It's not there. The problem is it's deeply connected with the meeting, because asking for something specific involves you understanding what it is I'm asking for. So I'm worried that there's an interesting circularity there, but I want to. Let me just probe one other theme that's come up. Second big theme is that the customer is changing you. Going from B2B data deals to patient relationships. A patient relationship is primary. With TEFCO's Individual Access Services identity verified through something that clear, a patient can now pull their own complete record and point it at an AI agent, or vice versa. Point the agent to it, right? Based on what you're actually seeing, where does that go? What is a patient going to do with the data?
C
So there's three rails that work, right? In my world, there's the treatment relationship, which is, you know, provider to provider data exchange for caring for A patient who's a third party in that there's bring your own data. So people who sit on large chunks of data and there's lots of dirty, I mean your MMRF example, you got lots of data is dirty. It needs to get cleaned up. And then ias, which you brought up, Leon here, which is Individual Access services. So that's my right under HIPAA to get my own data. It's no different than showing up in your doctor's office and saying, print out my files and give them to me. But I want them electronically. And I want them electronically so I can stick them in, I don't know, some AI decision support tool or my wellness app or list one of the million things. IIS is going through a transformation this year, which is like the networks did a couple years ago. And probably over the next 12 to 18 months, I think we will see IIS data. It's the provider participation issue. There's actually pushback from the America, from the AHA essentially to TEFCA right now, asking them to delay mandatory response to IAS because people are very worried about patient matching, right? Meaning if, if it's not a covered entity to a covered enemy governed by hipaa. If data is moving from HIPAA into consumer data, which it becomes the second it leaves the covered entity and gets given to the consumer, how do I know that I didn't breach, I didn't send the wrong Mike Smith data to the right person? And there's a lot of debate around that. Honestly, I feel my, this is a personal opinion. I feel like if clear is good enough for me to get on an airplane or ID me is good enough for the IRS to let me read my, my tax files. Those methods are probably good enough, right, for health data exchange. And the real question is, how do you manage the risk? Like somebody's going to have to absorb the risk and who should it be that's in there? So there's a little bit of legality right in. But I don't think the identity piece is a big problem.
B
It's super interesting. I got it. But it raises just one more business level question. I mean, first, to just summarize, right, consumer mediated data access is the big structural change here. This is not just like a random peripheral feature. This is the really big driver. And now that a patient can pull a complete record, they can point it and can point an AI agent added or vice versa. This creates a new shape to the entire relationship. What it opens though, is an uncomfortable question that the field tends to dance around which is data monetization. When the patient becomes the access point, who captures the value? Is it the networks, is it the vendors, is it the patient? And where should the patient sit in that conversation?
C
So the patient has to give explicit consent as part of the process and the consent has to be managed. So that's written into the rules. What the patient chooses to give consent to various companies is now that the patients need to. Patients, us, the consumers, not patients anymore. It's just regular consumers.
B
Just people.
C
Just people. You need to understand what you're signing up for. By the way, we've been through this in marketing and fintech and all the rest and we all just clicked everything that, whatever and we never read any of the fine print. Right. And anyone who tells me they did, I don't believe them. Right. So that data is all the way out there anyway. By the way, your consumer data is way scarier than your healthcare data. If you ever really dig into it
B
on the back end and it's for sale for like 20 bucks. I've done it as an exercise in a class and it's terrifying. Yeah, yeah.
C
So we're entering that world. It raises a lot of questions like you brought up, but really I think the question is if you going to go sign up for random AI bots to do healthcare stuff for you and you start giving out the rights to your healthcare data, all those random things, your day is going to get out there and be monetized by lots of other people really, really fast. And I think this is now consumer decision making. I think we live in a society where that's a risk that we've already taken. I've talked to people in healthcare who have a very paternalistic viewpoint who say basically you're not smart enough to handle your own data. It's been, I've been told that don't. I understand where they're coming from. I don't agree with them personally, but I have seen that pieces. You don't understand what it's.
B
But you're making a great case for the work that Steve and I are doing with the DCI network on data governance. So I think we're all sort of agree in that direction.
C
So.
A
Yeah, yeah.
C
But you know, if you think about it, what are you going to do with your data that's valuable. And I think you need to focus on what am I getting out of it. I don't think that any individual person, a lot of people, like, oh, people should participate in the monetization. Like, do I really care about getting 10 or 15 bucks for, like, someone, because I'm going to be part of some aggregate pool of data. Like, like, no, I'm not going to resell my data, you know, a million times. So it's $10 million worth of like, maybe I get a hundred bucks over my life. Like, I don't care that much. So what I care about is the service. Right. And if I'm procuring the service, I probably am less. Like I'm paying cash out of my own pocket. I'm going to be way less likely to let them use the data. But if they're going to give me something for free. Right. Maybe it's really valuable to me, then I'm going to be more inclined to let them use the data. But I'm going to want to understand where and how it gets used on the back end. There are always bad actors out there. Every industry has bad actors. I think it's really easy to focus on them as, like these edge cases of the horrible things that go wrong. Though I do believe that there's a very large majority of people who actually have the best of intentions here and are looking to build good businesses in healthcare that are going to be helpful for people. So there's no clean answer there. But I, I think it's just that. Just understand who you're in the room with.
A
Yeah.
C
What they want to do.
A
You know, I, I've been listening to all this and I've been sitting on my hands, not trying not to interrupt. But this is, you know, it opens up another. There's a contra question here which I'm going to put on the table, and we are running out of time, so we're going to need you to answer relatively quickly. You know, the patient privacy aspect of all this. I know that patients can walk into a hospital and opt out of these things. I don't know that. Patients all know that. I don't know that because, you know, in the real world, and this just literally happened to my wife, we talked about it, I think in the prep, you know, she went into a medical office for a preoperative procedure and was effectively told by a front office staffer that she had to sign a consent in order to be seen. Now, that was wrong. It was actually illegal. She wasn't supposed to be coerced into signing something. She did sign it. She was under pressure and all of a sudden her medical records in another institution became shared with the institution that she didn't want them to have. You know, your system is really all about building this layer of trusted data exchange. But what happens when the people, patients, don't want to participate? Can you get them flagged in your system so that you can be the arbiter of privacy, or do you even work at that level?
C
We don't work at that level. In the covered entity world, in the treatment relationship where it's. We're acting on behalf of a covered entity asking another covered entity for things they actually control, who, what they choose to exchange. We're simply in the middle of the transaction. In the IAS individual access services world, every data type and layer consent goes down to the type of data. It isn't just at, like all my data, it's by, you know, medications, allergies. Like, it goes down to the data type itself and you have an explicit right and it's required to be in there to delete your data. I was actually thinking about your wife's case, Steve, and I was wondering, can she just go back to these institutions and insist that they delete the data?
A
She's actually done that. She's gone to the state attorney general of Connecticut who gave her that advice, and she's actually executed on that. But, you know, just like everything else, as soon as it's out there, it's almost impossible to pull it back. And no matter what you do, listen, we're running short on time. I wanted to sort of like bring us home a little bit if you had, you know, how is this all going to play out over the next 12 months? If you're. If I'm basically a CIO at a health research institution, I'm a leader there, and I'm listening right now. What should I be doing over the next year or so to be ready for all this? Because what if I'm not part of it and I'm going to get into it? What should I be thinking about?
C
Yeah, I think the most important thing I think people doing is starting to realize that data is an incredibly valuable asset, especially in a world of AI. It is probably the single most valuable thing that you can have that's not money itself or compute is the data that's in there. And so what governs my access to data, it used to be that it was the other partners I had who I had agreed data exchange agreements with with. Right. Meaning my institution has an agreement with this other institution. We agree to exchange data. We have a data usage and agreement between us on who can do what with the data, et cetera. Well, now we live in a world where, as we talked about the data just moves around. Right. And who does it move around? It follows the patient. So your relationship with the patients and the services that you provide to patients, again, whether it's a treatment service, right. A clinical service or an AI service for individual access or whatever it is, that is the value. And I think for the first time I see patients having a voice in healthcare because they are the mediators of their data, whether it is through their own piece or through where they choose to get their care as to who gets to access them. And that's really different.
A
That's a serious sea change in how the world has shifted. And I think new companies are going to be evolving, new capabilities are going to be evolving. And, and I keep telling folks, we are at the very beginning of this whole AI journey and how the data is shifting. We're at the beginning, we're nowhere close to what stability is going to look like.
B
Yeah, and I love that ending on this note. I think it sort of brings a little synthesis to everything we've been discussing. Is that moving data used to be the hard part and it isn't anymore. Now making it mean something is the hard part and the value is migrating to whoever sits closest to the patient. That's a big change in the industry and I just want to encourage. Thank you for bringing it to our audience attention, Mika, and really appreciate you being here. What a delightful and informative conversation. With that, I just want to thank you for joining. Thanks Steve, for co hosting and thank our listeners and I hope you'll all join us next week at Practical AI in Healthcare.
A
Thank you for joining us this week on Practical AI in Healthcare. If you're ready to go beyond buzzwords and hype and explore how AI is truly transforming healthcare, stay tuned for more conversations that get us to what works. Until next time, stay pract.
Release Date: July 12, 2026
Hosts: Dr. Steven Labkoff & Dr. Leon Rozenblit
Guest: Mika Newton (CEO, xCures)
In this episode, Drs. Labkoff and Rozenblit speak with Mika Newton, CEO of xCures, about the real-world transformation occurring in healthcare data interoperability. The conversation moves beyond the “broken promise” of the last 20 years and shows how new standards (CDA, FHIR), regulatory momentum (21st Century Cures Act), and changing business models are enabling near-complete, nationwide electronic health record aggregation, and what “intelligent interoperability” means for providers, patients, researchers, and the entire healthcare data ecosystem.
“Interoperability has been kind of the broken promise of healthcare IT for more than 20 years.” – Dr. Labkoff (00:48)
“If you don’t have clean data, what do you do? You start again. Patient walks in, you start at ground zero….” – Mika Newton (11:40)
“We won’t accept 80%; it needs to be 95%+ ... and why we provide source verification with every document.” – Mika Newton (18:30)
“This week, we ran 600,000 people’s records.” – Mika Newton (24:18)
“As soon as we did that, a whole ton of people who didn’t want to do business with us were suddenly very open to doing business with us…” – Mika Newton (35:45)
“TEFCA and Carequality…these are like participating in the Visa card network. Hospitals are like the banks….” – Dr. Rozenblit (39:09)
“What I care about is the service…if they’re going to give me something for free, maybe it’s really valuable to me, then I’m going to be more inclined to let them use the data.” – Mika Newton (47:06)
“Moving data used to be the hard part and it isn’t anymore. Now making it mean something is the hard part and the value is migrating to whoever sits closest to the patient.” – Dr. Rozenblit (52:30)
For leaders in health research, clinical operations, and AI in healthcare: