
Loading summary
A
Welcome to Just Now Possible with Teresa Torres.
B
I'm Tulsi, I'm director of product and technology at Hotility.
C
I'm Lorna, I'm the head of data and AI at Hotility.
D
And I'm Jack. I'm head of engineering at Hertility. I've led the team here for the last two and a half years.
A
Excellent. And tell me a little bit about what does Hertility do?
B
Fertility is a women's health tech company. We focus on everything to do with female gynaecology. So that's not just fertility related as the name might suggest, but everything from menstruation through to menopause and we do at home diagnostics. So the idea is that you fill in an online health assessment with us which comprises of up to 50 questions. There is proprietary logic that sits behind those questions and the idea is that we gain a rich picture about you, your medical history, but also your feelings, your experience with the healthcare system. Things that potentially you wouldn't typically be asked in a consultation with a doctor. Off the back of that you have a bespoke hormone panel created for you. We only test you for things that you need and are actually going to provide us with data that will help you achieve your goal. And then you would go on to do an at home finger prick blood test. You only need 5 mils of blood and that can be posted back via the postal service through to our lab. Your sample is analyzed against the panel that we assigned you to and then a doctor will review your responses to the online health assessment and your results to provide you with a comprehensive report detailing each of your results, whether they're in range, out of range, what that means and then put you onto our care plan which contains next steps, either telemedicine related or in person clinic related, again to help you with your health journey.
A
Okay, so I'm going to guess based on what you described, this is not a consumer app where I just sign up. Is it something that a primary care doctor prescribes? And then I also am curious about like just region. I know a lot of medical stuff like is approved by different regulatory bodies. So give me a sense of where are your customers, what regions is this available?
B
It is only available in the UK and Ireland at the moment. We are accredited by the CQC and our product is CE marked as a medical device. So approved from a EU regulatory perspective with regards to how it reaches people. It is actually direct to consumer. In the UK a lot of the healthcare system is obviously government funded. We have the NHS So from a private sector point of view, this would be something that you could personally decide to purchase yourself. On our website, we also work with employers to have hertility products and services available to employees as a benefit. So your employer would purchase these upfront and then you would have the availability as an employee to redeem them. But we do also work with a lot of clinics, particularly in the ivf, egg freezing and pelvic scan space, who would potentially refer people to hertility as a product. It's less a prescription and more of a recommendation.
A
Okay. So for listeners, for women in the uk, they do have the option to just go direct and then it's possible they may have it through their employer or it's possible they may hear about it through their own clinical relationships.
B
Absolutely, yes. And we also operate with an insurer in Ireland as well. So for those who have insurance coverage with this particular insurer in Ireland, they'll be able to redeem the services through there, too.
A
Okay, great. And then I understand we're going to talk through two different AI products today. Before we get into the specific products, I'm really curious about for your team, for hertility in general, like, what was the. What's your experience with ML or AI products? Was this brand new for you? I know for a lot of teams, they're just getting started with this technology and everything's new, so I'd love to just hear a little bit about your background. And as you started to look at these two products, did you have experience with this or was it really brand new for you?
D
One of the really cool things about ATLITI is we've got amazing data. So that first touch point that all users will have with us is this online health assessment where you give an answer maybe like 50 to 70, up to 70 questions, to give a really rich sense of your medical history, your cycle data, exercise, lifestyle, that sort of thing. So it gives a sense of your symptoms. Often you'll then purchase a blood test, you get your results back. So immediately we're connecting symptoms with blood test results. And then a subset of our users will then go on to have a pelvic ultrasound scan with us. We'll speak to a gynaecologist, or we have a suite of telemedicine consultations like mental health or nutrition or fertility advice, that sort of thing. And with a pelvic ultrasound scan, you'll get an image at the end. So we get an actual visual representation of your reproductive anatomy. Or with alternative medicine, you get a letter. So you get this amazing ability to connect symptoms with blood results and then a diagnosis, an outcome, a letter or actual. An image of your ovaries and how many follicles you have. So we have this incredible rich data set which we've been building up over the last seven years. I think the number is a million women have started our health assessment and then a subset have. We have blood results and slightly lower numbers of vulgar scans, et cetera. So I would say it's one of the more unique women's health data sets in certainly in the UK and Europe, but possibly around the world. And our ability to connect, as I said, how you're feeling now with a diagnosis, I think is quite powerful. The obvious thing to do there is to build predictive models to. Lorna can probably speak a little bit more than I can, to try and predict based on your cycle length or your previous diagnoses or various other symptoms of how you're feeling with an actual diagnosis. For example, polycystic ovary system. Polycystic ovary syndrome. And that's something we've been one of the things that we. Is one of the core tenets of going AI and also feeds into the scan automation product that we built.
A
Yeah, yeah, yeah. One of the things I love about this is I know this is a space where, just from a research standpoint, there's. I feel like medicine is not invested a ton in women's health research. I know here in the us we have very limited time with our primary care and a lot of women, their gynecology care is through their primary care, not through a gynecologist. That's not true for everybody, but is often the case. And I think you're. I suspect, Jack, what you said is that this is a very unique data set simply because it hasn't been an area really focused on by researchers. So that's exciting to me that a company can help to fill that gap. Lorna, did you also want to jump in?
C
Yeah, for me, I think, like the way, I guess part of the Hotility mission, which is to improve health outcomes for women and reduce the gender health gap, which we definitely see, because, as you've already mentioned, there's no way near enough research spent on women's health outcomes. Plus the data that we've got, I think naturally leads us into looking at what machine learning or AI opportunities can we offer, like using the data to help provide more informed diagnoses, better diagnoses, more detailed diagnoses than you might get from your gp. And so, yeah, I think part of what we're doing with AI just naturally fell out of the. Like the research and the mission, ambition and fertility to just generally just not just give people diagnoses, but give people better and more informed diagnoses on top of it as well. And I think the other angle driving it is efficiency as well. So obviously we have quite a lot of moving parts. We work with clinicians, we want to be able to sell our product direct to consumers at a price point that people will be able to pay and that isn't as expensive as if you went through a full private health insurer in the uk, for example. To be able to do that, we need to make our processes as efficient as possible, make the doctor's time as well spent as possible. And if we can't do that, it's going to be very difficult for us to scale in the way that we want to. So I think the research aim is driving where we're looking at AI opportunities and also the efficiency aim as well. And it's kind of what we've done so far is like those two things coming together, I think.
A
Yeah. There's something I want to highlight for listeners that may not be aware of this, just to give some context to your problem space. I think for decades, a lot of medical research excluded women and they excluded women specifically because pregnancy and reproductive cycles were like confounding variables and they wanted to have clean data. But the consequence of that is women's health. We assumed women were just like men and we're learning that's not true. And so this is where this gender gap comes from in the research is that. And then the consequence of this, one major consequence of this is that women's reproductive health in particular was just sort of left out of the research. Now, I think in recent decades, we're starting to correct for this, but I know I still regularly see study results that are all men, or so it's not completely solved. And so I do really love that. I think also, like, when I first saw your application, I was like, 150 questions. I'm curious how many women answer all those questions. And then my next thought was like, I bet they do, because there's no other way to get help in this area and they'll get guidance. It's such a huge gap and I'm not sure, especially for men, that they're even aware that this gap exists. And I think even for a lot of women, they're never exposed, they never know that there's this whole field developed around this because their primary career doesn't have time to even get into it.
B
Absolutely. And I was going to say, I think that two things. Interesting you pointed out whether you believed people would actually fill in that many questions. We have run experiments. You're talking about your classic E comm journeys where you want to reduce the number of clicks to purchase because that drives sales. But actually, what we found when we increase the number of questions, and we made sure that it was a mandatory step before actually purchasing the test kit at all, that we've seen a massive uplift in conversion because probably for the first time a lot of these women felt heard, felt seen and that this was something that was being taken seriously and it's not just an off the shelf random purchase. And then secondly, from a clinician point of view, speaking particularly about the uk, there is so much pressure and demand on NHS services that the clinicians cannot possibly deliver the level of care that they would like to. Particularly in primary care, where your appointment a. It's impossible to get a booking in some areas and when you do, you only have 10 minutes with a professional who is supposed to be a generalist and therefore would not necessarily go down the same route of questioning as our assessment does. And for the woman not having the confidence or understanding to know what she should even be coming to the appointment prepared with. So actually, what we are finding with our research with these AI tools is that not only do they provide efficiency gains from a commercial point of view, which is great, but actually from a clinician experience perspective, from a clinician burnout point of view, how happy they feel doing their work. Are they able to provide that quality of care that they set out to when they went out to become a doctor, that we're able to drive improvements there as well, because women's health outcomes are not going to be improved solely by testing alone. You also need to take the clinicians with you and for them to feel a part of that process. So that's something that we're looking to improve with these products as well.
C
You mentioned that women, that it's. The women are complicated, so to speak, which is why they've not been included in clinical trials. To a certain extent, that is true because your cycling hormones and everything do really. They really change over the duration of your cycle. And that's. And then again, that's even more unpredictable if you don't have a regular cycle, which a lot of women don't. And so whilst I obviously don't agree with it, I do have some sympathy and understanding of how we've ended up at this point where there is such a gap in data and I think that kind of brings out the benefits and the power of a take home test like fertility offers. Because a lot of the things that we test for in your blood can only be done on day three of your cycle. And so if you're trying to book into a doctor's or a hospital to get those, to get that blood draw done and you're trying to manage an appointment and you only have one day that you can do it and maybe you've got other things going on at the day, at the time that they offer you, it becomes very difficult. Whereas if you can buy an at home test ahead of time, you know it's there, you just buy it, you wait for the right time in your cycle, you test it and you ship it off, you can get much more reliable results. Which again adds to the power of the data set that we've got.
A
Yeah, I think you're highlighting something that's pretty important. I don't think that research gap was malicious. Right. I think it was scientists trying to do good science and removing a variable that really confound a lot of things. Really unfortunate consequences from doing that in terms of just women's health being understudied. And I think it's why today we still see lots of studies where women aren't included because we haven't found a good way to account for that. What I wish is that we just then also had the similar study with just women.
D
Right.
A
Let's look at both sides of the human population. Okay. I heard Jack say that this, your company has been around for seven years, or maybe it's your product has been around for seven years. I can, I can imagine just your survey and the blood test, like the data that you're collecting on its own must be incredibly valuable for both the patient and the clinician. Is that what your product did before AI? Was it just all about collecting more data to educate both the patient and the clinician? Tell me a little bit about how did this work pre AI to set the table for then how you're now using AI?
B
Yeah, I think the process in general first of all was just to collect the data. In the first, our CEO and founder, Dr. Helen O', Neill, she wanted to run a research study but wasn't able to, so set about to collect that data herself initially. And then that's where the kind of product idea came from. And from a data collection point of view, what it's allowed is that in the back end we have this Observability anyway, in terms of the trends that we were starting to see. And we can discuss that with the clinicians who obviously review every online health assessment response and every blood test result before writing the report and start to identify some of those patterns. So pre AI, it's more like a templating process where, you know, typically for these types of patients you would tend to see these types of results. And so to reduce that workflow administrative burden, you can create these pre existing templates that then the clinicians can go on and edit themselves in the kind of back end before it's published to the patient. So it's like a kind of manual editing process that it was.
D
We have to add a couple of bits on that, what we were doing pre AI and to extend. We're still in the pre AI, which we're just starting to build with AI, but we're trying to reflect clinical best practice, helping clinicians who will review every blood. Blood result before we publish it to the user. We're just trying to reflect clinical best practice for diagnosis of. I think it was 25 different. 25, or is it 18? 18 reproductive health conditions. The other thing I wanted to say was a clinician will review and provide a diagnosis against every report. And that is a combination of health assessment, which is your symptoms, how you're feeling now and the blood results just like a snapshot of your biological reality. And so we have this incredible labeled data set where a clinician has gone in and labeled each combination of health assessment and blood result with a diagnosis and the power of that and then being able to build predictive diagnosis. And as I said, we're trying to do now is reflect clinical best practice. And the opportunity is actually to improve clinical best practice and actually look at the guidelines and look at how we can go even better than they are now.
C
Yeah. And I think as you say, we've been collecting this data for seven years and we are starting to develop AI products with it now. But as well as that, we've always been doing research and publications using this data. For example, our most recent paper that was published last week or the week before was on testosterone levels and what typical reference ranges are for women. Because the standard understanding of how testosterone level should be in a average woman, so to speak, is based on a handful of people in a lab decades ago. And that was like. And no one has updated the research since then. But then we were like, we. You test more testosterone in a day than was used in that original paper. We think we can do better so we've been able to update that with a more age specific regression model based so we could share back and give a more informed view of different analytes and how you should expect them to behave in women in a way that we couldn't do before. So yes, it is all about the AI, but alongside that as well, we're trying to give back and use our data for good in that sense and contribute back to the academic community of things that we're learning as we're collecting this data.
A
Yeah, my brain is going a million different directions and I'm going to pull it back from the very political topic of testosterone levels, particularly like in transgender athletes, where I'm sure we're making horrible policy decisions based on like very unreliable data. So I love hearing somebody out there is collecting better data and actually publishing it and sharing it with it. Okay, let's talk a little bit about. Okay, so it sounds like you've been collecting really rich data. That data is being whether it's survey responses, blood tests. That data is then shared with a clinician who is providing a report to the patient. So the patient is just learning a lot about their own sort of reproductive health. And then you've been sharing with the academic community what you're learning. You probably have a very unique data set that's not available anywhere else. Tell me a little bit about. Let's get into your first AI product. So what problem were you trying to solve and how did you start to evaluate that AI could help you solve that? Yeah.
B
So the first AI product being gyne AI is exactly that flow that you just described. And as Jack mentioned, being able to correlate between seemingly unrelated symptoms and results and an eventual diagnosis. If you think about the typical time it takes for a woman to receive a diagnosis, particularly for conditions like endometriosis, in the typical healthcare system, I think it's nearing up to 10 years from raised through to diagnosis received and then from there she then has to continue to fight there any kind of onward care that's necessary to manage that condition because it's something that's chronic. So with the data that we've collected, the development of gyne was to try and reduce that time to diagnosis because if you have enough data points that can provide you with an indicative diagnosis earlier on in the process. So say we'll have a certain percentage likelihood just from the online health assessment responses alone. You, you layer the blood test results on top of that. That gives you an even further percentage likelihood of receiving a certain diagnosis. And then if you do go on to have the onward care with the scan images, et cetera, or any other consultation, even the discussion that you have with the clinician that layered on top provides you an even a greater percentage likelihood of that being the diagnosis, you can increase that pathway to referral with the relevant information for that diagnosis to be received in a much shorter timeframe. So that was the problem that we were trying to solve, A from the woman's perspective, but B, also for the clinician that they are able to provide that kind of care with confidence much earlier in the process, where without that indicative percentage likelihood, they're having to follow these guidelines, as Jack mentioned, that were written a million years ago and were quite loose in terms of the criteria that you need to meet to be able to refer someone onward.
A
So it sounds like there was maybe this goal of reduce the time to diagnosis. That feels really big. How did you get a sense for what takes so much time? Like, how did you identify even like what needs to be resolved or what needs to be supported in order to have an impact on that outcome?
B
I think if you think about the typical UX journey and we're talking about the UK and Ireland at the moment, so that kind of healthcare setting, the biggest issue is that triage, where, as I mentioned, in primary care, the likelihood of you reaching a point that the collision can confidently triage you in 10 minutes for it being a specific gynecological issue is very slim. You typically have to have upwards of 10 appointments at a primary care level for even the picture to be complete enough to be referred. So that online health assessment removes friction with needing to go in person to an appointment, it's way more accessible. You can do it in the comfort of your own home. For communities that have typically had problematic experiences with the healthcare system, it removes that fear, that stress, that anxiety factor. And people tend to be a lot more open actually with a screen than when you're looking somebody dead in the eye and having to divulge quite personal and sensitive information. So we've really taken the burden out of that triage process through the online health assessment alone, which is typically the longest period before you're even able to be referred for bloods. And because that blood test process is automatic off the back of completing the assessment and you've already completed probably what would be the longest cycle in a typical health care setting in a matter of days as opposed to months or years.
A
Yeah, that makes a lot of sense and especially the comment about some of these areas are probably are not easy for people to talk about with another human. And so I think we've seen a few instances on this podcast where the, the anonymity that's almost created by interacting with AI allows for richer, more honest interaction, which is surprising. I think originally my bias would have been towards people don't want to talk to AI, people don't want to give feedback to a machine. But there are some instances where this is actually a really big positive.
C
I think it's really helpful for the doctors as well, because if you take it from the clinician side, they have multiple appointments a day and they only have 10 minutes to see you. They're context switching a lot between different patients, different ailments, and you're expecting them to be the expert in absolutely everything, which is a very difficult job for them to do as well. So I think the advantage for the doctor of having the AI assistant here as well is that it can pinpoint you to what is the most important information about this patient and about this person and the things that you can look at and suggesting the diagnosis where before with the greatest will in the world. I'm sure doctors aren't out to gaslight women or disbelieve them or anything like that, but I get that there's also some challenges with their job that makes this an inherent problem as well. So with the AI, you can help focus the doctor and give them the precise information that they need to help the patient suggest a follow on care, suggest a diagnosis that you would not be able to do otherwise. So it's useful for the patient as something to talk to. It's useful for the doctor on the other side for something to base their next decisions off.
A
Yeah. Okay, I want to dig into this a little bit. So, Tulsi, I think what you were describing, I didn't hear a difference between your 150 question survey, although Lorna may have just touched on it. Is the way that patients interact with the questions different now that there's an AI version? Is that like one piece of it? And then Lorna, your comment about we can use AI to present what are the most relevant factors? We may not have time to read all 150 questions. How do we steer you towards the most important data that seems absolutely like even maybe a machine learning or AI step where we can say, here's what's interesting here. Do you want to respond to both of those?
B
Yeah. The patient experience is exactly the same. So we're running these two things in parallel because we need to prove the efficacy of our AI products in order to get them certified and in order to scale them properly. So in terms of the patient experience, it's exactly the same. You would go onto our website and you would complete the online health assessment. There is conditional logic that sits within the assessment that we would ask you certain questions, if you select certain responses for the question before, etc. But in terms of the workflow, it's exactly the same. Where we then can typically look at whether or not the AI is having an effect is that clinician step in terms of actually assessing their responses versus having something summarized and to just approve or deny, essentially. But I'll hand over to Lorna to talk about that step.
C
Yes, the AI, so to speak, that we're using here, what we found to be most useful here, is actually something that's very transparent and understandable and gives a clear reasoning between inputs and how they fit used in the prediction. What we've actually, what we're actually building on at the moment is essentially a Sabbathian network, which is like a kind of causal directed graph based on things that we believe are linked to causes and consequences of other things across
B
the
C
whole spectrum of information that we collect on diagnoses, blood tests, symptoms, anything that you answer in your health assessment. And then we use that to be able to. Plus our information on how people have typically presented with these symptoms and everything over the time that we've been collecting data and the best knowledge that we have from clinical guidelines and everything to make sure that what we're doing is also going to be trusted by doctors. And so from the doctor's point of view, what you get on the outside is not just like a binary. This person, the doesn't have hypothyroidism or pcos now pmos, it's just been renamed. What you get is that you have a. This person has a 73% chance as having this diagnosis. And here is why. Because then that's useful for the doctors, because then it builds trust with the people who actually are going to be putting their name against the potential diagnoses that we're offering to people because they can understand it and it helps again, as Tulsi mentioned, with the triage process of really narrowing down. Okay, what is the useful information from the whole picture of someone's health that we've got to be able to give the person the answers that they need in the report. And it's also hopefully going to be helpful to the patients when we share that back. Because what you're getting is not just, oh, here is a label, go now, research that yourself and do something with it. But it, here is why we think you have this. And therefore, here is our recommended next steps for you. Whether it's a scan to investigate your symptoms more, whether it's a nutrition plan to manage something else, one of your particular symptoms, whatever it is. I think having that transparency in a medical setting I think is what also helps really build the trust both from the doctor side and the patient side of what we're doing with our diagnostic model.
A
Yeah, okay. I really love this because I think we have a lot of research that shows, especially for diagnosis decisions, Bayesian thinking is really critical. But I think most clinicians don't have. They're not statisticians, they're not data scientists. They're not really that well versed in Bayesian thinking at all. And so it's. And for probably many of our listeners aren't, it's a very hard thing to wrap your head around. Like the. I forget who it might have been one of the Atul Gawanda books, but he talks about a really salient example of this from a patient experiences. Let's say a woman gets a mammogram and there's a problematic finding and immediately the woman thinks she has cancer. But if you apply Bayesian probabilities to it, maybe for that type of finding, only 10% of women end up having cancer. And doctors don't understand these ratios enough to communicate that to the patient. They just communicate you have a finding that means you need a biopsy. And so patients spend an enormous amount of time thinking the worst case scenario when the probability of that worst case scenario is now higher than if you didn't have that finding, but it's still really low. And I think this is an incredibly. Humans don't really understand probability. You tell me 73% chance that sounds really high, but there's still the 27% where it doesn't have that. And then Even in that 73%, there's a range of outcomes within that. And I think this is a huge gap for both patients to understand, but also surprisingly for clinicians to understand because they're not well versed in data. That's not their job. Right. So this raises a number of questions for me in that view to the doctor, when you tell them that the diagnosis is 73% likely to be this, are you also giving them other potential diagnoses? Is there anything that's helping the doctor understand that data already? This seems like a giant step in the right direction. But I can also See, I can't tell you how many times I've heard a surgeon say 100% outcomes positive. And you're like, okay, I know that's not true. Like, how do you just bridge that kind of data gap for the clinician?
C
Yeah, so we do everything that we would give us a suggested diagnosis on, we would report on all of them. And the percentages are not totaling 100% either because often what you find is that there's more than one thing going on for women anyway. So you may have endometriosis and PMRs, and that's fine. They're two completely different systems. One's metabolic, one's an autoimmune type 3 condition that's affecting your endometrium. There's no biological reason why you can't have both. And so within the models as well, it's not like a multi classification either or in these situations it's like we're assessing them all independently based on the picture of symptoms and information that you've given us so far. And so it's possible that some people will come out with no indicative diagnosis. And we say, look, everything's in range, everything's fine, good for that person. There's also the case where you could find four or five potentially concerning things that you might want to recommend people to follow up on. And yeah, the way that we do it accounts for both situations.
A
This sounds like a really interesting UX challenge because it's almost like you could create sub clusters of these things. Could be this, these other. If you group them this way, it could be this. It's almost like you're giving a clinician a map of a territory to explore and it's still up to that clinician to do the follow on, actions to figure out. And every new piece of data changes the map. And maybe it converges, maybe it doesn't. Because humans are messy and data is messy. But already I can see just having
D
a map is helpful and it encourages, it's a really good point because it encourages some more nuanced thinking that is needed to make a diagnosis. With the existing guidelines are set up, we have, it's quite binary. If you hit a certain kind of number of criteria, blood results, symptoms, et cetera, medical history, then you tip over the threshold and that precipitates the diagnosis. Whereas what we're trying to build is something a little bit more nuanced and give the clinician the information they need to make a good decision to make a diagnosis. So you get that percentage like 73% 76%. So it's not just 1 and 0, it's somewhere in between, which I think probably reflects normal life a bit more.
A
I think that this also is really good for the patient because I know it's so easy to fall into this thinking of, like, medicine has the answer, whereas so often we don't. Like, it's very gray. And it helps both the clinician and the patient really understand the map, understand the territory, like what we're going to go explore and figure out what this could be. And I feel like that's very. I don't know if it originates in medical school, but my experience with most clinicians is they communicate in a very certain way and it doesn't really match the uncertainty of maybe what the data supports.
B
Absolutely. And I think just to touch on what you mentioned around human understanding of this type of topic, and it's not the clinician's job to be a statistician. We are developing these products, obviously for the end user, for the customer, so that they're receiving this information in an appropriate way. But equally, everything's being developed with clinicians. We appreciate that we have all this data and we think it's going to change the world and we think it's going to be amazing. But not if it's not adopted by clinicians in the correct types of workflows, not increasing the administrative burden on them. And also from their point of view, that they feel that it's data that they can trust and confidently move forward with that patient's kind of care plan. So all of these products that we're trialing are being developed with clinicians alongside us to ensure that they will actually unlock burden that we're trying to overcome. And the clinicians actually ultimately find it useful. We think it's exciting, but they also need to be brought along for the ride. Yeah.
A
I also want to highlight there's a really valid reason why clinicians communicate in certainty, even if they know it's not certain, and that there are humans out there that hear a 97% chance and they lock onto the 3% and think, I have to do nothing. Right. And so part of a clinician's job is to communicate in the way that they think the patient will do what's best for their health. And I know that this leads to a lot of gray and a lot of depending on the patient's personality and ability to process statistical data like that match has to really fit. For some people, communicating uncertainties is the best thing for their health. And for other people Communicating in the probabilities like they have the capacity to still make a good health decision. And Most clinicians have 10 minutes. They can't even assess the patient's capacity. So I don't want people to take any of this as me criticizing clinicians. I think these are systemic symptoms of just the way health care has evolved.
C
Yeah. And also here, everything that we're doing is designing to be a clinician aid rather than a replacement. I think we're always, at least for the foreseeable, we're going to have doctors in a loop in every patient decision that we make. It's just about how do we make that doctor as informed and as useful as possible and to give their assessment in as efficient time as possible.
A
Yeah. Okay, so this is great. I immediately see the benefit of just, let's add a little bit of Bayesian thinking to the diagnosis step. Now, my understanding this was just your first foray into sort of some intelligence in the process. Tell me about, are there other pieces of this first version that we haven't covered or should we dig into the second one?
B
I think we're good to dig into the second product, which, yeah, I find it the most exciting because I think it's as amazing as the Bayesian network work is. It is all very backend. So I think it's less tangible than the scan automation products that we're also trialing, which for the laypeople like me, it's more something that you can see, touch and feel. So again, when we're thinking about clinician workflows and that typical time to diagnosis, the burden that's placed on clinicians where they're having back to back appointments and they need to review people's results and write letters. And the propensity for human error is really high. From a litigation point of view, that's quite scary. From just a medical point of view, giving the wrong information or misleading information to a patient is obviously problematic. So we have also been looking at our pelvic scan process. So typically for most women who will come and see us, they will be recommended to have a pelvic scan, even if that's not for a diagnostic purpose. If you're typically suffering with symptoms that are unexplained, a pelvic scan is a good next step to get a physical view of what's going on. We work with scan providers to do this. We're not a scan provider ourselves. And so we are getting images from multiple different partners that then need to be received, reviewed and then sent to patients alongside a letter that is written by one of our clinicians, again taking the online health assessment responses, the blood test results, because typically most people will have done a blood test result before having their pelvic scan, and then the pelvic scan image review to give a comprehensive overview of what could be going on. And that process was, or still is, to be honest, for a lot of our providers, quite fragmented in that there's not a central way for us to receive those images. For the clinicians, they're having to open different windows on different screens and drag and drop things into different places to make sure that the patient's receiving the information. And they're typically trying to do this in between appointments, also writing letters where there is a likelihood that a mistake is made and there aren't necessarily the right guardrails to ensure that the patient is receiving their information and that it is correct. So we've developed and are trialling a scan automation flow where we receive all of these images into a central system that then can be automatically reviewed. It will create a summary table for the clinician with all of the findings, all of the objects that are in the scan, the parameters of those. They can then click a button and essentially generate a letter for that patient, which will need to be reviewed by them to ensure that the content is appropriate and accurate. And then it's sent to be approved by a second clinician before being published to the patient. And we're trialling this with scans at the moment because it's one of our highest volume products with all of these kind of administrative and operational issues that exist across the flow. Then you can see how that process could also be applied across other telemedicine services that we offer, where that scan step isn't necessarily in place. But we have the blood test results, the online health assessment responses, potentially a transcript of a consultation, and all of that pulls into a AI generated letter that's reviewed before being published. So that's the end to end flow from a scan point of view that we're looking to improve.
A
And am I understanding correctly that you're doing the image analysis or is that coming from whoever provides the scan?
B
We do that in house through our models in the back end.
A
And I assume that the provider that is giving you this scan, do they have a radiologist that has looked at it and provided their analysis and your image analysis is on top of that, or are you replacing. They're just collecting the data and you're doing all the analysis?
C
It's a little bit of both, yes. There is a little bit of both. The current way that we would be getting the scan images is just in a zip file where the images are not ordered in any particular way or they're not labeled in any particular way. So the first job for the clinician would be to be like, okay, this random hash file name. What's in here? Okay, it's an ovary. Call next one. Okay, that's also an ovary. And there's a lot of admin work that has the steps. So I think the first step in the image processing for the machine learning is classifying those images. So you can. So in the ui, we can easily lay out, okay, here are your Doppler images, which measure blood flow around the endometrium. Here's your left ovary, here's your right ovary, here's this type of view of your uterus, et cetera. Because then it makes it so much easier for someone to look at it and be like, okay, I know what I'm looking at now. That sort of thing of admin off my list. And then there's some steps within the images that we're looking at, like picking out, like, we're detecting different objects in there. So where the follicles are and the outline of the endometrium, for instance, on the uterus. And then we're adding contours around that with a, like, third refinement on the image processing to draw around the boundary of an ovary, let's say, or the endometrium. And then I think this is the bit where it's where we think we can do slightly better than the radiologists when they're looking at the images. Because when you're performing the scan to try and measure the volume of an ovary, for instance, you're taking like, you're trying to draw two axes on it to create, like, an ellipse. And then you calculate the area of that ellipse and be like, okay, that's the volume of your ovary. But in reality, ovaries are not nice ellipsoid shapes on scans. They're much more complicated. So if we can accurately draw a contour around the ovary, our calculation of the volume of it is going to be much more precise than what and what a radiologist can do with the tools that they have available to them. But then in other cases, less so, I think. So one thing we're toying with at the moment is like, your follicle count, because almost every scan will include how many follicles there are present across both ovaries. That's a good indication of your egg health, your egg reserve, and possible indications for polishes to go freeze and things like that. It's very difficult from the AI image to see which we can count the number of follicles that we can see, but that's not necessarily the same thing as those are the number of follicles that were present, because there could be one multiple slices of the same image, and you're looking at the same follicle slice on an ovary just because of the way that the images were taken. So in those cases, we would defer to the sonographer or the radiologist that's done the scan to say, okay, I counted nine follicles here, for instance, and pulling those out from the notes. So I think it's a little bit of both in terms of information that we think we can get better through our machine learning processes on AI because we can do it in a more sophisticated way than is available when you're doing the scan, plus information that really you need the proprietary information from the sonographer who was doing the scan at the time to tell you the story of what's going on in the images. So we've used both together to get the most accurate picture that we think of what's happening in the scan.
A
Jack, was there something you wanted to add?
D
Yeah, I guess what I was actually thinking about was the step before you do the image analysis, which I guess as an engineer, that was one of the kind of most interesting challenges we have with this project, which is how do we get an image that a sonographer's just made Using their wand. Sam wand piped up to wherever they saw it in the cloud and then into our infrastructure in a way that's secure and that has as few human clicks as physically possible. And we were able to do that. Annoying. There's still one human click in the process, and we're working to remove that human click. But that was a really fun infrastructure challenge. We had to be quite innovative in the way that we set our platform to then be able to receive the images. And the images are called dicom. The format is called dicom. They're quite sort of complex image type. There's a lot of metadata, so there's a lot of work to both get the raw images in and then feed that through to Lorna's team so that they can pass and process and analyze them.
A
Yeah, there's a lot of interesting bits to this flow. I wish we had two more hours to dig into all of it. One thing I want to highlight that I'm hearing in your story, and I've heard it consistently across every little detail that we've talked through, which I really like. So my background is in teaching teams how to do good product discovery, and I anchor it around this idea of start with the outcome, discover unmet needs, pain points, and desires, and then match solutions to them. And I'm hearing this structure through everything that you say. So a very clear outcome of reduced time to diagnosis. A lot of really rich detail around understanding your customers needs. Even Lorna, the level of detail of what a radiologist is doing to estimate the volume of an ovary or whatever they're trying to estimate and the assumptions there. And that's an area where the AI can actually do a better job because machines are better at calculating irregular shape volumes and that humans have to use a crude approximate approximation. This, like, level of detail of understanding the process and really designing targeted solutions for what you're uncovering. It's rare that I hear product teams talk about this at such a depth of, like, empathy for their customer and really just matching the right technology solution to the needs that they're uncovering. So I just wanted to acknowledge that I think your story is pretty amazing, and I'm sure we could dig into, like, many more pieces, but I want to zoom out a little bit. It seems like there's plenty of opportunity here to use technology to aid the humans in the process. And it sounds like you're finding success by really understanding the humans and what the humans are doing and then asking how can the technology help? Whereas I think a lot of teams reverse that. They just get excited by the technology and they start throwing technology at the problem. I think by having a clear understanding of what the humans need probably sets you up for what I want to dig into next, which is evals. So how do you know that your technology is doing the right thing? And I can imagine in this context it's really critical that the technology help and not create new problems. So I'm curious just to hear how you think about that, what role evals play in the process?
C
Yeah, so I think, I guess in different situations, yeah, we'll have different evals that we focus on. So going back to gang AI briefly, like, in that case, it's relatively easy to start with, to know we're doing the right thing because we have a nice. We have the benefit of a nice, expansive, labeled data set. So that makes it easy for us to be able to validate our models. And make sure that the expected probabilities that we're coming out with roughly in line with what a clinician would agree with. I think one of the challenges we're going to face there as it gets deployed and gets used a lot more by clinicians, is like automation bias. Because what you don't want is we could sit there and do nothing and it could look like our accuracy is going up because more and more clinicians are agreeing with what we're saying. But that doesn't necessarily mean that our model's improving. It just potentially means that clinicians are like, cool, I trust this. Now I'm just going to default to it and not think myself, which is a problem that we're going to face. And I think how we plan to mitigate that is we would always try and manage holdout sets, try and have new clinicians, fresh faces, independent eyes to just check what we're doing and use that as a basis for deciding where we need to retrain either the clinicians or the model or both, and how to work in that case. And then I think on the scan side, similar, I guess in terms of the image processing, we do have a labeled data set in terms of what's on the images and the contours, which took us a ridiculously long time to get because it had to be hand done by people who both could understand code enough to create the label the images that we need to, and also understand ultrasound scans enough to be able to reliably do it. But we have that now and then again, similar, I guess we would be evaluating how close that matches to the data set. And part of what we want to release with the scan automation tool as well is the ability for doctors to change contours and relabel images anything they're not happy with. If a classification is wrong, or we think we've not got the outline of a particular organ quite correct, the clinicians will have the ability to adjust that. So that's helping improve our machine learning over time. Or like an independent researcher outside of the clinician flow, would be able to do it as part of the product. I think that was really important so that we could be confident going forward that it was doing the right thing. We also have on the letter side, which I know we've not really touched on yet much, but as well as the image processing here, we are using AI to draft a clinical letter based on the findings from the scan. So once the doctor's seen all the nice images we have and made their assessment, added their Notes that they need to, to anything over and above what we've already done, added any potential diagnoses, we will use all of that information to draft a clinical letter, plus all of the information we know about the patient so far. So if they've had blood tests, that goes then the health assessment, trying to pull all of that in to give like a really rich picture of okay, what's going on for you. And then, yeah, and then that goes back to the clinician to review and then because QC'd and everything before we go out. But yeah, the eval process there looks interesting because we've, yeah, we're working on building for that an agentic loop that does like a few kind of AI checks before going to a human to check. So for there we're looking at, we're pulling in all of this information into the letter. It's a reasonable risk that your LLM could hallucinate while we're accidentally pulling in context from a different patient or a different letter in the way that the model has been trained and developed, which obviously we don't want to present to the patient, we only want to present what's going on for you. We have a kind of agentic loop step here where we're checking the content of a letter versus the facts that we know about that patient. And if they're not right, we're like, okay, let's send it through the loop again. Let's regenerate that letter. And we keep doing that until we get to the point where we know that the important information about the letter is correct. And that's how we're adding our own guardrails. And we're finding that the LLMs is actually quite good at that particular thing because they can be sometimes a bit rogue in terms of what they write for you and what they generate. Using an LLM as a checker is often like really effective that we're finding. So in that process we're like building trust with, for ourselves and with the clinicians that any draft letter that they'll be looking to review has gone through this rigorous process so that we make sure that it's accurate.
B
The only thing I would add is something that is quite critical for us in this highly regulated, highly sensitive space. And although we're using AI to speed up a lot of that process, it's go fast to go slow. We will be having a two step clinician review process before it ever reaches the patient. So if you do have that moment where the clinicians are becoming reliant on The AI and potentially not critically evaluating what they're seeing. There is always going to be a secondary person who is reviewing the review before it goes to the patient. And typically they haven't been involved in that patient's care up till this point. So it really is a fresh pair of eyes from the right area of expertise. And that's also like a physical guardrail that we can put in place in the product process because we have optimized the flow so much, it doesn't then actually affect it too much to have that two step. That two step review.
A
Yeah, I like this. It's like a very thoughtful way of thinking about human in the loop. How do we separate this first step and then have somebody with fresh eyes to look at that then Lorna, that loop that you described, this is a very common theme of how do you guard against hallucinations? How do you make sure everything's grounded in the data of having one model generate the output. So in this case, your letter and then having a second call really verifies everything in this letter grounded in the data. And that that back and forth or that iterative loop is quite effective. And surprisingly, you would think that second agent would have just as many problems or that second run would have just as many problems. There's this theme around having the agent show its work, right. So don't just say yes. Everything's excited but everything's cited. And here's where it came from. That really has been a pattern that's come up on episode after episode on our podcast. And so it's nice to see like teams stumble on these same patterns that just get us to much more reliable output with the LLM.
D
Yeah, totally agree. And maybe just a couple of more points on that one is that it's really interesting to see the data in that it's rare to get a third loop. So often if there's a first in the first loop, you get something wrong, the checker will get it right. That second loop, it rarely gets to the. Occasionally it does. So that's that really cool like live eval pipeline to make sure we're sending out the correct information to the patient. The second really useful thing is it gives us great data points then feedback to think how can we make changes to the harness to then improve the output and we can measure the number of loops over time and check to see whether a change to the harness improves. Improves the accuracy of that first go.
A
Yeah, I love this. I see the exact. So I have a very similar repair loop in one of My products and I see that exact same thing. It's one thing to fix the problem in the repair loop, but to also see what problems continue to emerge is a nice feedback loop to like, how do we solve those upstream so we go through the repair loop less often. It sounds like you're very thoughtful about understanding your both the patient needs, the clinician needs, you've got evals in place. I know for a lot of teams in healthcare, evals are challenging because of the sort of pii, the health data, you're getting scans from third parties. I'm just curious how you're able to do this. What has to be in place to make this doable. And I know a lot of teams are like thinking it's impossible and not even trying it, but we have talked to a few teams in healthcare on this podcast now where they're finding ways to overcome that and still follow all the regulations. So I'm curious what your experience with that has been and just what the learning curve to figure out how to make this feasible has been.
C
Yeah, I think one of the big principles is just trying to data minimization and then pseudo anonymization as quickly as possible. So for any, there's no need to put in personally identifiable data into a diagnostic algorithm or a letter even, because that's if anything, that's just noise that just adds. That makes the AI's job harder at trying to do what it wants to do. In the case of the gang AI, all of the data is pseudo anonymized and it's less of a risk anyway because the machine learning models are like, we're not putting them in, we're not outsourcing it, putting it into someone else's, waiting on someone else's cloud. They're all built and trained in house, but even then, nothing in there. There is health information in there, of course, there has to be. But none of it is alongside anything personally identifiable to people, which is important for our regulations and everything that we stick with that. And then in the case of a letter too, like all of the personally identifiable bits of a letter are at the start and the end of it. Really the body of the letter is what we're interested in getting the AI to generate. Yeah. So we just, again, we just chop it out and we're like, just focus on the information that we do. And we would never pass someone's name and date of birth into an LLM in which to draft a letter. We would just add that on afterwards after we've got the AI to do the body of the letter and then make it into an actual letter afterwards. So that's one of the big ways that we're managing it in that sense.
B
I would also say from a kind of. Yeah, from a, from an overall how have we gotten here Perspective, I think is don't ignore the regs. It's the longest and most difficult part of the process for any health tech. The earlier you do it, the quicker you can move. At a point where seven years ago, the type of technology that we're working with now didn't exist, exist in the same way. But we were steadfast in no, we must be properly regulated, making sure that we have all of the I's dotted, the t's crossed in order to even collect this data set in the first place. And equally, when working with partners, they need to meet a certain standard. For example, our laboratory partner, our scan partners, we don't just work with anyone. And, and again, that's to ensure that we never stumble over a regulatory hurdle because that is the thing that will stop you in your tracks from actually launching something like this at scale. The technology is there for people to do this without any of the regs, but actually the adoption en masse is only possible with the regs. So I would say that's something that if other teams are thinking about moving into the health tech space, don't leave the regs as a last minute scramble. Think about it as part of your product development process from the beginning and it might feel like you're moving a little bit more slowly, but actually what you end up developing is something that's really defensible, really scalable and ultimately investable as well. Because the first thing an investor is going to ask you is so can you scale this? And if you don't have the regs, you can't. So that's also something I'd say is really important.
C
Yeah, agree.
B
It's good.
C
It keeps you honest as well, to a certain extent. Because if you're always thinking about, if you're building things at the start with the regs in mind, you know that you can't cut this corner or you know that you have to think about where your data is stored here and who has access to it and all of these things, which are a lot harder to unpick afterwards and with everyone's greatest will in the world, is not maybe not something that you would think about as you're building the machine learning product to begin with. So always having the regs in mind means that it's easier to scale. Yes. But also it is going to be more secure and it's going to be better for the patient that you're doing things in the right way as well.
A
So it sounds like for the Bayesian model you built that in house, there's not a third party involved. I'm guessing for the scan image processing, are you using a third party vision model? And then for your letter, are you using a third party model?
C
For the image models we're using, essentially we're building on top of like pre trained Pytorch models to do different things. So our versions of the models we've built in house. And so no like.
A
So you're not sending it off to like no. OpenAI's vision model.
C
Exactly. Yeah. It's all trained in house and actually that is one bit of like personally identifiable information is difficult to remove. Is like the name, because every ultrasound scan has the name and the date of birth of a patient at the top. But also we found that we want to train the models on slightly messy data to try and account for every possible type of scan that we come in. So we do a lot of slicing of the images, slightly distorting them and using that in the training process to get better results. But it also, it's a good way to train the image models, but it also has the slightest advantage that you're often distorting the text and the patient's name before you put it in the model anyway.
A
Okay.
C
But again, yeah, that's all in house. The bits where we. Yeah, with LLMs, obviously we're not, we don't have our own proprietary LLM that we're using to train, but we, we're using everything that lives within our own infrastructure on aws.
A
So Through Bedrock.
C
Exactly. Yeah. So it's not. Which gives us like slightly more protections and like more security and overview because it all comes under the regulation and the governance that we have with our AWS platform in general, rather than having to do all that due diligence again via OpenAI or anthropic or which if we were sending it off to someone else.
A
Yeah. It's funny, when Bedrock first came out, I was confused about why it needed to exist. But the more I work with companies where PII and especially PHI is such a big deal, Bedrock is a no brainer. And especially recently with so many European companies rightfully wanting to keep their data in the eu, Bedrock's like the only way to do that. So it's clear that AWS found a need and is doing a reasonably good job at filling it, which is nice to see.
C
Yeah, for sure.
A
This has been amazing. It's been really delightful to hear more of your story. It's really clear you're a very thoughtful product team with a good understanding of your problem space and really focused on solving real needs and not getting distracted by the technology, which again, I'm going to reiterate, I think is really rare. I also hope that you expand to the US One day and that if you do, you let me know because I would love to try it out.
B
Thank you. Watch this space.
A
All right. Thank you both. I really appreciate it and give Jack my regards as well.
B
Will do. Thanks so much.
C
Thank you so much.
A
If you enjoyed this conversation, please subscribe in your favorite podcast app and give us a rating as it helps others find the show. Thanks. I appreciate it.
This episode features Teresa Torres in conversation with Tulsi (Director of Product & Technology), Lorna (Head of Data & AI), and Jack (Head of Engineering) from Hertility. They discuss building AI-driven products to close gaps in women’s health: notably, their work to combine Bayesian diagnostic models and automatic pelvic scan analysis, streamlining both patient and clinician experiences. The team digs deep into how rich, underutilized data sets can speed up diagnosis and improve care, all while meeting strict regulatory requirements and maintaining a human-centered approach.
Key Points:
This episode is a masterclass in problem-driven, human-centered AI product development in healthcare—and a must-listen for teams attempting to bring advanced AI into regulated, high-stakes environments.