
Loading summary
A
Imagine reaching into a desk drawer to find a specific key. Your hand brushes past loose paperclips, old receipts and disconnected cables. You know the key is in there, but the items are scattered, forcing you to rummage through the clutter to find what you need. Today, a parallel situation exists with personal health data. Individuals carry highly detailed physiological measurements, things like heart rate, sleep cycles and laboratory results directly in their own pockets. Yet when attempting to answer basic questions about physical well being, the experience remains identical to rummaging through that cluttered desk drawer. The data points are isolated, making it difficult to construct a coherent picture of health technology. Companies are now deploying artificial intelligence tools to organize this data. OpenAI's recent US launch of Health in ChatGPT allows users to connect their Apple Health and medical records to a large language model. To the general public, this appears like a very useful, straightforward feature launch. However, an analysis of the official launch documentation reveals a highly sophisticated exercise in regulatory navigation. We've previously considered some of the implications of large AI companies applying their products to healthcare issues, so I won't repeat over that again. I'd recommend that you look back at those previous episodes, but today we're going to look carefully in detail at what looks like a casual collection of user testimonials and product descriptions, but is in reality a very carefully engineered framework designed to really carefully avoid classification as a regulated medical device while signaling maximum clinical utility to consumers. In this product launch in most, if not all countries, software that actively diagnoses, triages or treats clinical conditions is classified as software as a medical device. This classification triggers very rigorous multi year regulatory review by bodies like the FDA or MHRA in the uk. To avoid this pathway, the language in the launch documentation is very carefully calibrated. A prime example is how the document uses early tester testimonials to introduce high risk clinical capabilities. Instead of OpenAI themselves claiming that AI can perform things like clinical screening, a testimonial from a nurse is introduced. The tester notes that the tool helped identify an unexpected chart entry, allowing for appropriate follow up. The this is a specific regulatory maneuver. If the software developer officially claimed that the tool could actively identify clinical anomalies, the software could be classified as a diagnostic screening device. By presenting this outcome through a user's subjective experience, though, the developer introduces the concept of clinical utility to the reader without incorporating it into the official product specifications. Another testimonial describes using the tool to infer non obvious patterns and make connections across multiple laboratory results in a clinical setting, synthesizing complex lab data to find non obvious patterns is a diagnostic act. The text presents this as a personal strategy used by a portfolio manager to navigate their health journey. This positioning distances the developer from making a formal claim of diagnostic capability elsewhere. The document also systematically replaces active clinical verbs with more passive administrative terminology. Medical devices are designed to diagnose, assess, treat or manage conditions throughout the launch text. These terms are replaced by understand, navigate, compare and summarize. Under FDA guidelines for clinical decision support, software tools that simply organize, display or summarise historical medical records don't fall under medical device regulations. Consequently, in the formal documentation, the AI is described as a tool to compare a new result with prior tests or to summarize changes since a last appointment. Comparing and summarizing are classified as clerical data processing tasks. The software acts as an automated filing clerk rather than a practicing clinician. This strategy extends to how the tool handles potentially serious health issues, shifting them into the category of general wellness. General wellness software, which assists with things like sleep, fitness and diet, is exempt from strict medical device oversight. When discussing a user with a physical injury, the text doesn't suggest that the AI will generate a clinical, rehabilitation or physical therapy plan. Instead, it describes the AI considering the injury when planning lower impact weekend activities with the family. By routing a clinical issue an injury into a family leisure activity, the software stays within the unregulated boundary of general lifestyle advice. Similarly, when managing dietary restrictions or allergies, the tools positioned as assistant for choosing a restaurant or looking up recipes. This frames the utility as a daily consumer convenience rather than medical nutrition therapy. The software is also explicitly framed as a preparatory tool. To remain classified as a non device, the AI must support rather than replace the independent judgment of a healthcare professional. The document repeatedly notes that the tool helps users prepare questions for a follow up appointment or understand what a doctor might have said. By positioning the AI as a homework assistant that prepares the patient for a human led consultation, the final clinical decision making responsibility remains firmly with the doctor. This is undoubtedly a really useful functionality and something that I'd encourage people to do. But it's not just this functionality that the tools have, which we'll come to a bit later. This positioning is also reflected in how OpenAI measures the performance of its models GPT 5.5 and 5.6. In a traditional clinical trial, a diagnostic tool is evaluated on its clinical sensitivity, specificity and its impact on patient health outcomes. In contrast, the performance metrics for these models are framed around cognitive and communication benchmarks. The models are graded on completeness, communicates clearly, and follows instructions. These benchmarks measure the linguistic performance of a language model rather than clinical efficacy. The primary clinical metric highlighted is the model's ability to recognize when professional care is needed. This positions the AI's clinical capability as a protective safety switch, knowing when to stop talking and refer to a human rather than as an active tool for clinical intervention. Finally, regulatory frameworks require that the responsibility for data accuracy rests with the user. If a clinical system displays incorrect medication doses, it poses a direct safety risk. To mitigate this liability, the setup flow requires the user to review synced medications, remove anything that's no longer relevant, and and manually add missing details. The user is warned that connected information may not be complete or current and must be verified against the original medical source. This requirement turns the consumer into the active editor of the data, legally separating the software developer from liability or inaccurate or outdated clinical records. The careful structuring of language, user feedback and data management represents a very sophisticated design pattern for deploying advanced technology within an extremely heavily regulated area. However, a fundamental discrepancy exists between the sanitized public facing text and the actual capabilities of the live software. While the launch materials are very carefully constructed to present the tool as a passive administrative assistant, the underlying general purpose model remains fully capable of performing highly regulated clinical acts. If a user bypasses the suggested prompts and directly asks the system to diagnose a complex set of symptoms or suggest a specific therapeutic drug regime, the model will usually confidently generate those clinical recommendations. The official marketing material simply omit mentions of these more deeper active diagnostic functions. The omission introduces an important regulatory question Is it legally compliant for a software tool to perform regulated diagnostic acts, provided that the developer doesn't explicitly advertise them? If a software system is capable of producing diagnostic conclusions and therapeutic recommendations, and the developer makes those capabilities accessible to the public, writing generic disclaimers or omitting the feature from a press release shouldn't necessarily shield the product from being classified as a medical device if its actual real world function behaves like one. The strategy of using polished compliant language in public documentation while maintaining less constrained, very capable diagnostic engines under the hood represents a significant regulatory gamble in the model digital health landscape, and it will certainly be really interesting to see how things develop as we go forwards. Thanks for listening, and don't forget to hit like and subscribe if you're interested to hear more of these sorts of practical health AI considerations in future.
Host: Stephen A
Date: July 27, 2026
This episode decodes the careful strategic language behind the US launch of OpenAI's ChatGPT Health product. Stephen A examines how OpenAI navigates regulatory barriers by framing its AI-powered personal health assistant as a non-medical, administrative, and wellness tool—despite its underlying clinical capabilities. The discussion focuses on the deliberate distinction between clinical functionality and how these features are publicly described, raising critical questions about regulatory compliance in the rapidly evolving digital health landscape.
On Regulatory Evasion:
On the Risk of Strategic Understatement:
Stephen A maintains a concise, clinical, and analytical tone throughout, targeting healthcare professionals who want meaningful insights without tech hype. The episode provides a critical, behind-the-scenes look at how leading AI companies strategically maneuver regulatory language. It highlights the significant, unresolved tensions between real-world AI capabilities and current regulatory frameworks—a must-listen for anyone interested in digital health policy, compliance, and the responsible integration of AI into clinical workflows.