Does a new AI study mean doctors are being made obsolete?

By 13 August 2026 Society, Technology

We came across a post on X suggesting that doctors would be made obsolete by AI, declaring them “officially cooked,” slang commonly used by Gen Z for being finished or in serious trouble. The post, which has drawn over 31,000 views within 48 hours, cites a new Google DeepMind and Google Research paper, which allegedly describes an “AI doctor” that has been trained through simulated medical residencies. It points to the reported performance of this system as evidence that the gap between human physicians and AI has effectively vanished.

Among the series of statistics highlighted is a 31% drop in missed red flags, 88% diagnostic accuracy, and a preference for the AI in 87.6% of blinded comparisons by expert clinicians.

With these results, the post describes the AI doctor as the “AlphaGo for medicine”. AlphaGo was the Google DeepMind program that, in 2016, became the first AI to beat a world champion at the board game, Go, a strategy game long considered too complex for computers to master, making it a landmark, widely recognised AI achievement.

That framing makes AI doctors seem plausible. For Singapore particularly, the claim lands as we are navigating a super-aged society, where over 20% of our population are past the age of 65, corresponding to a higher demand for healthcare services. With AI increasingly discussed as one way to meet such demands, a claim such as this could pique some interest here, leading us to look a little closer at it.

 

What does the paper actually show?

The study, ResidencyRL: Reinforcement Learning in Simulated Clinical Environments, was submitted to arXiv, an open research-sharing platform, on 7 August 2026 by a team of 35 researchers, most from Google DeepMind and Google Research. The paper describes a method for training an AI agent to handle simulated clinical conversations with patients, up to 60 back-and-forth exchanges and 8 actions per patient. Some of these simulated patients were designed to be difficult, exaggerating their symptoms, resisting the AI’s questions, or ignoring its advice, mirroring the kind of challenging encounters that can come up in real practice.

In tests, the trained agent reached 88% diagnostic accuracy under adversarial conditions, a 7-percentage point improvement over the untrained base model and missed 31% fewer red flags. In blinded, side-by-side comparisons, board-certified clinicians preferred the trained agent’s responses 87.6% of the time. The agent also outperformed the base model on external Google benchmarks it was not trained on.

These figures match what can be found in the social media post. However, every result reported was measured against simulated patients and other AI benchmarks, not real cases in a clinical setting. Training was also limited to text-based conversations with English-speaking, US-based patients, a narrow slice of real-world healthcare, which spans emergency care, surgery, chronic disease management, and end-of-life care, each with its own demands and context.

Moreover, the paper also states that “prospective validation with real-world workflows remains necessary to establish clinical utility”. In other words, the researchers who built and tested this system are explicit that it has not yet been shown to work outside of simulation, a caveat the post does not mention at all.

The researchers caveat that the study is not designed to show how AI can replace human doctors. Rather, its aim was to demonstrate that a particular training method works well in simulation, not that the system is ready to treat patients. The researchers describe their own goal in similarly modest terms: a model with broad medical knowledge “can be made into a meaningfully more competent and helpful clinician”. They call the current system “a first step,” with real-world use still requiring broader training data, richer clinical grounding, and the ability to track patients across visits over time.

This gap between strong benchmark results and real-world clinical performance is well documented and not unique to this paper. A review in Frontiers in Medicine published in October 2025 found that AI diagnostic tools with benchmark accuracies as high as 94.5% often see real-world performance drop by 15 to 30% once deployed, due to shifts in patient populations and integration challenges that lab testing does not capture.

 

What would need to happen next

The paper’s own limitation points to what should happen next: real-world clinical trials. In the United States, that would likely mean navigating the FDA’s regulatory pathways for AI-enabled medical devices.

In Singapore, a comparable tool would likely fall under the Health Sciences Authority’s (HSA) software-as-medical-device framework, and the AI in Healthcare Guidelines jointly issued by the Ministry of Health and HSA, which set out clear accountability requirements for developers, healthcare institutions and clinicians using such tools before deployment.

Hence, we rate the claim that doctors are “officially cooked” by an “AI doctor” as false. The study documents an AI system tested only in simulation, with its own authors stating that real-world validation is still needed, a caution the post ignores entirely.

The study is a legitimate piece of research that reports its own limitations. The distortion happened one step later, in how the results were framed for social media, using a recognisable name such as AlphaGo to lend credibility to the claim. It is worth being careful of claims that associate a new result with a famous or established name to suggest a level of certainty the result has not actually reached.

Leave a Reply