Scott Winters thought he had the flu. Instead, he nearly died because he trusted a chatbot to tell him otherwise.
The former Florida pastor is suing OpenAI and Sam Altman in San Francisco County Superior Court. His claim? ChatGPT’s medical advice contributed to the pulmonary embolism that nearly killed him.
Here is what happened in 2025. Winters complained of dizziness. His blood pressure was unstable. He turned to ChatGPT-4o.
The AI didn’t panic. It dismissed the symptoms as minor. It told Winters to stay recliner-bound. The bot calculated he needed eight to ten more “episodes” before his condition warranted real concern.
Weeks later, a blood clot lodged in his lungs. One of Winters’ own doctors linked the clot to the prolonged immobility that the chatbot had specifically recommended.
On the day the pain peaked, Winters asked if groin tenderness required an ER visit. The chatbot didn’t call an ambulance. It invoked his faith. It reportedly said: “God did not design your body to endlessy fail.”
Winters nearly died. OpenAI says its tools are not medical devices and users shouldn’t rely on them for diagnosis. Winters’ lawyers disagree. They want financial damages. They also want to pause ChatGPT Health. They want an independent safety evaluation.
This isn’t the first time an AI has been blamed for a health crisis.
In May, a Texas couple sued Open AI after their son overdosed. The argument: their son had searched for drug information on ChatGPT and ignored safety guardrails because the bot provided usable answers.
These lawsuits test a heavy question. How much responsibility do AI companies bear when users seek medical or psychological help in a crisis?
What the Research Actually Says About AI Diagnostics
The Winters case feels like a nightmare. But controlled research suggests the story is more nuanced. In rigid, scripted environments, AI is terrifyingly good at diagnosing.
A 2024 JAMA Internal Medicine study compared GPT-4 to human doctors. The test included 20 clinical cases rated by the r-IDEA clinical-reasoning scale.
GPT-4 scored a perfect 10 out of 10. Attendings averaged 9. Residents averaged 8.
The chatbot was actually less often completely wrong than the human residents.
A follow-up study in JAMA Network Open pushed the limits. Researchers tested 50 physicians on five extremely difficult cases. ChatGPT, operating alone, reached 90% accuracy. Physicians working alone scored 74%. Physicians using ChatGPT as a tool scored 76%.
The AI assistant didn’t help much. Why? Many doctors disregarded or second-guessed what the bot suggested.
Larger studies reinforce this gap. Nature published work in 2025 on Google’s AMIE model. It tested 302 complex real-world cases against 20 clinicians. AMIE found the correct diagnosis 59% of the time. Unassisted clinicians hit 34%. Doctors using AMIE outperformed those using standard search engines.
A meta-analysis of 50 studies across npj Digital Medicine concluded that AI generally performs comparably to clinicians, and sometimes better, on standardized tasks.
But there is a massive caveat. These studies use written vignettes. Structured prompts. Hand-selected cases. They are not real life.
Where AI Diagnosis Fails Real-World Users
Real-world advice looks different. It involves weeks of unscripted conversation. It involves incomplete information. It involves no physical exam.
Research shows AI struggles when the script disappears.
A NEJM AI study built a 750-question benchmark. It used script concordance testing, which measures how new information should shift a diagnosis under uncertainty. It pitted ten leading AI models against 1,500 medical professionals.
OpenAI’s o3, the top performer, managed only 68% accuracy. That is below the level of senior residents. This is ironic. These same models ace multiple-choice medical licensing exams.
High test scores don’t equal sound clinical judgment in the messy reality of a sick body.
The problem gets worse with hallucinations.
A study in Communications Medicine fed six chatbots—including GPT-4o and DeepSeek—clinical vignettes seeded with lies. Invented lab tests. Made-up diseases.
Under default conditions, the models accepted these lies 50% to 83% of the time. They confidently described invented diseases as facts. A simple prompt warning the AI that input might be false helped. But it didn’t stop the errors.
Then there is the omission risk.
A Stanford-led team scored 20 models on 1,100 cases. They weren’t just looking for diagnostic accuracy. They looked for potential harm.
Direct application of the advice risked severe harm or death in 24.6% of the cases. Over 80% of these errors were omissions. The AI failed to flag something dangerous. It stayed silent.
Winters’ experience mirrors this silence. Escalating reassurance instead of urgent referral. A chatbot telling you to sit still while your legs clot.
The Deployment Problem, Not Just the Model Problem
A 2025 Sermo survey of over 1,000 doctors confirms the fear. 94% of physicians are concerned about patients using AI for medical advice. Misdiagnosis and delayed care are the top worries.
The research creates a split picture.
In narrow, well-defined tasks? AI matches or beats doctors.
In open-ended, high-stakes human interactions? It fails. It hallucinates. It omits critical warnings.
For health systems and tech companies, this isn’t just a question of competence. It is a design failure. The same model that aced the test case can confidently narrate a fabricated death sentence or talk a frightened user into waiting too long.
AI’s diagnostic potential is real. In some structured metrics, it exceeds the average physician.
But whether that promise survives contact with human behavior is unknown. People don’t interact with AI like standardized test-takers. They type symptoms at 3 a.m. in pain. They want reassurance. They are vulnerable.
The courts, hospitals, and researchers still have a lot to figure out. The next wave of cases will decide if that gap closes, or if it widens until it becomes a chasm.



















