Skip to content
DopeSwagYolo

AI in Medicine

Can AI Diagnose Disease? What It Does Well and Where It Fails

AI performs well on narrow, tested tasks such as reading eye and breast images, but 2026 studies found chatbots unreliable for assessing symptoms and urgency.

By DopeSwagYolo4 min read

Researched and fact-checked by AI, with no human review. 13 sources listed below. How we verify

AI can diagnose some diseases, but only in narrow, well-tested situations. Software built to examine one kind of medical image for one condition has performed as well as clinicians, or better, in some studies. General-purpose chatbots are a different matter. When members of the public used them to assess symptoms in a controlled test, results were poor. This article is general information, not medical advice. Anyone worried about a symptom should consult a clinician.

What can AI diagnose reliably today?

In 2018 the US Food and Drug Administration (FDA) authorized IDx-DR. The software checks retinal photographs for diabetic retinopathy, an eye disease linked to diabetes. It was the first autonomous AI diagnostic system the agency authorized. It returns a result without a clinician interpreting the image. Its pivotal trial enrolled 900 people with diabetes at 10 primary care sites. Among the 819 whose results could be analyzed, it correctly flagged 87.2% of those with more than mild disease (sensitivity). It cleared 90.7% of those without it (specificity).

In breast screening, Sweden's MASAI trial randomly assigned 105,934 women to standard screening or to AI-supported screening. In standard screening, two radiologists read each mammogram. In AI-supported screening, software scored each exam, routed it to one or two radiologists and marked suspicious areas.

Results published in The Lancet in January 2026 showed 1.55 interval cancers per 1,000 women in the AI group and 1.76 with standard reading. Interval cancers are those diagnosed between screening rounds. Sensitivity was 80.5% with AI and 73.8% without. Specificity was 98.5% in both groups.

The trial tested whether AI-supported screening was no worse than standard practice, within one country's program. At least one radiologist still read every exam. The authors said the findings do not support replacing health professionals with AI.

Software the FDA authorized in 2018 checks retinal photographs for diabetic retinopathy, an eye disease linked to diabetes.

Can AI diagnose skin cancer?

AI can help identify skin cancer, with important limits. A 2024 meta-analysis in npj Digital Medicine pooled 19 studies comparing algorithms with clinicians. AI had a sensitivity of 87.0% and a specificity of 77.1%. For clinicians overall, the figures were 79.8% and 73.6%. Against expert dermatologists, performance was comparable.

The authors urged caution:

  • Only 5.7% of the 53 studies they reviewed were prospective (designed to collect new data).
  • Of the 53 studies, 73.6% tested algorithms on images held back from the same data used to develop them, not on independent data.
  • Some ethnic groups and skin types were underrepresented.

In January 2024 the FDA authorized DermaSensor, which uses light and an algorithm to assess suspicious lesions. It was authorized as an adjunctive device for physicians, not a stand-alone test. Trade publication MedTech Dive reported on a 22-site study of 1,005 patients. In that study, the device's sensitivity was 96%, compared with 83% for primary care physicians, it reported. MedTech Dive also reported that the device's specificity was 21%, meaning most harmless lesions were flagged too.

Skin cancer diagnosis: AI compared with clinicians in a 2024 meta-analysis
  • AI sensitivity87%
  • Clinician sensitivity79.8%
  • AI specificity77.1%
  • Clinician specificity73.6%

Pooled results from 19 studies. Clinician figures are for clinicians overall; against expert dermatologists, performance was comparable. The authors urged caution. Source: A systematic review and meta-analysis of artificial intelligence versus clinicians for skin cancer diagnosis

Can a chatbot diagnose my symptoms?

This is where the evidence is weakest. In a randomized study published in Nature Medicine in February 2026, 1,298 UK adults worked through medical scenarios. They identified likely conditions and chose what to do next. Some used a chatbot (GPT-4o, Llama 3 or Command R+). A control group used whatever they would normally turn to at home. Given the full scenarios, the models named a relevant condition in 94.9% of cases. People using the same models did so in fewer than 34.5% of cases. The control group had higher odds of doing so.

A separate Nature Medicine paper led by Mount Sinai researchers tested ChatGPT Health, a consumer product OpenAI launched in January 2026. Across 960 responses to 60 clinician-written cases, the tool undertriaged about 52% of emergency responses (33 of 64). That means it directed them to care within 24 to 48 hours instead of the emergency department. Those emergencies covered two conditions, asthma flare-ups and diabetic ketoacidosis. A supplementary analysis of four textbook emergencies, including stroke, found no undertriage. The test used written vignettes, so real-world error rates are unknown. OpenAI told NBC News the study did not reflect how the product is typically used.

ECRI, a patient safety nonprofit, ranked misuse of AI chatbots first among its health technology hazards for 2026.

Why does AI ace tests but struggle in practice?

In a 2024 randomized trial in JAMA Network Open, 50 physicians worked through written cases. Those given GPT-4 alongside their usual references scored a median 76% on diagnostic reasoning. That was not significantly better than the 74% of colleagues without it. Yet the model alone scored 92%. The trial used six cases and a single model.

The World Health Organization's 2024 guidance on generative AI in health warns of automation bias. This means people may trust a system's output so readily that its mistakes go unnoticed.

Diagnostic reasoning scores in a 2024 trial of physicians and GPT-4
  • Physicians with GPT-476%
  • Physicians without GPT-474%
  • GPT-4 alone92%

Physicians with GPT-4 did not score significantly better than those without it. The trial used six cases and a single model. Source: Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial

The bottom line

The strongest evidence is for narrow tools: one condition, one type of image, tested in trials. As of October 2026, the studies described here do not support relying on a chatbot alone to assess symptoms. Things to watch include:

  • MASAI follow-up on later screening rounds and cost-effectiveness
  • prospective skin cancer studies covering more skin types
  • independent tests of consumer health chatbots with real users

None of these tools replaces an examination by a qualified clinician.

Sources

  1. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices, npj Digital Medicine (full text at PubMed Central)
  2. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial, The Lancet
  3. Randomized Trial Shows AI-Supported Mammography Improves Sensitivity and Lowers Interval Cancer Rate, The ASCO Post
  4. AI support in breast cancer screening: Fewer missed cancer cases, Lund University
  5. A systematic review and meta-analysis of artificial intelligence versus clinicians for skin cancer diagnosis, npj Digital Medicine (full text at PubMed Central)
  6. Device Classification Under Section 513(f)(2)(De Novo): DermaSensor (DEN230008), U.S. Food and Drug Administration
  7. Dermasensor wins FDA clearance for AI-enabled skin cancer detection device, MedTech Dive
  8. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study, Nature Medicine (full text at PubMed Central)
  9. ChatGPT Health performance in a structured test of triage recommendations, Nature Medicine
  10. ChatGPT Health 'under-triaged' half of medical emergencies in a new study, NBC News
  11. Misuse of AI chatbots tops annual list of health technology hazards, ECRI
  12. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial, JAMA Network Open
  13. WHO releases AI ethics and governance guidance for large multi-modal models, World Health Organization

More from AI in Medicine

See all in AI in Medicine