Open models found psychosis-risk signals—and over-pathologized ordinary experience.
Researchers evaluated 11 locally deployed open-weight models on 678 partial clinical interview transcripts from 373 participants. The best model reached 80% classification accuracy and 93% sensitivity, but only 58% specificity. The peer-reviewed study reported clinically relevant confabulation in about 3% of reviewed summaries and a recurring tendency to over-score non-clinical experiences.
Verified 12:07 AM PDT · 2 original sources
The evidence
What the reporting establishes
What happened
Researchers evaluated 11 locally deployed open-weight models on 678 partial clinical interview transcripts from 373 participants. The best model reached 80% classification accuracy and 93% sensitivity, but only 58% specificity. The peer-reviewed study reported clinically relevant confabulation in about 3% of reviewed summaries and a recurring tendency to over-score non-clinical experiences.
Pressure point
This was a retrospective transcript study, not a prospective diagnostic trial. Most participants were already in a high-risk research cohort, site-level performance varied, false positives could burden services or harm patients, and the authors explicitly do not recommend immediate clinical deployment.
What to watch
Prospective trials in real referral populations, calibration across sites and languages, patient-consent and privacy controls, false-positive burden, clinician override behavior, and whether smaller models preserve performance outside the research dataset.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
- npj Digital MedicineEvaluating large language models for assessment of psychosis risk ↗
- King's College LondonAI could help identify patients at risk of psychosis more quickly ↗
Continue the morning