Five runs per medical AI case raised the cost about five times
A Nature Medicine study tested an on-premises diagnostic agent on retrospective benchmarks built from medical records. The strongest local model reached 90.0 percent accuracy on a seven-disease task. The cloud-model baseline reached 90.7 percent. The local model reached 83.8 percent on a second four-disease task.
Verified 12:52 PM PDT · 2 original sources
Consistency across five separate runs gave the clearest warning signal. The system used a 0.90 consistency threshold. It kept 49.4 percent of cases for autonomous handling at 98.9 percent accuracy. Humans would review the remaining cases.
The study used retrospective simulations, not patient care. The main benchmarks came from one institution and used text-only cases. Accuracy was lower in older age groups, and some confident errors remained. Running each case five times increased token use and compute by about five times. A separate Nature Medicine comment said clinical trust requires prospective real-world studies.
The next test is a prospective clinical workflow. It could measure patient outcomes, the burden on reviewers and results across subgroups. Researchers also need to see whether the thresholds still work after changes to the model, prompt or hospital. Cost and latency records could show whether five runs per case remain practical when many cases arrive together.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
Continue the edition