Cortico Launches Clinical AI Safety Benchmark With an Accuracy–Caution Trade-Off
The free benchmark examines escalation, dangerous reassurance and uncertainty. Its initial results show why a strong diagnostic score alone may not reveal how a model handles clinical risk.
Listen to this story
The audio brief
Story brief
3 key pointsCortico’s MedSafe-Dx gives developers and health systems an open way to test clinical AI judgment beyond diagnostic accuracy. In an initial evaluation of 11 models across 250 simulated adult cases, GPT-5.2 achieved the strongest safety score but escalated 71% of routine cases, while Gemini 3 Pro Preview led on diagnostic recall and ranked lowest on safety. The framework is auditable and reusable, but its results are...
- 01
The benchmark tests escalation, unsafe reassurance, and uncertainty handling—not just whether a diagnosis appears plausible.
- 02
GPT-5.2 had the top Safety Pass Rate but escalated 71% of routine cases.
- 03
Gemini 3 Pro Preview posted the highest Top-3 diagnostic recall and the lowest Safety Pass Rate.
Cortico has launched MedSafe-Dx, a free open benchmark designed to test whether clinical AI systems can make safer decisions, not merely recognize the right diagnosis. Its first evaluation found a sharp trade-off: the model with the strongest safety score also escalated 71% of routine cases, while the model with the best diagnostic recall had the lowest safety pass rate.
That distinction matters in clinical decision support, where a plausible diagnosis alone does not show whether a system will flag a potentially life-threatening condition, falsely reassure someone at risk, or communicate doubt when symptoms are unclear. MedSafe-Dx is aimed at those judgment calls.
A test of decisions, not exam answers
The benchmark’s launch paper evaluated 11 frontier models using 250 simulated adult cases from the DDxPlus dataset. The results compare how models behave in this test environment; they are not clinical deployment outcomes.
The three behaviors MedSafe-Dx measures
- Escalating cases that may involve life-threatening conditions.
- Avoiding reassurance that could be unsafe for a patient at risk.
- Showing appropriate uncertainty when the available symptoms are ambiguous.
Cortico designed the test to be auditable rather than having one AI system grade another. It uses deterministic rules, a frozen dataset and standardized outputs. Cortico has also released the benchmark’s code and dataset alongside a preprint, allowing others to inspect or reuse the framework.
Caution can create a workload problem
The results make the benchmark’s central tension concrete. Broad escalation may reduce the chance of missing a dangerous case, but excessive warnings can add work for clinical teams and risk becoming easier to ignore. Every model in the initial evaluation missed at least some cases categorized as requiring escalation.
The benchmark does not settle which model should be used in care. Its cases are simulated, and its escalation labels are proxy labels derived from dataset severity ratings rather than assessments by practising clinicians. Cortico describes MedSafe-Dx as a comparative safety test, not clinical validation or proof that a model is ready for deployment.
Still, the release gives health systems and model developers a clearer way to examine a question that exam-style benchmarks can miss: not only whether an AI reaches a likely diagnosis, but whether its recommendation is safe when the stakes, evidence and uncertainty differ from case to case.
Sources
- techcouver.comCortico Launches Benchmark to Test Clinical AI Safety - Techcouver.com
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.