Cortico Launches Clinical AI Safety Benchmark With an Accuracy–Caution Trade-Off

The free benchmark examines escalation, dangerous reassurance and uncertainty. Its initial results show why a strong diagnostic score alone may not reveal how a model handles clinical risk.

By 2 min read
Cortico Launches Clinical AI Safety Benchmark With an Accuracy–Caution Trade-Off
Cortico Launches Clinical AI Safety Benchmark With an Accuracy–Caution Trade-Off

Listen to this story

The audio brief

About 1:41
0:001:41
Read transcript
Cortico’s new clinical AI benchmark exposes a sharp trade-off: GPT-5.2 earned the strongest safety score, but escalated 71 percent of cases classified as routine. Meanwhile, Gemini 3 Pro Preview achieved the best Top-3 diagnostic recall—and the lowest safety pass rate in the evaluation. The free benchmark, called MedSafe-Dx, is designed to test more than whether an AI can name a plausible diagnosis. It checks whether the system escalates cases that could be life-threatening, avoids dangerous reassurance, and communicates uncertainty when symptoms are ambiguous. In its initial evaluation, Cortico tested 11 frontier models on 250 simulated adult cases from the DDxPlus dataset. The results show why diagnostic accuracy alone can be misleading. A system that escalates broadly may miss fewer dangerous cases, but it can also create extra work for clinical teams—and warnings may become easier to ignore. At the other extreme, confident reassurance can be risky when the evidence is incomplete. Every model tested missed at least some cases labeled as requiring escalation. MedSafe-Dx is built for inspection and reuse, with deterministic rules, a frozen dataset, standardized outputs, and released code and data. But this is a comparison tool, not clinical validation: the cases are simulated, and escalation labels are proxies based on dataset severity ratings rather than practising clinicians’ assessments. The key constraint is that the benchmark still cannot say which model is safe to deploy. It does offer a clearer test of whether diagnostic confidence matches clinical risk.

Story brief

3 key points

Cortico’s MedSafe-Dx gives developers and health systems an open way to test clinical AI judgment beyond diagnostic accuracy. In an initial evaluation of 11 models across 250 simulated adult cases, GPT-5.2 achieved the strongest safety score but escalated 71% of routine cases, while Gemini 3 Pro Preview led on diagnostic recall and ranked lowest on safety. The framework is auditable and reusable, but its results are...

  1. 01

    The benchmark tests escalation, unsafe reassurance, and uncertainty handling—not just whether a diagnosis appears plausible.

  2. 02

    GPT-5.2 had the top Safety Pass Rate but escalated 71% of routine cases.

  3. 03

    Gemini 3 Pro Preview posted the highest Top-3 diagnostic recall and the lowest Safety Pass Rate.

Cortico has launched MedSafe-Dx, a free open benchmark designed to test whether clinical AI systems can make safer decisions, not merely recognize the right diagnosis. Its first evaluation found a sharp trade-off: the model with the strongest safety score also escalated 71% of routine cases, while the model with the best diagnostic recall had the lowest safety pass rate.

That distinction matters in clinical decision support, where a plausible diagnosis alone does not show whether a system will flag a potentially life-threatening condition, falsely reassure someone at risk, or communicate doubt when symptoms are unclear. MedSafe-Dx is aimed at those judgment calls.

A test of decisions, not exam answers

The benchmark’s launch paper evaluated 11 frontier models using 250 simulated adult cases from the DDxPlus dataset. The results compare how models behave in this test environment; they are not clinical deployment outcomes.

The three behaviors MedSafe-Dx measures

  • Escalating cases that may involve life-threatening conditions.
  • Avoiding reassurance that could be unsafe for a patient at risk.
  • Showing appropriate uncertainty when the available symptoms are ambiguous.

Cortico designed the test to be auditable rather than having one AI system grade another. It uses deterministic rules, a frozen dataset and standardized outputs. Cortico has also released the benchmark’s code and dataset alongside a preprint, allowing others to inspect or reuse the framework.

Caution can create a workload problem

The results make the benchmark’s central tension concrete. Broad escalation may reduce the chance of missing a dangerous case, but excessive warnings can add work for clinical teams and risk becoming easier to ignore. Every model in the initial evaluation missed at least some cases categorized as requiring escalation.

The benchmark does not settle which model should be used in care. Its cases are simulated, and its escalation labels are proxy labels derived from dataset severity ratings rather than assessments by practising clinicians. Cortico describes MedSafe-Dx as a comparative safety test, not clinical validation or proof that a model is ready for deployment.

Still, the release gives health systems and model developers a clearer way to examine a question that exam-style benchmarks can miss: not only whether an AI reaches a likely diagnosis, but whether its recommendation is safe when the stakes, evidence and uncertainty differ from case to case.

Sources

  1. techcouver.comCortico Launches Benchmark to Test Clinical AI Safety - Techcouver.com

Loading discussion...