Modelspublished

OpenEvidence Releases Three Clinician AI Models, Keeps Darwin Behind Research Access

The rollout gives verified clinicians a choice between answers measured in seconds and literature investigations measured in minutes. Its strongest performance claims, however, come from company-run evaluations, while Darwin remains gated over stated dual-use concerns.

By 3 min read
OpenEvidence Releases Three Clinician AI Models, Keeps Darwin Behind Research Access
OpenEvidence Releases Three Clinician AI Models, Keeps Darwin Behind Research Access

Listen to this story

The audio brief

About 1:40
0:001:40
Read transcript
OpenEvidence is giving verified clinicians three medical AI models for free, with no usage limit—but keeping its most capable system, Darwin, behind a research-only gate. The production choices are built around depth and speed. Osler is the default, returning an answer in about five seconds. Sackett takes roughly 30 seconds for questions that need a deeper evidence search. Snow, which replaces the company’s Deep Consult feature, can spend around five minutes investigating the medical literature before producing a report. All three are available on OpenEvidence’s website and mobile apps, and the company says they share the same clinical-accuracy standard; the difference is how long they reason and how deeply they search. Darwin is different. Approved researchers can apply for access, while OpenEvidence works with institutional and academic partners to evaluate its safeguards. The company cites dual-use concerns, especially around virology, immunology, human genetics, bioweapons-relevant research, and germline editing. OpenEvidence says Darwin got all 660 questions right on its final, physician-reviewed MedQA set, but that was a reduced and re-annotated evaluation—not the original 1,273-question test split. Its other scores are also company-supplied. And a Nature Medicine study found general-purpose models outperforming earlier specialized clinical tools, calling for independent, real-world validation. The key question is whether Darwin’s restricted evaluation supports broader access—and whether these faster clinician models hold up outside OpenEvidence’s own benchmarks.

Story brief

3 key points

OpenEvidence is expanding its clinician product while withholding Darwin, its most capable model, from general clinical access. Osler, Sackett, and Snow are now free and unlimited for verified clinicians, with response times ranging from roughly five seconds to five minutes. Darwin is limited to approved researchers because of stated dual-use concerns in areas including virology and human genetics. OpenEvidence...

  1. 01

    Osler, Sackett, and Snow take about five seconds, 30 seconds, and five minutes, respectively, reflecting increasing search depth.

  2. 02

    Darwin’s research-only access will remain until institutional and academic partners help assess its safeguards and capabilities.

  3. 03

    OpenEvidence reports Darwin answered 660 physician-reviewed MedQA questions correctly, but that was a reduced, re-annotated evaluation set.

Verified clinicians can now use three new OpenEvidence medical AI models without a usage limit or fee, choosing between a default answer engine, a deeper evidence search and a five-minute literature investigation. The company has kept Darwin, the model it calls its most advanced, in a research preview instead of putting it into the same clinical rollout.

The split creates a practical boundary in the product line. Osler is the default and fastest option, with answers taking about five seconds; Sackett is designed for questions where weighing evidence matters and takes about 30 seconds; Snow is the deepest production option and takes roughly five minutes to produce a report after investigating medical literature.

OpenEvidence is rolling the three production models out on its website and iOS and Android apps. Clinicians can select among them in the interface. The company says the models are held to the same clinical-accuracy standard; their intended distinction is how long they reason and how deeply they search.

Darwin is available only to approved researchers who apply through OpenEvidence's website. The company says the restriction reflects dual-use risks: advanced reasoning in virology, immunology and human genetics could accelerate tightly governed work, including bioweapons-relevant research and germline editing outside mainstream scientific oversight.

OpenEvidence says it will evaluate Darwin's safeguards and capabilities with institutional partners, research collaborators and accredited academic AI researchers before widening access. It also says Darwin's capabilities will move into the three clinician models as those safeguards are validated.

What OpenEvidence says Darwin scored

  • It answered all 660 questions correctly in OpenEvidence's final physician-reviewed MedQA evaluation set.
  • It scored 72.8% on MedXpertQA, 82.7% on HealthBench Professional and 87.2% on NOHARM, according to the company.

The reported Darwin results are company-supplied evaluations, including a MedQA set narrowed to 660 questions after physician re-annotation and review. That makes the perfect score a result on OpenEvidence's final evaluation set, rather than a result on the original 1,273-question test split.

Independent context points to the remaining test. A June Nature Medicine study evaluated an earlier OpenEvidence tool and UpToDate Expert AI against three frontier general-purpose models using MedQA, HealthBench and a real-clinical-queries benchmark reviewed by 12 U.S. clinicians. It found the general-purpose models outperformed the specialized tools across all three evaluations and called for independent, real-world assessment before clinical use. That study does not evaluate Darwin.

OpenEvidence says it plans specialty-focused models beginning with oncology, radiology and clinical genetics. For now, the more immediate question is whether its restricted Darwin evaluation produces evidence that supports broader access—and whether the clinician-facing models deliver dependable results beyond the company's benchmark framework.

Sources

  1. unite.aiOpenEvidence Launches Medical AI Model Family With Darwin Preview
  2. nature.comGeneral-purpose large language models outperform specialized clinical AI tools on medical benchmarks - Nature Medicine