Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper Replication

The company’s result centers on reproducing known findings, not making discoveries. Its next test is whether that training approach can move from replication to reliable new science.

By 2 min read
Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper Replication
Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper Replication

Listen to this story

The audio brief

About 1:29
0:001:29
Read transcript
Inherent says its Faraday research agent has beaten Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at reproducing findings from published scientific papers. The London startup says Faraday received no answers in advance and independently recovered the results. But this is a replication test, not a discovery test: the system was asked to reproduce known work, not generate new scientific knowledge. Faraday’s core is a 27-billion-parameter Qwen 3.6 model, trained with reinforcement learning to improve its choices about which experiments to run and how to design them. Inherent calls that capability “research taste”—the judgment to spend effort on experiments that are likely to be informative. There’s an important wrinkle in the comparison. Faraday delegates coding to OpenAI’s GPT-5.5 Codex, rather than using a coding tool built by Inherent. So the reported result measures the design of a complete research system—its model, training process, and software choices—not Qwen 3.6 by itself. And the evidence is difficult to independently assess: Inherent has not released numerical scores, the test papers, or the evaluation method. The company has raised 50 million dollars in seed funding and has about a dozen employees, with plans to reach 20 to 25 by year-end. The key question is whether Faraday’s experimental judgment can carry from reproducing published results to producing reliable new science.

Story brief

3 key points

Inherent’s Faraday is a research-agent system built around a 27B Qwen 3.6 model and reinforcement learning for experimental choices. The London startup claims it outperformed Claude Opus 4.8 and GPT-5.5 on unpublished details of a replication evaluation, but provided no scores, papers, or methodology. Because Faraday delegates coding to OpenAI’s GPT-5.5 Codex, the result measures system design rather than a...

  1. 01

    Faraday reproduces existing paper findings; it has not yet demonstrated new scientific discovery.

  2. 02

    Inherent disclosed no scores, test papers, or evaluation methodology, limiting independent assessment of the claimed win.

  3. 03

    The system combines Qwen 3.6, reinforcement learning, and GPT-5.5 Codex for coding.

Inherent says its Faraday agent independently reproduced findings from published scientific papers without receiving the answers in advance, outperforming Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 in the company’s evaluation. That comparison puts a young London lab’s bet on research-agent design against much larger frontier systems.

A rehearsal for the larger goal

Faraday was asked to recover findings already reported in papers, rather than generate new scientific knowledge. That distinction is central: Inherent’s stated long-term goal is an agent that contributes to discovery across scientific fields. Cofounder and chief scientist Edward Hughes described replication as a common starting exercise for researchers, saying many PhD students begin there.

The company’s approach is built around what it calls research taste: deciding which experiments are worth running and how to design them. Inherent used reinforcement learning, which rewards desired outcomes, to train Faraday toward that experimental judgment. It is betting this method will generalize across fields more effectively than training an agent primarily on descriptions of scientific practice.

A smaller core model, with outside tools

Inherent chose not to build its own coding tool, instead having Faraday use GPT-5.5 Codex. That design choice makes the reported win less a clean base-model contest than a test of how a research agent combines a core model, reward-based training and specialized software.

The disclosed Faraday setup

  • Qwen 3.6 as the 27-billion-parameter core model described by Inherent.
  • Reinforcement learning aimed at experimental selection and design.
  • GPT-5.5 Codex for coding instead of an Inherent-built coding tool.

The evidence still stops at replication

The disclosed evaluation includes no numerical scores, named test papers or methodology, limiting what can be concluded from the comparison. It supports Inherent’s early case for its research-agent approach, but does not yet demonstrate that Faraday can produce reliable new science.

Inherent emerged from stealth with a $50 million seed round and has about a dozen employees. It plans to grow to roughly 20 to 25 people by year-end; its most consequential next proof point will be whether experimental judgment carries beyond reproducing published results.

Sources

  1. techcrunch.comInherent, founded by DeepMind alumni, says its AI 'teammate' just outperformed Anthropic and OpenAI at replicating research | TechCrunch

Loading discussion...