Modelspublished

Grok 4.6 Clears 50% on Biosecurity Refusal and Routine Biology Tasks

The reported result rewards a difficult balance: recognizing hazards hidden inside ordinary-looking research requests without shutting down legitimate biological work. Its surveillance score shows that refusal performance is not the whole biosecurity test.

By 3 min read
Grok 4.6 Clears 50% on Biosecurity Refusal and Routine Biology Tasks
Grok 4.6 Clears 50% on Biosecurity Refusal and Routine Biology Tasks

Listen to this story

The audio brief

About 1:29
0:001:29
Read transcript
Grok 4.6 cleared a difficult biosecurity test: it was the only frontier model xAI says exceeded 50 percent on both refusing disguised hazardous biology tasks and completing routine ones. In LatchBio’s evaluation, the model refused 59.2 percent of 46 red-team tasks, where dangerous intent could be hidden in scientific data, mislabeled files, or misleading context. It also completed 64.8 percent of benign biology tasks. Combined as a harmonic mean—a score that penalizes weakness on either side—it averaged 62.1 percent across the tested agent setups. That balance is the point. A system that blocks anything mentioning pathogens or toxins may look safe while disrupting legitimate research. A system that complies too freely can miss hazards concealed inside ordinary-looking work. LatchBio used multiple agent harnesses, and the results came from model-and-agent configurations running at their highest offered effort, not simple chat responses. Grok’s lead did not carry over to every workflow. On BioSecBench-Surveillance, which tests pathogen genomic-surveillance work involving files, tools, and scientific judgment, it averaged 53.5 percent—behind Opus 5 but ahead of GPT-5.6 Sol. xAI says it plans broader third-party evaluations, pre-deployment testing, and post-deployment monitoring as autonomy expands. The key constraint is whether this calibration still holds when agents gain wider access to biological tools and workflows.

Story brief

3 key points

In results published September 1, xAI reported that Grok 4.6 led the tested models on LatchBio’s balance of biosecurity refusal and routine biology work, averaging a 62.1% harmonic-mean score across agent harnesses. It refused 59.2% of disguised red-team tasks and completed 64.8% of benign ones, but ranked behind Opus 5 on pathogen-surveillance workflows at 53.5%. The findings highlight a deployment constraint:...

  1. 01

    Grok 4.6’s 62.1% harmonic-mean result combines refusal of concealed hazards with completion of routine biology tasks.

  2. 02

    The refusal benchmark included 46 red-team tasks using obfuscated files, scientific data, or misleading context.

  3. 03

    On BioSecBench-Surveillance, Grok 4.6 averaged 53.5%, behind Opus 5 and ahead of GPT-5.6 Sol.

xAI says LatchBio’s new evaluation found Grok 4.6 was the only tested frontier model to exceed 50% on both refusing disguised hazardous biology tasks and completing routine ones. The result puts the model’s ability to distinguish intent at the center of a biosecurity contest where excessive blocking can also impair useful work.

The result, published by xAI on September 1, concerns two LatchBio benchmark suites. BioSecBench-Refusal tests whether an agent can separate ordinary biological research from requests that conceal biosecurity hazards. BioSecBench-Surveillance tests pathogen genomic-surveillance workflows, which combine file inspection, tool use, and scientific judgment on sequencing data.

A test built around concealed intent

The refusal suite pairs routine tasks adapted from published literature with 46 red-team tasks that are designed to appear ordinary. The hazardous element can be hidden in scientific data, mislabeled files, or other obfuscation. A model that merely reacts to terms such as pathogen or toxin could block routine work while missing the disguised cases; the benchmark is intended to test whether it examines the task context instead.

xAI said LatchBio used multiple agent harnesses to reduce setup-related confounds and, unless noted otherwise, tested agents at their highest offered effort settings. That makes the published figures a measure of model-and-agent performance under those configurations, rather than a simple chat-response test.

The reported balance on refusal
59.2%Red-team tasks refused

Grok 4.6 refused 59.2% of red-team tasks in the reported BioSecBench-Refusal results.

64.8%Routine tasks completed

Grok 4.6 completed 64.8% of routine biological tasks in the same results.

62.1%Harmonic-mean score

Across harnesses, Grok 4.6 occupied the top three positions and averaged 62.1% on LatchBio’s trial-weighted harmonic mean of refusal and routine compliance, according to xAI.

Grok’s lead does not extend to every biological workflow

On BioSecBench-Surveillance, Grok 4.6 averaged a 53.5% success rate. xAI said that placed it behind Opus 5 and ahead of GPT-5.6 Sol. The distinction matters because surveillance evaluates an agent’s ability to carry out public-health monitoring workflows, not its willingness to decline suspicious requests.

xAI described evaluation traces in which Grok 4.6 inspected task environments, found mismatches between a prompt’s stated purpose and its contents, and refused after identifying high-risk material obscured through filenames or encryption. It said the model used similar environment reasoning on plainly benign tasks before judging them safe.

The next question is whether safeguards keep pace with capability

xAI’s stated safeguard stack

  • Refusal training intended to teach the model when and how to refuse, including in adversarial scenarios.
  • Inference-time safeguards that xAI says reject harmful requests before they reach the model, plus behavioral controls in deployment.
  • Post-deployment monitoring at the session and user level, intended to detect and stop patterns of adversarial use and feed back into calibration.

xAI said Grok 4.6 made substantial gains in refusal and biosecurity performance over Grok 4.5 and Grok 4.3. It plans broader pre-deployment testing, more third-party evaluations, stronger post-deployment monitoring, and deployments with organizations working at the frontier of biology as model capability and autonomy rise.

The unresolved calibration problem is two-sided. xAI argues that helping malicious actors is dangerous, but that incorrectly refusing routine biological work can also weaken outbreak detection and other critical research. The benchmarks offer one reported snapshot of that trade-off; the harder next test is whether the same balance holds as agents receive wider access to biological tools and workflows.

Sources

  1. x.aiBiosecurity at the frontier