Grok 4.6 Clears 50% on Biosecurity Refusal and Routine Biology Tasks
The reported result rewards a difficult balance: recognizing hazards hidden inside ordinary-looking research requests without shutting down legitimate biological work. Its surveillance score shows that refusal performance is not the whole biosecurity test.
Listen to this story
The audio brief
Story brief
3 key pointsIn results published September 1, xAI reported that Grok 4.6 led the tested models on LatchBio’s balance of biosecurity refusal and routine biology work, averaging a 62.1% harmonic-mean score across agent harnesses. It refused 59.2% of disguised red-team tasks and completed 64.8% of benign ones, but ranked behind Opus 5 on pathogen-surveillance workflows at 53.5%. The findings highlight a deployment constraint:...
- 01
Grok 4.6’s 62.1% harmonic-mean result combines refusal of concealed hazards with completion of routine biology tasks.
- 02
The refusal benchmark included 46 red-team tasks using obfuscated files, scientific data, or misleading context.
- 03
On BioSecBench-Surveillance, Grok 4.6 averaged 53.5%, behind Opus 5 and ahead of GPT-5.6 Sol.
xAI says LatchBio’s new evaluation found Grok 4.6 was the only tested frontier model to exceed 50% on both refusing disguised hazardous biology tasks and completing routine ones. The result puts the model’s ability to distinguish intent at the center of a biosecurity contest where excessive blocking can also impair useful work.
The result, published by xAI on September 1, concerns two LatchBio benchmark suites. BioSecBench-Refusal tests whether an agent can separate ordinary biological research from requests that conceal biosecurity hazards. BioSecBench-Surveillance tests pathogen genomic-surveillance workflows, which combine file inspection, tool use, and scientific judgment on sequencing data.
A test built around concealed intent
The refusal suite pairs routine tasks adapted from published literature with 46 red-team tasks that are designed to appear ordinary. The hazardous element can be hidden in scientific data, mislabeled files, or other obfuscation. A model that merely reacts to terms such as pathogen or toxin could block routine work while missing the disguised cases; the benchmark is intended to test whether it examines the task context instead.
xAI said LatchBio used multiple agent harnesses to reduce setup-related confounds and, unless noted otherwise, tested agents at their highest offered effort settings. That makes the published figures a measure of model-and-agent performance under those configurations, rather than a simple chat-response test.
Grok 4.6 refused 59.2% of red-team tasks in the reported BioSecBench-Refusal results.
Grok 4.6 completed 64.8% of routine biological tasks in the same results.
Across harnesses, Grok 4.6 occupied the top three positions and averaged 62.1% on LatchBio’s trial-weighted harmonic mean of refusal and routine compliance, according to xAI.
Grok’s lead does not extend to every biological workflow
On BioSecBench-Surveillance, Grok 4.6 averaged a 53.5% success rate. xAI said that placed it behind Opus 5 and ahead of GPT-5.6 Sol. The distinction matters because surveillance evaluates an agent’s ability to carry out public-health monitoring workflows, not its willingness to decline suspicious requests.
xAI described evaluation traces in which Grok 4.6 inspected task environments, found mismatches between a prompt’s stated purpose and its contents, and refused after identifying high-risk material obscured through filenames or encryption. It said the model used similar environment reasoning on plainly benign tasks before judging them safe.
The next question is whether safeguards keep pace with capability
xAI’s stated safeguard stack
- Refusal training intended to teach the model when and how to refuse, including in adversarial scenarios.
- Inference-time safeguards that xAI says reject harmful requests before they reach the model, plus behavioral controls in deployment.
- Post-deployment monitoring at the session and user level, intended to detect and stop patterns of adversarial use and feed back into calibration.
xAI said Grok 4.6 made substantial gains in refusal and biosecurity performance over Grok 4.5 and Grok 4.3. It plans broader pre-deployment testing, more third-party evaluations, stronger post-deployment monitoring, and deployments with organizations working at the frontier of biology as model capability and autonomy rise.
The unresolved calibration problem is two-sided. xAI argues that helping malicious actors is dangerous, but that incorrectly refusing routine biological work can also weaken outbreak detection and other critical research. The benchmarks offer one reported snapshot of that trade-off; the harder next test is whether the same balance holds as agents receive wider access to biological tools and workflows.
Sources
- x.aiBiosecurity at the frontier