Toolspublished

Hack The Box Benchmark: Best Human Team Cleared 36 Challenges; Top AI Team Reached 32

The latest benchmark points to a useful division of labor in cyber work: agents can accelerate solving, but elite performance still depends on people choosing paths and checking results.

By 3 min read
Hack The Box Benchmark: Best Human Team Cleared 36 Challenges; Top AI Team Reached 32
Hack The Box Benchmark: Best Human Team Cleared 36 Challenges; Top AI Team Reached 32

Listen to this story

The audio brief

About 1:40
0:001:40
Read transcript
The best human team in Hack The Box’s latest cyber benchmark cleared all 36 challenges. The best AI-augmented team reached 32, even though the strongest AI-assisted teams worked roughly three to four times faster than human-only peers. That gap captures the benchmark’s main lesson: agents can accelerate the work, but they have not yet replaced expert judgment. Across three years of Hack The Box’s Global Cyber Skills Benchmark, agents appeared on 17 of the top 25 teams. That sounds significant, but agents represented only 2.7 percent of registered accounts, and the data shows association, not causation. Ninety-three designated agent accounts across 54 teams produced 4.2 percent of submitted flags and 4.6 percent of awarded points; 46 of those accounts were active. A November 2025 NeuroGrid CTF comparison offers a cleaner test because augmented and human-only teams faced the same challenges. Across the field, AI-augmented teams solved them 3.2 times faster. Among the top five percent, the advantage narrowed to 1.69 times. Meanwhile, median solve time across the competition fell from 26.1 hours in 2024 to 13.8 in 2026, but changing challenge boards make the cause unclear. The practical takeaway is narrower than replacement: agents can compress research, coding, troubleshooting, and first-pass analysis. Operators still choose routes, test weak outputs, and validate results. The constraint to watch is whether teams can direct and check an agent’s work at the speed it creates.

Story brief

3 key points

Hack The Box’s three-year cyber-skills data shows AI agents are concentrated among elite teams, but not yet superior to the best humans. In a November 2025 comparison, AI-augmented teams solved challenges 3.2 times faster overall, while the top human team still finished the full 36-challenge board versus 32 for the best AI-augmented team. The practical takeaway is deployment discipline: agents may accelerate...

  1. 01

    Agents appeared on 17 of the top 25 teams despite representing just 2.7% of registered accounts; this shows association, not causation.

  2. 02

    Ninety-three designated agent accounts across 54 teams generated 4.2% of flags and 4.6% of awarded points; 46 accounts were active.

  3. 03

    In the NeuroGrid CTF comparison, AI-augmented teams solved challenges 3.2× faster overall, but the top-5% advantage narrowed to 1.69×.

The strongest AI-augmented teams in Hack The Box’s benchmark worked three to four times faster than their human-only peers. Yet the only team to clear all 36 challenges was human; the best AI-augmented team stopped at 32.

Where agents show up

Hack The Box’s latest three-year data places AI agents disproportionately near the top of its Global Cyber Skills Benchmark. Agents appeared on 17 of the top 25 teams, or 68%, while representing 2.7% of registered accounts overall. Thirty-three of the top 100 teams used an agent.

The footprint was small but productive. The event counted 93 designated agent accounts across 54 teams, with 46 active. Those accounts produced 4.2% of submitted flags and 4.6% of awarded points—more than their share of registered accounts.

That pattern is association, not proof that agents created the performance gap. Hack The Box explicitly says its observational benchmark cannot establish that AI caused teams to perform better; it shows that many strong competitors have incorporated it into their toolkit.

The test that separates association

A November 2025 NeuroGrid CTF comparison offers a more direct view because AI-augmented and human-only teams worked on the same challenge set. Across the field, AI-augmented teams solved challenges at 3.2 times the human-only rate. Among the top 5%, that advantage narrowed to 1.69 times, even as the strongest AI-augmented teams retained their three-to-fourfold speed edge.

The measurement problem

The benchmark also records a faster competition overall. Median time to solve fell from 26.1 hours in 2024 to 13.8 hours in 2026, while full-board completions rose from two teams in 2024 and three in 2025 to 15 this year.

Neither trend is a clean measure of AI’s effect or a simple score for technical skill. Hack The Box says the multi-year data cannot isolate causes behind the shorter solve times. Challenge-board design changed as well, which means completion counts reflect more than one variable, including the breadth and coordination needed to work through an event.

The operating question

The evidence supports a narrower conclusion than either replacement or irrelevance. In Hack The Box’s framing, agents can compress research, coding, troubleshooting and first-pass analysis. Operators still choose the route, test weak output and decide whether a result is valid.

For security leaders, that leaves a practical training and deployment question: whether a team can direct and validate an agent’s work at the pace the agent creates. The benchmark shows adoption among highly capable practitioners. It does not yet show that giving agents to less experienced teams closes the gap.

Sources

  1. helpnetsecurity.comThe best human hacking team still out-solved the best AI team - Help Net Security
  2. itpro.comTop security teams use AI agents, says Hack The Box