Modelspublished

DeepSeek V4 Pro Lands on Fireworks With $2.50 CyberGym Solves, but Kimi K3 Scores Higher

The provider’s results frame security-agent selection as a three-way choice among raw solve rate, cost per completed task, and whether a model will reliably execute the required tool workflow.

By 3 min read
DeepSeek V4 Pro Lands on Fireworks With $2.50 CyberGym Solves, but Kimi K3 Scores Higher

Listen to this story

The audio brief

About 1:38
0:001:38
Read transcript
Fireworks AI has put DeepSeek V4 Pro into production, and its accompanying security results make the model’s appeal clear: not the highest solve rate, but a much lower cost for long, tool-heavy jobs. In a 697-task CyberGym cohort, DeepSeek solved 53.7 percent of tasks at an average cost of $2.502 per success. Kimi K3 solved 68.4 percent, leading on raw performance, but cost nearly twice as much: $4.914 per completed task. CyberGym is testing more than code generation. Across 1,507 real vulnerabilities in 188 open-source projects, an agent has to inspect a vulnerable codebase, create an executable exploit, and then produce a patch. Fireworks says DeepSeek’s runs averaged 89 turns and 25 minutes, so sustained tool use is central to the result. The model entered validation on 97.3 percent of tasks, and completed the patch in 94.7 percent of those cases. That puts the choice into three dimensions: solve rate, cost per success, and dependable execution. DeepSeek also beat GPT-5.5, at 47.6 percent and $9.641 per success, and Claude Opus 4.8, at just 5.9 percent and $33.275. A separate 920-task diagnostic found DeepSeek solving 50.2 percent at $2.60 per success. The constraint is that these are provider-reported results. The key question is whether the same cost and completion pattern survives on each team’s repositories, harnesses, and security workflows.

Story brief

3 key points

Fireworks’ release gives security-agent teams a lower-cost alternative to Kimi K3, not the top CyberGym solver. On 697 common-valid tasks, DeepSeek V4 Pro achieved a 53.7% reward rate at $2.502 per success, versus K3’s 68.4% at $4.914. It also entered validation on 97.3% of tasks and completed patches in 94.7% of those cases. Because the results come from the model’s platform provider, teams should validate the cost...

  1. 01

    CyberGym tests 1,507 real vulnerabilities across 188 open-source projects, requiring executable exploits and patches—not just generated code.

  2. 02

    DeepSeek V4 Pro’s 89-turn, 25-minute average tasks make sustained tool use central to the reported performance.

  3. 03

    In a separate 920-task diagnostic, DeepSeek solved 50.2% at $2.60 per success, or $1.31 per run.

Fireworks AI has made DeepSeek-V4-Pro-0813 available on its platform, pairing the release with security-benchmark results that favor the model on cost and tool reliability rather than outright solve rate. In Fireworks’ 697-task CyberGym cohort, DeepSeek V4 Pro reached a 53.7% reward rate at $2.502 per solved task; Kimi K3 scored 68.4%, but cost $4.914 per success.

A benchmark built around finding and fixing real flaws

CyberGym is designed around 1,507 real vulnerabilities from 188 open-source projects. For each task, an agent receives a vulnerable codebase and a short description, then must locate the flaw, create an input that triggers it, and patch it. The benchmark grades by execution: the proof of concept must actually crash the target.

That setup makes the numbers more consequential than a one-shot code-generation score. Fireworks said its CyberGym runs averaged 89 turns and 25 minutes per task, making tool use and sustained task completion part of the test rather than incidental features.

The central trade-off: K3’s hit rate versus V4 Pro’s cost

0153.7% reward rate

DeepSeek V4 Pro

DeepSeek V4 Pro achieved a 53.7% reward rate and cost $2.502 per solved task in Fireworks’ 697-task common-valid cohort.

0268.4% reward rate

Kimi K3

Kimi K3 posted the cohort’s highest reported reward rate, 68.4%, while costing $4.914 per solved task.

DeepSeek sits between the closed-model results and K3

The comparison does not establish DeepSeek V4 Pro as the highest-performing model in the cohort. K3 led it by 14.7 percentage points in reward rate. But DeepSeek outperformed GPT-5.5, which recorded a 47.6% reward rate and $9.641 per solved task, and Claude Opus 4.8, which recorded 5.9% and $33.275, respectively.

For a team choosing a model strictly by completed-task rate, Fireworks’ figures point to K3. For a workload constrained by spending per resolved vulnerability, they point to DeepSeek V4 Pro. Those are different optimization targets, and the reported results do not eliminate the trade-off between them.

Reliability may be the sharper distinction for agent builders

Fireworks reported native tool calls from DeepSeek V4 Pro on all 840 traced primary tasks, with no observed refusals or output-length truncations. In the 697-task cohort, it entered validation on 678 tasks, or 97.3%.

Where the reported DeepSeek results break down

  • After reaching stage-one validation, DeepSeek V4 Pro completed the full patch in 94.7% of cases, according to Fireworks.
  • In Fireworks’ separate 920-task diagnostic, it solved 462 tasks, or 50.2%, at $1.31 per run and $2.60 per success.

The results are provider-reported testing from Fireworks, which now offers the model, rather than an independent benchmark release. Still, they isolate a practical decision for security-agent deployments: a model can be less likely to solve a task than the leader and yet be materially cheaper to run through a long, tool-dependent workflow. The unanswered deployment question is whether the reported cost and completion pattern holds across individual teams’ repositories, harnesses, and security workflows.

Sources

  1. fireworks.ai8/26/2026 DeepSeek V4 Pro is Redefining Security Agent Economics