DeepSeek V4 Pro Lands on Fireworks With $2.50 CyberGym Solves, but Kimi K3 Scores Higher
The provider’s results frame security-agent selection as a three-way choice among raw solve rate, cost per completed task, and whether a model will reliably execute the required tool workflow.
Listen to this story
The audio brief
Story brief
3 key pointsFireworks’ release gives security-agent teams a lower-cost alternative to Kimi K3, not the top CyberGym solver. On 697 common-valid tasks, DeepSeek V4 Pro achieved a 53.7% reward rate at $2.502 per success, versus K3’s 68.4% at $4.914. It also entered validation on 97.3% of tasks and completed patches in 94.7% of those cases. Because the results come from the model’s platform provider, teams should validate the cost...
- 01
CyberGym tests 1,507 real vulnerabilities across 188 open-source projects, requiring executable exploits and patches—not just generated code.
- 02
DeepSeek V4 Pro’s 89-turn, 25-minute average tasks make sustained tool use central to the reported performance.
- 03
In a separate 920-task diagnostic, DeepSeek solved 50.2% at $2.60 per success, or $1.31 per run.
Fireworks AI has made DeepSeek-V4-Pro-0813 available on its platform, pairing the release with security-benchmark results that favor the model on cost and tool reliability rather than outright solve rate. In Fireworks’ 697-task CyberGym cohort, DeepSeek V4 Pro reached a 53.7% reward rate at $2.502 per solved task; Kimi K3 scored 68.4%, but cost $4.914 per success.
A benchmark built around finding and fixing real flaws
CyberGym is designed around 1,507 real vulnerabilities from 188 open-source projects. For each task, an agent receives a vulnerable codebase and a short description, then must locate the flaw, create an input that triggers it, and patch it. The benchmark grades by execution: the proof of concept must actually crash the target.
That setup makes the numbers more consequential than a one-shot code-generation score. Fireworks said its CyberGym runs averaged 89 turns and 25 minutes per task, making tool use and sustained task completion part of the test rather than incidental features.
The central trade-off: K3’s hit rate versus V4 Pro’s cost
DeepSeek V4 Pro
DeepSeek V4 Pro achieved a 53.7% reward rate and cost $2.502 per solved task in Fireworks’ 697-task common-valid cohort.
Kimi K3
Kimi K3 posted the cohort’s highest reported reward rate, 68.4%, while costing $4.914 per solved task.
DeepSeek sits between the closed-model results and K3
The comparison does not establish DeepSeek V4 Pro as the highest-performing model in the cohort. K3 led it by 14.7 percentage points in reward rate. But DeepSeek outperformed GPT-5.5, which recorded a 47.6% reward rate and $9.641 per solved task, and Claude Opus 4.8, which recorded 5.9% and $33.275, respectively.
For a team choosing a model strictly by completed-task rate, Fireworks’ figures point to K3. For a workload constrained by spending per resolved vulnerability, they point to DeepSeek V4 Pro. Those are different optimization targets, and the reported results do not eliminate the trade-off between them.
Reliability may be the sharper distinction for agent builders
Fireworks reported native tool calls from DeepSeek V4 Pro on all 840 traced primary tasks, with no observed refusals or output-length truncations. In the 697-task cohort, it entered validation on 678 tasks, or 97.3%.
Where the reported DeepSeek results break down
- After reaching stage-one validation, DeepSeek V4 Pro completed the full patch in 94.7% of cases, according to Fireworks.
- In Fireworks’ separate 920-task diagnostic, it solved 462 tasks, or 50.2%, at $1.31 per run and $2.60 per success.
The results are provider-reported testing from Fireworks, which now offers the model, rather than an independent benchmark release. Still, they isolate a practical decision for security-agent deployments: a model can be less likely to solve a task than the leader and yet be materially cheaper to run through a long, tool-dependent workflow. The unanswered deployment question is whether the reported cost and completion pattern holds across individual teams’ repositories, harnesses, and security workflows.
Sources
- fireworks.ai8/26/2026 DeepSeek V4 Pro is Redefining Security Agent Economics