Fireworks AI has made DeepSeek-V4-Pro-0813 available on its platform, pairing the release with security-benchmark results that favor the model on cost and tool reliability rather than outright solve rate. In Fireworks’ 697-task CyberGym cohort, DeepSeek V4 Pro reached a 53.7% reward rate at $2.502 per solved task; Kimi K3 scored 68.4%, but cost $4.914 per success.
A benchmark built around finding and fixing real flaws
CyberGym is designed around 1,507 real vulnerabilities from 188 open-source projects. For each task, an agent receives a vulnerable codebase and a short description, then must locate the flaw, create an input that triggers it, and patch it. The benchmark grades by execution: the proof of concept must actually crash the target.
That setup makes the numbers more consequential than a one-shot code-generation score. Fireworks said its CyberGym runs averaged 89 turns and 25 minutes per task, making tool use and sustained task completion part of the test rather than incidental features.
The central trade-off: K3’s hit rate versus V4 Pro’s cost
0153.7% reward rateDeepSeek V4 Pro
DeepSeek V4 Pro achieved a 53.7% reward rate and cost $2.502 per solved task in Fireworks’ 697-task common-valid cohort.
0268.4% reward rateKimi K3
Kimi K3 posted the cohort’s highest reported reward rate, 68.4%, while costing $4.914 per solved task.
DeepSeek sits between the closed-model results and K3
The comparison does not establish DeepSeek V4 Pro as the highest-performing model in the cohort. K3 led it by 14.7 percentage points in reward rate. But DeepSeek outperformed GPT-5.5, which recorded a 47.6% reward rate and $9.641 per solved task, and Claude Opus 4.8, which recorded 5.9% and $33.275, respectively.
For a team choosing a model strictly by completed-task rate, Fireworks’ figures point to K3. For a workload constrained by spending per resolved vulnerability, they point to DeepSeek V4 Pro. Those are different optimization targets, and the reported results do not eliminate the trade-off between them.
Reliability may be the sharper distinction for agent builders
Fireworks reported native tool calls from DeepSeek V4 Pro on all 840 traced primary tasks, with no observed refusals or output-length truncations. In the 697-task cohort, it entered validation on 678 tasks, or 97.3%.
Where the reported DeepSeek results break down
- After reaching stage-one validation, DeepSeek V4 Pro completed the full patch in 94.7% of cases, according to Fireworks.
- In Fireworks’ separate 920-task diagnostic, it solved 462 tasks, or 50.2%, at $1.31 per run and $2.60 per success.
The results are provider-reported testing from Fireworks, which now offers the model, rather than an independent benchmark release. Still, they isolate a practical decision for security-agent deployments: a model can be less likely to solve a task than the leader and yet be materially cheaper to run through a long, tool-dependent workflow. The unanswered deployment question is whether the reported cost and completion pattern holds across individual teams’ repositories, harnesses, and security workflows.
Reader comments
Newest comments first. Replies stay oldest first.