After the Answer-Key Leak: Which SWE-bench Pro Scores Still Identify a Verifiable Dataset Version?
The public record establishes leakage and task-quality concerns, but the decisive procurement question is score-level provenance. In the bounded public sample, no visible SWE-bench Pro row preserves the complete metadata chain needed for strict historical or pre/post-remediation comparison.
This is a frozen public-record audit at the September 15, 2026 cutoff, not a rerun or a claim about unpublished vendor records. The sample is limited to 25 visible public leaderboard rows.
01Evidence matrix plan: treat the 25 rows visible on Scale’s public leaderboard at the September 15, 2026 cutoff as the bounded sample.
02For each score, test for five public receipts: row-level run date; immutable dataset/task manifest; immutable container or image revision; pinned harness, configuration, and tool or internet policy; and a row-specific downloadable run artifact.
03Record remediation status separately: whether a row is explicitly marked pre- or post-remediation for git-history leakage or task-quality changes.
04Compare visible leaderboard disclosures with Scale’s repository, historical leaderboard, dataset commit history, issue and pull-request records, OpenAI’s task-quality audit, and SWE-Bench Pro Verified.
05Classify a score as strictly reproducible only if the complete public receipt chain is present; otherwise classify it as inconclusive rather than invalid.
Sources
Evidence
5 publishers supporting 30 records. Expand a publisher to inspect its cited pages.