{"title":"Vibe Code Bench’s Score Reset: Model Regression or a Different Test?","description":"The apparent collapse is not interpretable as a model-performance reversal from the public record alone: the earlier benchmark and VCB 1–100 change the task, scoring, budget, state, and evaluation structure.","dataset_id":"spd:vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565","canonical_url":"https://superpowerdaily.com/research/vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565","version_url":"https://superpowerdaily.com/research/vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565","version":"v1","snapshot_hash":"48cd49840a187d0f2cf1f763772b94060227be23c642e1c431b4d1a3fb474ac0","date_created":"2026-09-20T00:11:51.439Z","date_modified":"2026-09-20T00:11:51.439Z","license":{"name":"Superpower Daily data reuse terms","url":"https://superpowerdaily.com/terms"},"license_url":"https://superpowerdaily.com/terms","temporal_coverage":"2026-09-20","coverage_note":"The evidence matrix supports material changes in task structure, scoring, budget, state semantics, and test content. Public records do not provide a complete version-and-harness crosswalk or matched reruns that isolate model movement.","measurement_technique":["Evidence matrix: compared the earlier Vibe Code Bench paper, VCB 1–100 methodology and leaderboard, earlier scaffold and run-artifact repositories, a grading-version issue, and a secondary leaderboard mirror.","Classified each comparison field as materially changed, partially matched, or not publicly crosswalked: task structure, test content, scoring, budget, state/reset behavior, model labels, harness, and artifacts.","Used only records available through the September 20, 2026 UTC cutoff. Recovering saved public evidence was not a new collection or experiment.","Did not infer a model-only effect where benchmark and system conditions changed together."],"methodology":["Evidence matrix: compared the earlier Vibe Code Bench paper, VCB 1–100 methodology and leaderboard, earlier scaffold and run-artifact repositories, a grading-version issue, and a secondary leaderboard mirror.","Classified each comparison field as materially changed, partially matched, or not publicly crosswalked: task structure, test content, scoring, budget, state/reset behavior, model labels, harness, and artifacts.","Used only records available through the September 20, 2026 UTC cutoff. Recovering saved public evidence was not a new collection or experiment.","Did not infer a model-only effect where benchmark and system conditions changed together."],"metrics":[{"label":"Verified observations","value":"13","detail":"11 measured fields"},{"label":"Supported claims","value":"7","detail":"7 material findings"},{"label":"Cited sources","value":"6","detail":"5 primary or authoritative"},{"label":"Research score","value":"87","detail":"Automated topic and evidence score"}],"columns":[{"key":"entity","label":"Entity"},{"key":"metric","label":"Metric"},{"key":"value","label":"Value"},{"key":"unit","label":"Unit"},{"key":"observed","label":"Observed"},{"key":"source","label":"Source"},{"key":"transform","label":"Transform"}],"data":[{"unit":"arcs","value":"100","entity":"Vibe Code Bench 1–100","metric":"application arcs","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-20","transform":null},{"unit":"applications","value":"100","entity":"Vibe Code Bench","metric":"application specifications","source":"https://arxiv.org/html/2603.04601","observed":"2026-05-13","transform":null},{"unit":"workflows","value":"11,299","entity":"Vibe Code Bench 1–100","metric":"authored workflows","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-20","transform":null},{"unit":"workflows","value":"964","entity":"Vibe Code Bench","metric":"browser workflows","source":"https://arxiv.org/html/2603.04601","observed":"2026-05-13","transform":null},{"unit":"percent","value":"32.03%","entity":"Gemini 3.1 Pro Preview","metric":"earlier Vibe Code Bench accuracy","source":"https://arxiv.org/html/2603.04601","observed":"2026-05-13","transform":null},{"unit":"percent","value":"61.77%","entity":"GPT-5.3-Codex","metric":"earlier Vibe Code Bench accuracy","source":"https://arxiv.org/html/2603.04601","observed":"2026-05-13","transform":null},{"unit":"hours","value":"5 hours","entity":"Vibe Code Bench","metric":"generation wall-clock budget per application","source":"https://arxiv.org/html/2603.04601","observed":"2026-05-13","transform":null},{"unit":"iterations","value":"10","entity":"Vibe Code Bench 1–100","metric":"maximum sequential changes per arc","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-20","transform":null},{"unit":"requests","value":"939","entity":"Vibe Code Bench 1–100","metric":"ordered change requests","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-20","transform":null},{"unit":"percent","value":"33.16%","entity":"Vibe Code Bench 1–100","metric":"regression workflow share","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-20","transform":null},{"unit":"hours","value":"10 hours","entity":"Vibe Code Bench 1–100","metric":"shared generation budget per arc","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-20","transform":null},{"unit":"percent","value":"28.53%","entity":"Claude Opus 5","metric":"VCB 1–100 score","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-16","transform":null},{"unit":"percent","value":"6.69%","entity":"Gemini 3.1 Pro Preview (02/26)","metric":"VCB 1–100 score","source":"https://www.vals.ai/benchmarks/vcb-1-100","observed":"2026-09-16","transform":null}],"sources":[{"url":"https://arxiv.org/html/2603.04601","name":"arXiv / Vals AI","title":"arxiv.org","records":5},{"url":"https://github.com/vals-ai/create-benchmark-service/issues/138","name":"Vals AI","title":"evaluator_version ignores the get_service_version() override · Issue #138 · vals-ai/create-benchmark-service","records":0},{"url":"https://github.com/vals-ai/vibe-code-bench-cais-2026-artifacts","name":"Vals AI","title":"GitHub - vals-ai/vibe-code-bench-cais-2026-artifacts","records":0},{"url":"https://github.com/vals-ai/VibeCodeBench-Openhands-Scaffold","name":"Vals AI","title":"GitHub - vals-ai/VibeCodeBench-Openhands-Scaffold: A forked version of OpenHands designed for use with the Vals AI Vibe Code Bench benchmark.","records":0},{"url":"https://www.vals.ai/benchmarks/vcb-1-100","name":"Vals AI","title":"www.vals.ai","records":8},{"url":"https://benchlm.ai/benchmarks/vibe-code-bench-1-100","name":"BenchLM","title":"Vibe Code Bench 1-100 Leaderboard & Scores — September 2026","records":0}],"provenance":{"publisher":"Superpower Daily","source_count":6,"source_urls":["https://arxiv.org/html/2603.04601","https://github.com/vals-ai/create-benchmark-service/issues/138","https://github.com/vals-ai/vibe-code-bench-cais-2026-artifacts","https://github.com/vals-ai/VibeCodeBench-Openhands-Scaffold","https://www.vals.ai/benchmarks/vcb-1-100","https://benchlm.ai/benchmarks/vibe-code-bench-1-100"],"methodology":["Evidence matrix: compared the earlier Vibe Code Bench paper, VCB 1–100 methodology and leaderboard, earlier scaffold and run-artifact repositories, a grading-version issue, and a secondary leaderboard mirror.","Classified each comparison field as materially changed, partially matched, or not publicly crosswalked: task structure, test content, scoring, budget, state/reset behavior, model labels, harness, and artifacts.","Used only records available through the September 20, 2026 UTC cutoff. Recovering saved public evidence was not a new collection or experiment.","Did not infer a model-only effect where benchmark and system conditions changed together."],"snapshot_hash":"48cd49840a187d0f2cf1f763772b94060227be23c642e1c431b4d1a3fb474ac0"},"distributions":{"csv":"https://superpowerdaily.com/api/research/vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565?format=csv","json":"https://superpowerdaily.com/api/research/vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565?format=json"}}