Research observatory 19 / Evaluations under pressure
AI Benchmark Saturation Tracker
Tracks reported benchmark results, leaderboard movement, evaluation releases, and reliability concerns.
Public editionLive dataset
- Evaluation records
- 17
Published source-backed records
- Source references
- 34
34 distinct URLs shown
- Tracked dimensions
- 4
Agent and coding / Safety evaluations / Leaderboards
- Median importance
- 54
Editorial importance score
Data status17 verified records across 1 period
Snapshot only. There is not enough history to claim a trend yet.
Verified observationHover or focus any mark for exact valuesData and chart downloads share this dataset
Coverage note
This tracker records reported benchmark developments, not a recomputed universal model ranking. Scores from unlike evaluations are never normalized into one synthetic capability number.
- Dataset ID
- spd:benchmark-saturation
- Coverage
- 2026-08/2026-08-19
- Records
- 17
- Fields
- 7
- Formats
- CSV / JSON
- Updated
| Published | Record | Dimension | Category | Importance | Sources | Story URL |
|---|---|---|---|---|---|---|
| 2026-08-17 | Meta Launches Muse Code With a 92% Discount for Training Rights | Agent and coding | models | 69 | 1 | /posts/meta-launches-muse-code-with-a-92-discount-for-training-rights |
| 2026-08-17 | Anthropic’s Dario Amodei Says AI Must Deliver, Not Advertise, Its Way Out of a Trust Crisis | Safety evaluations | culture | 39 | 3 | /posts/anthropic-s-dario-amodei-says-ai-must-deliver-not-advertise-its-way-out-of-a-trust-crisis |
| 2026-08-17 | Zuckerberg’s superintelligence promise runs into AI’s trust test | Safety evaluations | culture | 65 | 1 | /posts/zuckerberg-s-superintelligence-promise-runs-into-ai-s-trust-test |
| 2026-08-17 | A Naming Error Let Anthropic Models Reach a Real Production Database | Safety evaluations | models | 46 | 2 | /posts/a-naming-error-let-anthropic-models-reach-a-real-production-database |
| 2026-08-17 | Alibaba Launches HappyShrimp for One-Prompt Songs, but Leaves Key Use Terms Unstated | Leaderboards | models | 46 | 1 | /posts/alibaba-launches-happyshrimp-for-one-prompt-songs-but-leaves-key-use-terms-unstated |
| 2026-08-18 | Zhipu Says GLM-5.3 Can Find Bugs Across an Exploitation Chain | Safety evaluations | models | 72 | 2 | /posts/zhipu-says-glm-5-3-can-find-bugs-across-an-exploitation-chain |
| 2026-08-18 | Alibaba Releases Downloadable Qwen3.8-Max Weights Alongside a Laptop Model | Agent and coding | models | 63 | 3 | /posts/alibaba-puts-qwen-on-laptops-and-opens-its-biggest-model-as-meta-courts-the-same-developers |
| 2026-08-18 | Nvidia’s Reported Lancium Deal Would Put Grid Connections at the Center of Its AI Buildout | Leaderboards | business | 35 | 3 | /posts/nvidia-s-ai-infrastructure-bet-reaches-the-grid-not-just-the-data-center |
| 2026-08-18 | ByteDance’s Hollywood Copyright Truce Leaves the Guardrails Private | Leaderboards | policy | 53 | 4 | /posts/hollywood-and-bytedance-trade-a-copyright-fight-for-an-unseen-ai-guardrail-deal |
| 2026-08-18 | Snowflake Wants AI Apps to Stop Paying Frontier-Model Prices for Every Task | Agent and coding | products | 77 | 2 | /posts/snowflake-wants-ai-apps-to-stop-paying-frontier-model-prices-for-every-task |
| 2026-08-18 | OpenAI Funds a Study of How AI Could Shift Tax Revenue | Safety evaluations | policy | 41 | 1 | /posts/openai-funds-a-study-of-how-ai-could-shift-tax-revenue |
| 2026-08-18 | OpenAI Paused Deployment-Bound Training as It Tightened Frontier Security | Safety evaluations | policy | 8 | 1 | /posts/openai-paused-deployment-bound-training-as-it-tightened-frontier-security |
| 2026-08-18 | Etched’s $700 Million Round Puts a $21 Billion Price on Inference Hardware | Capability evaluations | startups | 54 | 3 | /posts/etched-s-700-million-round-puts-a-21-billion-price-on-inference-hardware |
| 2026-08-18 | Wispr’s $280 Million Bet Is That Voice Can Move Beyond Dictation | Leaderboards | startups | 56 | 4 | /posts/wispr-s-280-million-bet-is-that-voice-can-move-beyond-dictation |
| 2026-08-19 | AgentX Replays Claude Code Sessions to Test AI Serving Systems | Agent and coding | tools | 72 | 1 | /posts/agentx-replays-claude-code-sessions-to-test-ai-serving-systems |
| 2026-08-19 | TestMu’s New Agent Test Shows What It Can’t Verify | Safety evaluations | products | 43 | 1 | /posts/testmu-s-new-agent-test-shows-what-it-can-t-verify |
| 2026-08-19 | LMCache Reworks the Cache Plumbing That Can Stall Long-Running Agents | Agent and coding | tools | 65 | 1 | /posts/lmcache-reworks-the-cache-plumbing-that-can-stall-long-running-agents |
Methodology
How to read this report
- 01Records require a named evaluation, benchmark result, leaderboard change, or benchmark-integrity development.
- 02The series counts developments and groups them by evaluation role; it does not compare incompatible score scales.
- 03Primary benchmark papers and maintainers remain authoritative for exact protocols and revisions.
Sources
Evidence
28 publishers supporting 34 records. Expand a publisher to inspect its cited pages.