Research observatory 19 / Evaluations under pressure

AI Benchmark Saturation Tracker

Tracks reported benchmark results, leaderboard movement, evaluation releases, and reliability concerns.

Public editionLive dataset
Evaluation records
17

Published source-backed records

Source references
34

34 distinct URLs shown

Tracked dimensions
4

Agent and coding / Safety evaluations / Leaderboards

Median importance
54

Editorial importance score

Interactive figureBenchmark saturation
CSV JSON
Data status17 verified records across 1 period

Snapshot only. There is not enough history to claim a trend yet.

Verified observationHover or focus any mark for exact valuesData and chart downloads share this dataset
Coverage note

This tracker records reported benchmark developments, not a recomputed universal model ranking. Scores from unlike evaluations are never normalized into one synthetic capability number.

Dataset ID
spd:benchmark-saturation
Coverage
2026-08/2026-08-19
Records
17
Fields
7
Formats
CSV / JSON
Updated

Read the data

The records behind the figure

CSV JSON
AI Benchmark Saturation Tracker data records
PublishedRecordDimensionCategoryImportanceSourcesStory URL
2026-08-17Meta Launches Muse Code With a 92% Discount for Training RightsAgent and codingmodels691/posts/meta-launches-muse-code-with-a-92-discount-for-training-rights
2026-08-17Anthropic’s Dario Amodei Says AI Must Deliver, Not Advertise, Its Way Out of a Trust CrisisSafety evaluationsculture393/posts/anthropic-s-dario-amodei-says-ai-must-deliver-not-advertise-its-way-out-of-a-trust-crisis
2026-08-17Zuckerberg’s superintelligence promise runs into AI’s trust testSafety evaluationsculture651/posts/zuckerberg-s-superintelligence-promise-runs-into-ai-s-trust-test
2026-08-17A Naming Error Let Anthropic Models Reach a Real Production DatabaseSafety evaluationsmodels462/posts/a-naming-error-let-anthropic-models-reach-a-real-production-database
2026-08-17Alibaba Launches HappyShrimp for One-Prompt Songs, but Leaves Key Use Terms UnstatedLeaderboardsmodels461/posts/alibaba-launches-happyshrimp-for-one-prompt-songs-but-leaves-key-use-terms-unstated
2026-08-18Zhipu Says GLM-5.3 Can Find Bugs Across an Exploitation ChainSafety evaluationsmodels722/posts/zhipu-says-glm-5-3-can-find-bugs-across-an-exploitation-chain
2026-08-18Alibaba Releases Downloadable Qwen3.8-Max Weights Alongside a Laptop ModelAgent and codingmodels633/posts/alibaba-puts-qwen-on-laptops-and-opens-its-biggest-model-as-meta-courts-the-same-developers
2026-08-18Nvidia’s Reported Lancium Deal Would Put Grid Connections at the Center of Its AI BuildoutLeaderboardsbusiness353/posts/nvidia-s-ai-infrastructure-bet-reaches-the-grid-not-just-the-data-center
2026-08-18ByteDance’s Hollywood Copyright Truce Leaves the Guardrails PrivateLeaderboardspolicy534/posts/hollywood-and-bytedance-trade-a-copyright-fight-for-an-unseen-ai-guardrail-deal
2026-08-18Snowflake Wants AI Apps to Stop Paying Frontier-Model Prices for Every TaskAgent and codingproducts772/posts/snowflake-wants-ai-apps-to-stop-paying-frontier-model-prices-for-every-task
2026-08-18OpenAI Funds a Study of How AI Could Shift Tax RevenueSafety evaluationspolicy411/posts/openai-funds-a-study-of-how-ai-could-shift-tax-revenue
2026-08-18OpenAI Paused Deployment-Bound Training as It Tightened Frontier SecuritySafety evaluationspolicy81/posts/openai-paused-deployment-bound-training-as-it-tightened-frontier-security
2026-08-18Etched’s $700 Million Round Puts a $21 Billion Price on Inference HardwareCapability evaluationsstartups543/posts/etched-s-700-million-round-puts-a-21-billion-price-on-inference-hardware
2026-08-18Wispr’s $280 Million Bet Is That Voice Can Move Beyond DictationLeaderboardsstartups564/posts/wispr-s-280-million-bet-is-that-voice-can-move-beyond-dictation
2026-08-19AgentX Replays Claude Code Sessions to Test AI Serving SystemsAgent and codingtools721/posts/agentx-replays-claude-code-sessions-to-test-ai-serving-systems
2026-08-19TestMu’s New Agent Test Shows What It Can’t VerifySafety evaluationsproducts431/posts/testmu-s-new-agent-test-shows-what-it-can-t-verify
2026-08-19LMCache Reworks the Cache Plumbing That Can Stall Long-Running AgentsAgent and codingtools651/posts/lmcache-reworks-the-cache-plumbing-that-can-stall-long-running-agents

Methodology

How to read this report

  1. 01Records require a named evaluation, benchmark result, leaderboard change, or benchmark-integrity development.
  2. 02The series counts developments and groups them by evaluation role; it does not compare incompatible score scales.
  3. 03Primary benchmark papers and maintainers remain authoritative for exact protocols and revisions.

Sources

Evidence

28 publishers supporting 34 records. Expand a publisher to inspect its cited pages.

finance.yahoo.comfinance.yahoo.com3 sources / 3 records
TechCrunchtechcrunch.com3 sources / 3 records
CNBCcnbc.com2 sources / 2 records
inferencex.semianalysis.cominferencex.semianalysis.com2 sources / 2 records
businesswire.combusinesswire.com1 source / 1 record
cnbctv18.comcnbctv18.com1 source / 1 record
Show 22 more publishers
cryptonews.netcryptonews.net1 source / 1 record
entrackr.comentrackr.com1 source / 1 record
finance.biggo.comfinance.biggo.com1 source / 1 record
forbes.comforbes.com1 source / 1 record
fortune.comfortune.com1 source / 1 record
latimes.comlatimes.com1 source / 1 record
marketscale.commarketscale.com1 source / 1 record
mediapost.commediapost.com1 source / 1 record
news.bloombergtax.comnews.bloombergtax.com1 source / 1 record
nypost.comnypost.com1 source / 1 record
openai.comopenai.com1 source / 1 record
pymnts.compymnts.com1 source / 1 record
qz.comqz.com1 source / 1 record
scmp.comscmp.com1 source / 1 record
securityweek.comsecurityweek.com1 source / 1 record
snowflake.comsnowflake.com1 source / 1 record
startupfortune.comstartupfortune.com1 source / 1 record
techbuzz.aitechbuzz.ai1 source / 1 record
thenextweb.comthenextweb.com1 source / 1 record
theregister.comtheregister.com1 source / 1 record
tvtechnology.comtvtechnology.com1 source / 1 record
wmbdradio.comwmbdradio.com1 source / 1 record
Next report / 20AI Partnership and Dependency Graph All research reports