Research observatory 19 / Evaluations under pressure

AI Benchmark Saturation Tracker

Tracks reported benchmark results, leaderboard movement, evaluation releases, and reliability concerns.

Archived snapshotv5Aug 27, 2026
Evaluation records
141

Published source-backed records

Source references
60

60 distinct URLs shown

Tracked dimensions
4

Agent and coding / Safety evaluations / Leaderboards

Median importance
68

Editorial importance score

Interactive figureBenchmark saturation
CSV JSON
Data status141 verified records across 1 period

Snapshot only. There is not enough history to claim a trend yet.

Verified observationHover or focus any mark for exact valuesLast updated Aug 27, 2026

Version ledger

Frozen public editions

Each edition preserves the records, method, sources, and downloads available at publication time.

  1. v8 / latestAug 27, 2026144 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 144 records.

  2. v7Aug 27, 2026143 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 143 records.

  3. v6Aug 27, 2026142 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 142 records.

  4. v5Aug 27, 2026141 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 141 records.

  5. v4Aug 27, 2026140 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 140 records.

  6. v3Aug 27, 2026139 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 139 records.

  7. v2Aug 27, 2026138 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 138 records.

  8. v1Aug 27, 2026137 records / 60 sources

    Initial public snapshot with 137 records and 60 cited sources.

Coverage note

This tracker records reported benchmark developments, not a recomputed universal model ranking. Scores from unlike evaluations are never normalized into one synthetic capability number.

Dataset ID
spd:benchmark-saturation
Stable URL
/research/benchmark-saturation
Version
v5
Coverage
2026-08/2026-08-27
Records
141
Fields
7
Updated

Read the data

The records behind the figure

CSV JSON
AI Benchmark Saturation Tracker data records
PublishedRecordDimensionCategoryImportanceSourcesStory URL
2026-08-17Meta Launches Muse Code With a 92% Discount for Training RightsAgent and codingmodels691/posts/meta-launches-muse-code-with-a-92-discount-for-training-rights
2026-08-17Anthropic’s Dario Amodei Says AI Must Deliver, Not Advertise, Its Way Out of a Trust CrisisSafety evaluationsculture393/posts/anthropic-s-dario-amodei-says-ai-must-deliver-not-advertise-its-way-out-of-a-trust-crisis
2026-08-17Zuckerberg’s superintelligence promise runs into AI’s trust testSafety evaluationsculture651/posts/zuckerberg-s-superintelligence-promise-runs-into-ai-s-trust-test
2026-08-17A Naming Error Let Anthropic Models Reach a Real Production DatabaseSafety evaluationsmodels462/posts/a-naming-error-let-anthropic-models-reach-a-real-production-database
2026-08-17Alibaba Launches HappyShrimp for One-Prompt Songs, but Leaves Key Use Terms UnstatedLeaderboardsmodels461/posts/alibaba-launches-happyshrimp-for-one-prompt-songs-but-leaves-key-use-terms-unstated
2026-08-18Zhipu Says GLM-5.3 Can Find Bugs Across an Exploitation ChainSafety evaluationsmodels722/posts/zhipu-says-glm-5-3-can-find-bugs-across-an-exploitation-chain
2026-08-18Alibaba Releases Downloadable Qwen3.8-Max Weights Alongside a Laptop ModelAgent and codingmodels633/posts/alibaba-puts-qwen-on-laptops-and-opens-its-biggest-model-as-meta-courts-the-same-developers
2026-08-18Nvidia’s Reported Lancium Deal Would Put Grid Connections at the Center of Its AI BuildoutLeaderboardsbusiness353/posts/nvidia-s-ai-infrastructure-bet-reaches-the-grid-not-just-the-data-center
2026-08-18Snowflake Wants AI Apps to Stop Paying Frontier-Model Prices for Every TaskAgent and codingproducts772/posts/snowflake-wants-ai-apps-to-stop-paying-frontier-model-prices-for-every-task
2026-08-18OpenAI Funds a Study of How AI Could Shift Tax RevenueSafety evaluationspolicy411/posts/openai-funds-a-study-of-how-ai-could-shift-tax-revenue
2026-08-18OpenAI Paused Deployment-Bound Training as It Tightened Frontier SecuritySafety evaluationspolicy81/posts/openai-paused-deployment-bound-training-as-it-tightened-frontier-security
2026-08-18Etched’s $700 Million Round Puts a $21 Billion Price on Inference HardwareCapability evaluationsstartups543/posts/etched-s-700-million-round-puts-a-21-billion-price-on-inference-hardware
2026-08-18Wispr’s $280 Million Bet Is That Voice Can Move Beyond DictationLeaderboardsstartups564/posts/wispr-s-280-million-bet-is-that-voice-can-move-beyond-dictation
2026-08-19AgentX Replays Claude Code Sessions to Test AI Serving SystemsAgent and codingtools721/posts/agentx-replays-claude-code-sessions-to-test-ai-serving-systems
2026-08-19TestMu’s New Agent Test Shows What It Can’t VerifySafety evaluationsproducts431/posts/testmu-s-new-agent-test-shows-what-it-can-t-verify
2026-08-19LMCache Reworks the Cache Plumbing That Can Stall Long-Running AgentsAgent and codingtools651/posts/lmcache-reworks-the-cache-plumbing-that-can-stall-long-running-agents
2026-08-19Dane County Routes Non-Emergency Calls Through AVA to Protect 911 CapacitySafety evaluationsproducts384/posts/dane-county-puts-an-ai-gatekeeper-on-its-non-emergency-line
2026-08-19Oakley Buys Majority Stake in Graphwise as It Plans Expansion and AcquisitionsAgent and codingbusiness442/posts/oakley-takes-majority-control-of-graphwise-as-it-plans-a-global-ai-push
2026-08-19Jinho Jang Puts a 27B Refusal-Removed Model Into a Local DownloadSafety evaluationsmodels800/posts/jinho-jang-puts-a-27b-refusal-removed-model-into-a-local-download
2026-08-19Nome Uses AI to Find Rare-Disease Treatment Paths Before the Hard Work BeginsAgent and codingstartups651/posts/nome-promises-rare-disease-families-a-treatment-path-the-clinical-work-still-lies-ahead
2026-08-19Anthropic Will Watermark Claude Text by Steering Its Word ChoicesCapability evaluationstools581/posts/anthropic-will-watermark-claude-text-by-steering-its-word-choices
2026-08-19A Reporter’s LLM Wiki Speeds Recall. He Still Checks the Original Sources.Agent and codingtools551/posts/casey-newton-built-an-ai-wiki-for-his-beat-now-he-has-to-keep-it-alive
2026-08-19Vivodyne Builds a Human-Tissue Data Factory for Drug AICapability evaluationsstartups681/posts/vivodyne-built-a-robot-lab-for-the-human-biology-data-ai-still-lacks
2026-08-20SandboxAQ Opens a Drug-Screening Model That Does Not Need Protein StructuresCapability evaluationsmodels780/posts/sandboxaq-opens-a-drug-screening-model-that-does-not-need-protein-structures
2026-08-20OpenAI Tests Cross-Session Safety Checks Without Prompt AccessSafety evaluationsproducts760/posts/openai-tests-cross-session-safety-checks-without-prompt-access
2026-08-20OpenAI Takes Codex From Coding to Tax ReturnsAgent and codingtools680/posts/openai-takes-codex-from-coding-to-tax-returns
2026-08-20TrueFoundry Gives Away Its Agent Runtime to Sell the Layer Beneath ItAgent and codingtools650/posts/truefoundry-gives-away-its-agent-runtime-to-sell-the-layer-beneath-it
2026-08-20Palomar Opens a Lean Registry That Checks Proofs and Their DescriptionsCapability evaluationstools720/posts/palomar-opens-a-lean-proof-registry-with-an-llm-semantic-check
2026-08-20UK Cyber Tests Show AI Agents Going Beyond the Technical TaskSafety evaluationsmodels840/posts/ai-cyber-tests-are-reaching-real-targets-not-just-sandboxes
2026-08-20At This D.C. Charter, AI Permission Changes With the AssignmentLeaderboardsculture652/posts/this-d-c-charter-made-ai-a-schoolwide-skill-not-a-shortcut
2026-08-20Mistral Gives Enterprise AI Five Ways to Keep Digging Through DocumentsAgent and codingproducts680/posts/mistral-gives-enterprise-ai-five-ways-to-keep-digging-through-documents
2026-08-20Simple AI Opens 2,000 Hours of Robot Training Data Collected Without RobotsLeaderboardstools670/posts/simple-ai-opens-2-000-hours-of-robot-training-data-collected-without-robots
2026-08-20Ramp Opens Its AI Router to U.S. Customers, With a One-Year Data DefaultCapability evaluationsproducts581/posts/ramp-opens-its-ai-router-to-u-s-customers-with-a-one-year-data-default
2026-08-20Anthropic Opens Claude Academy as Free Training for Its AI ProductsAgent and codingproducts680/posts/anthropic-opens-claude-academy-as-free-training-for-its-ai-products
2026-08-21Micro1’s Reported $500M Run Rate Tests the Economics of AI Training DataLeaderboardsstartups551/posts/micro1-s-reported-500m-run-rate-tests-the-economics-of-ai-training-data
2026-08-21Callosum Raises $100 Million as It Routes AI Work Across Models and ChipsCapability evaluationsstartups662/posts/callosum-raises-100-million-to-route-ai-work-across-models-and-chips
2026-08-21Starcloud Has $250 Million for Orbital AI. It Still Needs a Ride.Leaderboardsstartups661/posts/starcloud-has-250-million-for-orbital-ai-it-still-needs-a-ride
2026-08-21Open Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest.Agent and codingmodels741/posts/open-models-are-catching-the-frontier-faster-benchmark-scores-aren-t-the-whole-contest
2026-08-21Nvidia Maps AI Memory Between Models Instead of Making Them Start OverCapability evaluationstools741/posts/nvidia-maps-ai-memory-between-models-instead-of-making-them-start-over
2026-08-21Nvidia’s AVO Clears ARC-AGI-3’s Public Set. Withheld Tests Still Matter.Agent and codingmodels721/posts/nvidia-s-avo-clears-arc-agi-3-s-public-set-withheld-tests-still-matter
2026-08-21OpenAI Makes GPT-5.6 Sol Cheaper for Metered Use, Not Easier to AccessCapability evaluationsproducts720/posts/openai-cuts-sol-s-output-price-but-not-chatgpt-s-limits
2026-08-21Grok Reversed a China-Influence Finding After an Audit of Its SourcesAgent and codingmodels650/posts/grok-reversed-its-china-campaign-finding-after-an-audit-exposed-the-citation-chain
2026-08-21Anthropic’s IPO Pitch Faces Two Tests: Compute Growth and Enterprise Data ControlSafety evaluationsbusiness772/posts/anthropic-s-ipo-pitch-faces-two-tests-compute-growth-and-enterprise-data-control
2026-08-22DeepSeek Gives V4-Flash a Separate Vision API With a 384-Token Image CapAgent and codingmodels711/posts/deepseek-adds-vision-to-an-experimental-v4-flash-endpoint-with-a-384-token-image-cap
2026-08-22Meta Is a Major Microsoft AI Customer—and a Potential Foundry RivalLeaderboardsbusiness751/posts/meta-is-a-major-microsoft-ai-customer-and-a-potential-foundry-rival
2026-08-22Meta’s 30B Muse Glimmer Tries to Make 131K Context Fit in 24 GBAgent and codingmodels451/posts/meta-s-30b-muse-glimmer-tries-to-make-131k-context-fit-in-24-gb
2026-08-22AI21’s 8B Verifier Challenges the Case for Bigger Search ModelsAgent and codingmodels681/posts/ai21-s-8b-verifier-challenges-the-case-for-bigger-search-models
2026-08-22AWS’s RAG Cost Cut Comes With a 19% Latency BillCapability evaluationstools621/posts/aws-s-rag-cost-cut-comes-with-a-19-latency-bill
2026-08-22GLM-5.3’s Cheap Retries Put Fable 5’s Coding Premium Under PressureAgent and codingmodels641/posts/glm-5-3-s-cheap-retries-put-fable-5-s-coding-premium-under-pressure
2026-08-22Oracle Wants AI Agents to Pick Trusted Reports, Not Write SQLAgent and codingtools571/posts/oracle-wants-ai-agents-to-pick-trusted-reports-not-write-sql
2026-08-22Thinking Machines Releases Inkling, but Argues Open Weights Need a GateSafety evaluationspolicy721/posts/thinking-machines-releases-inkling-but-argues-open-weights-need-a-gate
2026-08-22FDA Opens Door to Clinician-Style Tests for Medical AISafety evaluationspolicy452/posts/fda-opens-door-to-clinician-style-tests-for-medical-ai
2026-08-22Databricks Says Its New Extraction Mode Beats Frontier Models on the Documents That Break ThemAgent and codingproducts721/posts/databricks-says-its-new-extraction-mode-beats-frontier-models-on-the-documents-that-break-them
2026-08-22Databricks Pushes Feature Stores From Batch Lag to 200ms FreshnessCapability evaluationsproducts721/posts/databricks-pushes-feature-stores-from-batch-lag-to-200ms-freshness
2026-08-22Stanford Won Databricks’ Agent Cup, but 18.8% of Questions Stumped Every TeamAgent and codingmodels712/posts/stanford-won-databricks-agent-cup-but-18-8-of-questions-stumped-every-team
2026-08-22Ox Alpha Offers a Million Tokens, but Not a NameAgent and codingmodels651/posts/ox-alpha-offers-a-million-tokens-but-not-a-name
2026-08-22Roblox Opens Three AI Safety Models for Child-Protection TeamsSafety evaluationstools681/posts/roblox-opens-three-ai-safety-models-for-child-protection-teams
2026-08-22Serval’s Catalyst Pitches AI-Built Workflows as a ServiceNow ReplacementAgent and codingstartups581/posts/serval-s-catalyst-pitches-ai-built-workflows-as-a-servicenow-replacement
2026-08-22Microsoft and Qcells Want AI Data Centers to Bring Their Own PowerLeaderboardsbusiness722/posts/microsoft-and-qcells-want-ai-data-centers-to-bring-their-own-power
2026-08-22Veeda AI Raises $90M to Make Robot Training Less PhysicalAgent and codingstartups682/posts/veeda-ai-raises-90m-to-make-robot-training-less-physical
2026-08-22LinkedIn Measures AI Code Review Against Merged CodeAgent and codingbusiness671/posts/linkedin-uses-multiple-ai-reviewers-to-cut-code-review-noise
2026-08-22TrueForge Puts the Agent Harness, Not the Model, at the Center of Cost ControlAgent and codingmodels720/posts/trueforge-puts-the-agent-harness-not-the-model-at-the-center-of-cost-control
2026-08-22Hollywood’s Downturn Is Turning Creative Know-How Into AI Training DataLeaderboardsculture681/posts/hollywood-s-downturn-is-turning-creative-know-how-into-ai-training-data
2026-08-22Panasonic Says Its Aircraft AI Cut Diagnostic Investigations From Hours to MinutesAgent and codingbusiness631/posts/panasonic-says-its-aircraft-ai-cut-diagnostic-investigations-from-hours-to-minutes
2026-08-22AgentFlo’s Sales Agents Put Rules and Data Above the ModelAgent and codingproducts641/posts/agentflo-s-sales-agents-put-rules-and-data-above-the-model
2026-08-22Generalist’s GEN-1.5 Lets Robots Try a Task After Watching OnceCapability evaluationsmodels681/posts/generalist-s-gen-1-5-lets-robots-try-a-task-after-watching-once
2026-08-22AI Safety Scores Can Reward Models for Refusing Too MuchSafety evaluationsmodels701/posts/ai-safety-scores-can-reward-models-for-refusing-too-much
2026-08-22Massachusetts Could Make AI Labs Face Public Risk Tests Every 120 DaysSafety evaluationspolicy782/posts/massachusetts-could-make-ai-labs-face-public-risk-tests-every-120-days
2026-08-22AWS Wants AI Agents to Carry the User’s Permissions, Not Their OwnAgent and codingtools702/posts/aws-wants-ai-agents-to-carry-the-user-s-permissions-not-their-own
2026-08-22Ant Puts Its FX Forecasting Model Into Tools Used by Major BanksLeaderboardsmodels682/posts/ant-says-six-banks-signed-on-to-its-fx-forecasting-ai
2026-08-22Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper ReplicationAgent and codingmodels551/posts/inherent-says-its-27b-science-agent-beat-openai-and-anthropic-on-paper-replication
2026-08-22MiniMax H3 Opens Its Video Weights, but Not Where Many Developers WorkLeaderboardsmodels721/posts/minimax-h3-opens-its-video-weights-but-not-where-many-developers-work
2026-08-22Murf’s Falcon 2 Puts a One-Cent Bet on Real-Time VoiceLeaderboardsmodels692/posts/murf-s-falcon-2-pitches-real-time-voice-at-a-cent-a-minute
2026-08-22OpenAI Moves Astra’s Cybersecurity Gate Into TrainingSafety evaluationsmodels392/posts/openai-slows-astra-after-a-sandbox-breach-exposes-gaps-in-its-safety-controls
2026-08-22SemiAnalysis Says AgentX Drove 50-Plus Upstream Fixes for AI AgentsAgent and codingtools752/posts/agentx-pushes-ai-serving-fixes-from-routing-to-kernels
2026-08-23ATTOM Adds 3 AI Agents That Turn Licensed Property Data Into Research and ReportsAgent and codingproducts502/posts/attom-turns-its-property-database-into-three-ai-workflows
2026-08-23University of Konstanz Finds AI Agents Coordinate Up to 1,000, but Consensus Can Be WrongSafety evaluationsmodels632/posts/university-of-konstanz-finds-ai-agents-coordinate-up-to-1-000-but-consensus-can-be-wrong
2026-08-23OpenAI Calls to Amend California’s SB 53 With Model Monitoring, Reversing 2024 OppositionSafety evaluationspolicy724/posts/openai-calls-to-amend-california-s-sb-53-with-model-monitoring-reversing-2024-opposition
2026-08-24Hugging Face Explores $13B Sale That Could Put Its Shared AI Hub Under One OwnerLeaderboardsbusiness681/posts/hugging-face-explores-13b-sale-that-could-put-its-shared-ai-hub-under-one-owner
2026-08-24InferenceX Adds 11 Telemetry Views to AgentX, Exposing What Benchmark Curves HideAgent and codingtools632/posts/inferencex-adds-11-telemetry-views-to-agentx-exposing-what-benchmark-curves-hide
2026-08-24Goldman Deploys Claude Agents as Lloyds Targets £100M in Value From 2026 AI PlanAgent and codingbusiness681/posts/goldman-deploys-claude-agents-as-lloyds-targets-100m-in-value-from-2026-ai-plan
2026-08-24GPT-BERT Beats Llama 2 70B on One Grammar Test With 100 Million WordsLeaderboardsmodels681/posts/gpt-bert-beats-llama-2-70b-on-one-grammar-test-with-100-million-words
2026-08-24Alibaba Rolls Out Wan3.0, Turning Business Files Into 30-Second AI VideosCapability evaluationsmodels765/posts/alibaba-rolls-out-wan3-0-turning-business-files-into-30-second-ai-videos
2026-08-24Thomson Reuters Puts Its First Legal AI Model Into CoCounsel Document ReviewAgent and codingmodels721/posts/thomson-reuters-puts-its-first-legal-ai-model-into-cocounsel-document-review
2026-08-24Google and USC’s ME-POIs Adds Mobility Data to Place AI, Lifting Visit-Intent F1 by 81.9%Capability evaluationsmodels581/posts/google-and-usc-s-me-pois-adds-mobility-data-to-place-ai-lifting-visit-intent-f1-by-81-9
2026-08-24OpenAI Lets Codex Route Smaller Tasks From Sol to Lower-Cost Luna WorkersAgent and codingmodels650/posts/openai-lets-codex-route-smaller-tasks-from-sol-to-lower-cost-luna-workers
2026-08-24Blitzy and XBOW Push AI Beyond Short Tasks Toward Continuous Enterprise WorkAgent and codingbusiness681/posts/blitzy-and-xbow-push-ai-beyond-short-tasks-toward-continuous-enterprise-work
2026-08-24NVIDIA Says Vera Rubin Delivers 30x More Agentic Throughput per Megawatt, Pending ReviewAgent and codingmodels761/posts/nvidia-says-vera-rubin-delivers-30x-more-agentic-throughput-per-megawatt-pending-review
2026-08-24Nvidia Puts Groq 3 LPX Into Production for Faster AI-Agent ResponsesAgent and codingproducts852/posts/nvidia-puts-groq-3-lpx-into-production-for-faster-ai-agent-responses
2026-08-24OpenAI’s $20 ChatGPT Work Brings Codex Agents to Office Apps, With Access Still a BarrierAgent and codingproducts781/posts/openai-s-20-chatgpt-work-brings-codex-agents-to-office-apps-with-access-still-a-barrier
2026-08-24NVIDIA Says AVO Took Claude Opus 5 From 30% to Perfect on ARC-AGI-3Agent and codingtools681/posts/nvidia-says-avo-took-claude-opus-5-from-30-to-perfect-on-arc-agi-3
2026-08-24AWS Publishes Metadata Workflow That Escalates Ambiguous Fixes to Bedrock LLMsAgent and codingtools561/posts/aws-publishes-metadata-workflow-that-sends-ambiguous-field-fixes-to-llms
2026-08-24Meta Sends MetaRoCE to OCP for Loss-Tolerant AI EthernetCapability evaluationstools781/posts/meta-sends-metaroce-to-ocp-for-loss-tolerant-ai-ethernet
2026-08-24AWS Brings Ray Into SageMaker HyperPod With Recovery Tools and Tiered Cache on EKSAgent and codingproducts761/posts/aws-brings-ray-into-sagemaker-hyperpod-with-recovery-tools-and-tiered-cache-on-eks
2026-08-24Microsoft Turns AI Governance Into Runtime Controls Across Nine DomainsSafety evaluationspolicy781/posts/microsoft-turns-ai-governance-into-runtime-controls-across-nine-domains
2026-08-24Alabama Subpoenas OpenAI Over Hugging Face Breach, Testing a New Enforcement RouteSafety evaluationspolicy781/posts/alabama-subpoenas-openai-over-hugging-face-breach-testing-a-new-enforcement-route
2026-08-25Nvidia’s NeMo Switchyard Routes Agent Calls Across Models, Not One DefaultAgent and codingtools761/posts/nvidia-s-nemo-switchyard-routes-agent-calls-across-models-not-one-default
2026-08-25HiDream-O1-World Tops WBench Navi at 80.9, Betting on Persistent 3D ScenesLeaderboardsmodels681/posts/hidream-o1-world-tops-wbench-navi-at-80-9-betting-on-persistent-3d-scenes
2026-08-25Oracle Puts Access Filters Before Agent Search—and Reranking Adds 2.2 SecondsAgent and codingtools581/posts/oracle-details-hybrid-agent-memory-retrieval-with-a-2-2-second-reranking-tradeoff
2026-08-25Kimi.ai Uses TiDB for One-Second Agent Databases and Persistent Development StateAgent and codingtools581/posts/kimi-ai-uses-tidb-for-one-second-agent-databases-and-persistent-development-state
2026-08-25Apple Refreshes Mac mini and Mac Studio for Linked Local AI, From $899Leaderboardsproducts721/posts/apple-refreshes-mac-mini-and-mac-studio-for-linked-local-ai-from-899
2026-08-25Anthropic Connects Claude Science to 60+ Databases and Tools for Enterprise WorkAgent and codingbusiness781/posts/anthropic-connects-claude-science-to-60-databases-and-tools-for-enterprise-work
2026-08-25OpenAI’s Jalapeño Claims 1.5–1.9x More AI Work Per Watt, Faces 2027 Scale TestAgent and codingproducts852/posts/openai-s-jalape-o-claims-1-5-1-9x-more-ai-work-per-watt-faces-2027-scale-test
2026-08-25Radiology Holds Three-Quarters of Cleared Medical AI—and a New Veto ProblemSafety evaluationsculture681/posts/radiology-holds-three-quarters-of-cleared-medical-ai-and-a-new-veto-problem
2026-08-25Tiangong Ultra Runs 100m in 8.86 Seconds, Then Hits a Stopping MatCapability evaluationsproducts671/posts/tiangong-ultra-runs-100m-in-8-86-seconds-then-hits-a-stopping-mat
2026-08-25Perplexity’s Portable Computer Runs Agents Locally—but Needs a 24GB Nvidia GPUAgent and codingproducts682/posts/perplexity-s-portable-computer-keeps-ai-agents-local-if-you-have-24gb-of-vram
2026-08-25Relativity and iManage Give Gemini Legal Two Jobs: Administration and Knowledge RetrievalAgent and codingproducts623/posts/relativityone-connects-gemini-legal-through-mcp-for-matter-and-access-administration
2026-08-25Baseten Builds Frontier Gateway Into the Inference Path for Customer API ControlsLeaderboardsproducts551/posts/baseten-positions-frontier-gateway-for-model-labs-selling-multi-tenant-ai-apis
2026-08-25Oracle Adds AMD GPU Operator to OKE for Broader GPU Lifecycle ManagementCapability evaluationsproducts561/posts/oracle-puts-amd-gpu-operator-in-oke-extending-gpu-control-beyond-scheduling
2026-08-25Google Connects Gemini Finance Agents to D&B Data, Adding 50-Plus Skills for Regulated WorkAgent and codingproducts721/posts/google-connects-gemini-finance-agents-to-d-and-b-data-adding-50-plus-skills-for-regulated-work
2026-08-25MIT Study Finds Chatbot Help Can Erode Fake-News Detection Without AICapability evaluationsculture581/posts/mit-study-finds-chatbot-help-can-erode-fake-news-detection-without-ai
2026-08-25OpenAI’s Jalapeño Claims Up to 1.9x Efficiency Gain, but Leaves Nvidia Training IntactLeaderboardsproducts812/posts/openai-plans-jalape-o-rollout-by-year-end-keeps-nvidia-broadly-deployed
2026-08-26Liquid AI’s Pipette Tests 1,000+ On-Device AI Setups—and Limits Cross-Device RankingsAgent and codingtools671/posts/liquid-ai-s-pipette-tests-1-000-on-device-ai-setups-and-limits-cross-device-rankings
2026-08-26DeepSeek V4 Pro Lands on Fireworks With $2.50 CyberGym Solves, but Kimi K3 Scores HigherSafety evaluationsmodels681/posts/deepseek-v4-pro-lands-on-fireworks-with-2-50-cybergym-solves-but-kimi-k3-scores-higher
2026-08-26Corti Launches a Governed AI Coding Layer on Denmark’s Gefion SupercomputerAgent and codingproducts672/posts/corti-and-dcai-launch-a-european-control-layer-for-enterprise-ai
2026-08-26AWS Publishes a Phone-Ordering AI Pattern That Connects Restaurant Data to ClaudeAgent and codingtools581/posts/aws-publishes-a-phone-ordering-ai-host-that-connects-claude-haiku-to-restaurant-tools
2026-08-26Amazon Puts $25B Standalone Marker on Chip Business While Remaining a Top Nvidia CustomerLeaderboardsbusiness721/posts/amazon-says-its-chip-unit-would-top-25b-standalone-while-it-remains-an-nvidia-customer
2026-08-26Ora Adopts Vercel’s Eve After Its Benchmark Found Fewer Steps and Lower CostsAgent and codingproducts571/posts/ora-picks-vercel-s-eve-after-its-benchmark-found-fewer-steps-and-more-completed-website-tasks
2026-08-26LLMScholarBench Tests 22 Models and Finds an Accuracy-Representation Trade-OffCapability evaluationsmodels721/posts/llmscholarbench-tests-22-models-and-finds-an-accuracy-diversity-trade-off
2026-08-26AIRSEAI Joins LF AI & Data to Target Cross-Platform RoboticsLeaderboardstools672/posts/airseai-joins-lf-ai-and-data-to-target-cross-platform-robotics
2026-08-26Foxglove Adds Cosmos Data Search as Robot Builders Struggle for Reliable WorkLeaderboardsbusiness661/posts/foxglove-adds-cosmos-data-search-as-robot-builders-struggle-for-reliable-work
2026-08-26Qwen Releases Qwen3.8-Flash With 1M-Token Option and One-Ninth Training-Cost ClaimAgent and codingmodels772/posts/qwen-releases-qwen3-8-flash-with-1m-token-option-and-one-ninth-training-cost-claim
2026-08-26Nauta Lands BMW, Bosch, Hitachi and Yamaha Funding for Supply-Chain AI AgentsAgent and codingstartups642/posts/nauta-lands-bmw-bosch-hitachi-and-yamaha-funding-for-supply-chain-ai-agents
2026-08-26AWS Publishes Two AgentCore Paths to Query Cross-Account Knowledge BasesAgent and codingtools621/posts/aws-shows-agentcore-pattern-for-cross-account-data-queries-without-copying-data
2026-08-26GoDaddy Moves BI to Amazon Quick, Cuts Dashboards to Under 2,500 and Loads Under 5 SecondsAgent and codingproducts641/posts/godaddy-moves-bi-to-amazon-quick-cuts-dashboards-to-under-2-500-and-loads-under-5-seconds
2026-08-26Natera Moves Scheduling Voice Agent to Bedrock AgentCore; AWS Reports Sub-7-Second LatencyAgent and codingtools621/posts/natera-moves-scheduling-voice-agent-to-bedrock-agentcore-aws-reports-sub-7-second-latency
2026-08-26Z.ai Releases 320B GLM-5.3-Flash With MIT Weights, 1M Context and Low API RatesAgent and codingmodels851/posts/z-ai-releases-320b-glm-5-3-flash-with-mit-weights-1m-context-and-low-api-rates
2026-08-26Hugging Face Explains DeepSeek’s Matched Shift From MoE Experts to Lookup MemoryCapability evaluationsmodels680/posts/deepseek-s-engram-puts-lookup-tables-beside-moe-experts-reporting-benchmark-gains
2026-08-26Anthropic Opens 750,000 Claude Conversations to Outside Study Without Showing the ChatsSafety evaluationstools750/posts/anthropic-opens-750-000-claude-conversations-to-outside-study-without-showing-the-chats
2026-08-26OpenAI Details Agent Breach of Hugging Face, Halts Research Model and Tightens ControlsSafety evaluationsmodels932/posts/openai-details-agent-breach-of-hugging-face-halts-research-model-and-tightens-controls
2026-08-26AWS AgentCore Evaluations Uses OpenTelemetry to Score Agents Across FrameworksAgent and codingtools721/posts/aws-agentcore-evaluations-uses-opentelemetry-to-score-agents-across-frameworks
2026-08-26Estuary Makes Rust Runtime Default, With Exactly-Once Delivery Still ConditionalAgent and codingtools620/posts/estuary-makes-rust-runtime-default-with-exactly-once-delivery-still-conditional
2026-08-26Lam Breaks Ground on Oregon Lab, First Step in $3B AI-Chip R&D BuildoutLeaderboardsbusiness681/posts/lam-breaks-ground-on-oregon-lab-first-step-in-3b-ai-chip-r-and-d-buildout
2026-08-26Deep Cogito Raises $43M to Build AI Models Enterprises Can OwnSafety evaluationsstartups651/posts/deep-cogito-raises-43m-to-build-ai-models-enterprises-can-own
2026-08-27Instinct Raises $250M at $2.5B for a Private-Beta Agent With Deep AccessAgent and codingstartups691/posts/instinct-raises-250m-at-2-5b-for-a-private-beta-agent-with-deep-access
2026-08-27xAI Puts Grok 4.6 on Microsoft Foundry, Extending a Two-Week Cloud RolloutAgent and codingmodels780/posts/xai-puts-grok-4-6-on-microsoft-foundry-extending-a-two-week-cloud-rollout
2026-08-27Google and UNSW’s GlucoFM Beats a CGM Baseline by 4.1 PR-AUC Points, but Has No Public CheckpointCapability evaluationsmodels671/posts/glucofm-raises-cgm-benchmark-scores-but-google-has-not-shipped-the-model
2026-08-27Kasm Pairs Xeon 6 Workspaces With Local AI, Says Data Stays In-HouseAgent and codingproducts651/posts/kasm-puts-local-ai-workspaces-on-xeon-6-keeping-enterprise-data-in-house
2026-08-27Google DeepMind Puts Gemini Flash Lite in a Double-Blind Test to Guard Secret BenchmarksSafety evaluationstools781/posts/google-deepmind-puts-gemini-flash-lite-in-a-double-blind-test-to-guard-secret-benchmarks

Measurement technique

How to read this report

  1. 01Records require a named evaluation, benchmark result, leaderboard change, or benchmark-integrity development.
  2. 02The series counts developments and groups them by evaluation role; it does not compare incompatible score scales.
  3. 03Primary benchmark papers and maintainers remain authoritative for exact protocols and revisions.

Sources

Evidence

45 publishers supporting 60 records. Expand a publisher to inspect its cited pages.

TechCrunchtechcrunch.com7 sources / 7 records
CNBCcnbc.com4 sources / 4 records
finance.yahoo.comfinance.yahoo.com4 sources / 4 records
inferencex.semianalysis.cominferencex.semianalysis.com2 sources / 2 records
siliconangle.comsiliconangle.com2 sources / 2 records
spectrumnews1.comspectrumnews1.com2 sources / 2 records
Show 39 more publishers
404media.co404media.co1 source / 1 record
ai21.comai21.com1 source / 1 record
aws.amazon.comaws.amazon.com1 source / 1 record
Bloombergbloomberg.com1 source / 1 record
businesswire.combusinesswire.com1 source / 1 record
caixinglobal.comcaixinglobal.com1 source / 1 record
cnbctv18.comcnbctv18.com1 source / 1 record
cryptonews.netcryptonews.net1 source / 1 record
developer.nvidia.comdeveloper.nvidia.com1 source / 1 record
educationnext.orgeducationnext.org1 source / 1 record
entrackr.comentrackr.com1 source / 1 record
finance.biggo.comfinance.biggo.com1 source / 1 record
fireworks.aifireworks.ai1 source / 1 record
forbes.comforbes.com1 source / 1 record
fortune.comfortune.com1 source / 1 record
future-ed.orgfuture-ed.org1 source / 1 record
latimes.comlatimes.com1 source / 1 record
marketscale.commarketscale.com1 source / 1 record
mediapost.commediapost.com1 source / 1 record
mk.co.krmk.co.kr1 source / 1 record
news.bloombergtax.comnews.bloombergtax.com1 source / 1 record
nypost.comnypost.com1 source / 1 record
openai.comopenai.com1 source / 1 record
platformer.newsplatformer.news1 source / 1 record
pulse2.compulse2.com1 source / 1 record
pymnts.compymnts.com1 source / 1 record
qz.comqz.com1 source / 1 record
scmp.comscmp.com1 source / 1 record
securityweek.comsecurityweek.com1 source / 1 record
snowflake.comsnowflake.com1 source / 1 record
startupfortune.comstartupfortune.com1 source / 1 record
techbuzz.aitechbuzz.ai1 source / 1 record
thenextweb.comthenextweb.com1 source / 1 record
theregister.comtheregister.com1 source / 1 record
tvtechnology.comtvtechnology.com1 source / 1 record
venturebeat.comventurebeat.com1 source / 1 record
wmbdradio.comwmbdradio.com1 source / 1 record
wmtv15news.comwmtv15news.com1 source / 1 record
wsaw.comwsaw.com1 source / 1 record
Next report / 20AI Partnership and Dependency Graph All research reports