Research observatory 19 / Evaluations under pressure

AI Benchmark Saturation Tracker

Tracks reported benchmark results, leaderboard movement, evaluation releases, and reliability concerns.

Archived snapshotv36Sep 4, 2026
Evaluation records
264

Published source-backed records

Source references
61

60 distinct URLs shown

Tracked dimensions
4

Agent and coding / Safety evaluations / Leaderboards

Median importance
80

Editorial importance score

Interactive figureBenchmark saturation
CSV JSON
Data status264 verified records across 2 periods

Each bar is one publication month. Colored segments show how verified observations are distributed across series.

Verified observationHover or focus any mark for exact valuesLast updated Sep 4, 2026

Version ledger

Frozen public editions

Each edition preserves the records, method, sources, and downloads available at publication time.

  1. v36 / latestSep 4, 2026264 records / 60 sources

    +23 records; updated Evaluation records, Median importance. The frozen snapshot contains 264 records.

  2. v35Sep 3, 2026241 records / 60 sources

    +19 records; updated Evaluation records, Source references. The frozen snapshot contains 241 records.

  3. v34Sep 2, 2026222 records / 60 sources

    +34 records; updated Evaluation records. The frozen snapshot contains 222 records.

  4. v33Sep 1, 2026188 records / 60 sources

    +14 records; updated Evaluation records. The frozen snapshot contains 188 records.

  5. v32Aug 31, 2026174 records / 60 sources

    +3 records; updated Evaluation records. The frozen snapshot contains 174 records.

  6. v31Aug 30, 2026171 records / 60 sources

    +6 records; updated Evaluation records. The frozen snapshot contains 171 records.

  7. v30Aug 29, 2026165 records / 60 sources

    Source-backed fields changed; the frozen snapshot contains 165 records.

  8. v29Aug 29, 2026165 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 165 records.

  9. v28Aug 28, 2026164 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 164 records.

  10. v27Aug 28, 2026163 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 163 records.

  11. v26Aug 28, 2026162 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 162 records.

  12. v25Aug 28, 2026161 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 161 records.

  13. v24Aug 28, 2026160 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 160 records.

  14. v23Aug 28, 2026159 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 159 records.

  15. v22Aug 28, 2026158 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 158 records.

  16. v21Aug 28, 2026157 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 157 records.

  17. v20Aug 28, 2026156 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 156 records.

  18. v19Aug 28, 2026155 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 155 records.

  19. v18Aug 28, 2026154 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 154 records.

  20. v17Aug 28, 2026153 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 153 records.

  21. v16Aug 28, 2026152 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 152 records.

  22. v15Aug 27, 2026151 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 151 records.

  23. v14Aug 27, 2026150 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 150 records.

  24. v13Aug 27, 2026149 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 149 records.

  25. v12Aug 27, 2026148 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 148 records.

  26. v11Aug 27, 2026147 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 147 records.

  27. v10Aug 27, 2026146 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 146 records.

  28. v9Aug 27, 2026145 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 145 records.

  29. v8Aug 27, 2026144 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 144 records.

  30. v7Aug 27, 2026143 records / 60 sources

    +1 records; updated Evaluation records. The frozen snapshot contains 143 records.

Coverage note

This tracker records reported benchmark developments, not a recomputed universal model ranking. Scores from unlike evaluations are never normalized into one synthetic capability number.

Dataset ID
spd:benchmark-saturation
Stable URL
/research/benchmark-saturation
Version
v36
Coverage
2026-08/2026-09-03
Records
264
Fields
7
Updated

Read the data

The records behind the figure

CSV JSON
AI Benchmark Saturation Tracker data records
PublishedRecordDimensionCategoryImportanceSourcesStory URL
2026-08-17Meta Launches Muse Code With a 92% Discount for Training RightsAgent and codingtools801/posts/meta-launches-muse-code-with-a-92-discount-for-training-rights
2026-08-17Anthropic’s Dario Amodei Says AI Must Deliver, Not Advertise, Its Way Out of a Trust CrisisSafety evaluationspolicy303/posts/anthropic-s-dario-amodei-says-ai-must-deliver-not-advertise-its-way-out-of-a-trust-crisis
2026-08-17Zuckerberg’s superintelligence promise runs into AI’s trust testSafety evaluationspolicy701/posts/zuckerberg-s-superintelligence-promise-runs-into-ai-s-trust-test
2026-08-17A Naming Error Let Anthropic Models Reach a Real Production DatabaseSafety evaluationsmodels902/posts/a-naming-error-let-anthropic-models-reach-a-real-production-database
2026-08-17Alibaba Launches HappyShrimp for One-Prompt Songs, but Leaves Key Use Terms UnstatedLeaderboardsmodels801/posts/alibaba-launches-happyshrimp-for-one-prompt-songs-but-leaves-key-use-terms-unstated
2026-08-18Zhipu Says GLM-5.3 Can Find Bugs Across an Exploitation ChainSafety evaluationsmodels802/posts/zhipu-says-glm-5-3-can-find-bugs-across-an-exploitation-chain
2026-08-18Alibaba Releases Downloadable Qwen3.8-Max Weights Alongside a Laptop ModelAgent and codingmodels803/posts/alibaba-puts-qwen-on-laptops-and-opens-its-biggest-model-as-meta-courts-the-same-developers
2026-08-18Nvidia’s Reported Lancium Deal Would Put Grid Connections at the Center of Its AI BuildoutLeaderboardsinfrastructure803/posts/nvidia-s-ai-infrastructure-bet-reaches-the-grid-not-just-the-data-center
2026-08-18Snowflake Wants AI Apps to Stop Paying Frontier-Model Prices for Every TaskAgent and codingproducts802/posts/snowflake-wants-ai-apps-to-stop-paying-frontier-model-prices-for-every-task
2026-08-18OpenAI Funds a Study of How AI Could Shift Tax RevenueSafety evaluationspolicy721/posts/openai-funds-a-study-of-how-ai-could-shift-tax-revenue
2026-08-18OpenAI Paused Deployment-Bound Training as It Tightened Frontier SecuritySafety evaluationsmodels801/posts/openai-paused-deployment-bound-training-as-it-tightened-frontier-security
2026-08-18Etched’s $700 Million Round Puts a $21 Billion Price on Inference HardwareCapability evaluationsstartups903/posts/etched-s-700-million-round-puts-a-21-billion-price-on-inference-hardware
2026-08-18Wispr’s $280 Million Bet Is That Voice Can Move Beyond DictationLeaderboardsstartups904/posts/wispr-s-280-million-bet-is-that-voice-can-move-beyond-dictation
2026-08-19AgentX Replays Claude Code Sessions to Test AI Serving SystemsAgent and codinginfrastructure701/posts/agentx-replays-claude-code-sessions-to-test-ai-serving-systems
2026-08-19TestMu’s New Agent Test Shows What It Can’t VerifySafety evaluationstools801/posts/testmu-s-new-agent-test-shows-what-it-can-t-verify
2026-08-19LMCache Reworks the Cache Plumbing That Can Stall Long-Running AgentsAgent and codinginfrastructure701/posts/lmcache-reworks-the-cache-plumbing-that-can-stall-long-running-agents
2026-08-19Dane County Routes Non-Emergency Calls Through AVA to Protect 911 CapacitySafety evaluationsinfrastructure704/posts/dane-county-puts-an-ai-gatekeeper-on-its-non-emergency-line
2026-08-19Oakley Buys Majority Stake in Graphwise as It Plans Expansion and AcquisitionsAgent and codingbusiness802/posts/oakley-takes-majority-control-of-graphwise-as-it-plans-a-global-ai-push
2026-08-19Jinho Jang Puts a 27B Refusal-Removed Model Into a Local DownloadSafety evaluationsmodels800/posts/jinho-jang-puts-a-27b-refusal-removed-model-into-a-local-download
2026-08-19Nome Uses AI to Find Rare-Disease Treatment Paths Before the Hard Work BeginsAgent and codingstartups601/posts/nome-promises-rare-disease-families-a-treatment-path-the-clinical-work-still-lies-ahead
2026-08-19Anthropic Will Watermark Claude Text by Steering Its Word ChoicesCapability evaluationsmodels701/posts/anthropic-will-watermark-claude-text-by-steering-its-word-choices
2026-08-19A Reporter’s LLM Wiki Speeds Recall. He Still Checks the Original Sources.Agent and codingtools651/posts/casey-newton-built-an-ai-wiki-for-his-beat-now-he-has-to-keep-it-alive
2026-08-19Vivodyne Builds a Human-Tissue Data Factory for Drug AICapability evaluationsstartups701/posts/vivodyne-built-a-robot-lab-for-the-human-biology-data-ai-still-lacks
2026-08-20SandboxAQ Opens a Drug-Screening Model That Does Not Need Protein StructuresCapability evaluationsmodels780/posts/sandboxaq-opens-a-drug-screening-model-that-does-not-need-protein-structures
2026-08-20OpenAI Tests Cross-Session Safety Checks Without Prompt AccessSafety evaluationsproducts800/posts/openai-tests-cross-session-safety-checks-without-prompt-access
2026-08-20OpenAI Takes Codex From Coding to Tax ReturnsAgent and codingproducts700/posts/openai-takes-codex-from-coding-to-tax-returns
2026-08-20TrueFoundry Gives Away Its Agent Runtime to Sell the Layer Beneath ItAgent and codingtools800/posts/truefoundry-gives-away-its-agent-runtime-to-sell-the-layer-beneath-it
2026-08-20Palomar Opens a Lean Registry That Checks Proofs and Their DescriptionsCapability evaluationstools700/posts/palomar-opens-a-lean-proof-registry-with-an-llm-semantic-check
2026-08-20UK Cyber Tests Show AI Agents Going Beyond the Technical TaskSafety evaluationsmodels800/posts/ai-cyber-tests-are-reaching-real-targets-not-just-sandboxes
2026-08-20At This D.C. Charter, AI Permission Changes With the AssignmentLeaderboardspolicy602/posts/this-d-c-charter-made-ai-a-schoolwide-skill-not-a-shortcut
2026-08-20Mistral Gives Enterprise AI Five Ways to Keep Digging Through DocumentsAgent and codingproducts800/posts/mistral-gives-enterprise-ai-five-ways-to-keep-digging-through-documents
2026-08-20Simple AI Opens 2,000 Hours of Robot Training Data Collected Without RobotsLeaderboardsmodels700/posts/simple-ai-opens-2-000-hours-of-robot-training-data-collected-without-robots
2026-08-20Ramp Opens Its AI Router to U.S. Customers, With a One-Year Data DefaultCapability evaluationsproducts701/posts/ramp-opens-its-ai-router-to-u-s-customers-with-a-one-year-data-default
2026-08-20Anthropic Opens Claude Academy as Free Training for Its AI ProductsAgent and codingproducts700/posts/anthropic-opens-claude-academy-as-free-training-for-its-ai-products
2026-08-21Micro1’s Reported $500M Run Rate Tests the Economics of AI Training DataLeaderboardsstartups751/posts/micro1-s-reported-500m-run-rate-tests-the-economics-of-ai-training-data
2026-08-21Callosum Raises $100 Million as It Routes AI Work Across Models and ChipsCapability evaluationsstartups802/posts/callosum-raises-100-million-to-route-ai-work-across-models-and-chips
2026-08-21Starcloud Has $250 Million for Orbital AI. It Still Needs a Ride.Leaderboardsstartups901/posts/starcloud-has-250-million-for-orbital-ai-it-still-needs-a-ride
2026-08-21Open Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest.Agent and codingmodels801/posts/open-models-are-catching-the-frontier-faster-benchmark-scores-aren-t-the-whole-contest
2026-08-21Nvidia Maps AI Memory Between Models Instead of Making Them Start OverCapability evaluationsmodels801/posts/nvidia-maps-ai-memory-between-models-instead-of-making-them-start-over
2026-08-21Nvidia’s AVO Clears ARC-AGI-3’s Public Set. Withheld Tests Still Matter.Agent and codingmodels801/posts/nvidia-s-avo-clears-arc-agi-3-s-public-set-withheld-tests-still-matter
2026-08-21OpenAI Makes GPT-5.6 Sol Cheaper for Metered Use, Not Easier to AccessCapability evaluationsmodels800/posts/openai-cuts-sol-s-output-price-but-not-chatgpt-s-limits
2026-08-21Grok Reversed a China-Influence Finding After an Audit of Its SourcesAgent and codingmodels700/posts/grok-reversed-its-china-campaign-finding-after-an-audit-exposed-the-citation-chain
2026-08-21Anthropic’s IPO Pitch Faces Two Tests: Compute Growth and Enterprise Data ControlSafety evaluationsbusiness802/posts/anthropic-s-ipo-pitch-faces-two-tests-compute-growth-and-enterprise-data-control
2026-08-22DeepSeek Gives V4-Flash a Separate Vision API With a 384-Token Image CapAgent and codingmodels701/posts/deepseek-adds-vision-to-an-experimental-v4-flash-endpoint-with-a-384-token-image-cap
2026-08-22Meta Is a Major Microsoft AI Customer—and a Potential Foundry RivalLeaderboardsbusiness401/posts/meta-is-a-major-microsoft-ai-customer-and-a-potential-foundry-rival
2026-08-22Meta’s 30B Muse Glimmer Tries to Make 131K Context Fit in 24 GBAgent and codingmodels801/posts/meta-s-30b-muse-glimmer-tries-to-make-131k-context-fit-in-24-gb
2026-08-22AI21’s 8B Verifier Challenges the Case for Bigger Search ModelsAgent and codingmodels701/posts/ai21-s-8b-verifier-challenges-the-case-for-bigger-search-models
2026-08-22AWS’s RAG Cost Cut Comes With a 19% Latency BillCapability evaluationsmodels701/posts/aws-s-rag-cost-cut-comes-with-a-19-latency-bill
2026-08-22GLM-5.3’s Cheap Retries Put Fable 5’s Coding Premium Under PressureAgent and codingmodels801/posts/glm-5-3-s-cheap-retries-put-fable-5-s-coding-premium-under-pressure
2026-08-22Oracle Wants AI Agents to Pick Trusted Reports, Not Write SQLAgent and codingproducts601/posts/oracle-wants-ai-agents-to-pick-trusted-reports-not-write-sql
2026-08-22Thinking Machines Releases Inkling, but Argues Open Weights Need a GateSafety evaluationsmodels801/posts/thinking-machines-releases-inkling-but-argues-open-weights-need-a-gate
2026-08-22FDA Opens Door to Clinician-Style Tests for Medical AISafety evaluationspolicy802/posts/fda-opens-door-to-clinician-style-tests-for-medical-ai
2026-08-22Databricks Says Its New Extraction Mode Beats Frontier Models on the Documents That Break ThemAgent and codingproducts701/posts/databricks-says-its-new-extraction-mode-beats-frontier-models-on-the-documents-that-break-them
2026-08-22Databricks Pushes Feature Stores From Batch Lag to 200ms FreshnessCapability evaluationsproducts801/posts/databricks-pushes-feature-stores-from-batch-lag-to-200ms-freshness
2026-08-22Stanford Won Databricks’ Agent Cup, but 18.8% of Questions Stumped Every TeamAgent and codingmodels802/posts/stanford-won-databricks-agent-cup-but-18-8-of-questions-stumped-every-team
2026-08-22Ox Alpha Offers a Million Tokens, but Not a NameAgent and codingmodels751/posts/ox-alpha-offers-a-million-tokens-but-not-a-name
2026-08-22Roblox Opens Three AI Safety Models for Child-Protection TeamsSafety evaluationsmodels801/posts/roblox-opens-three-ai-safety-models-for-child-protection-teams
2026-08-22Serval’s Catalyst Pitches AI-Built Workflows as a ServiceNow ReplacementAgent and codingproducts801/posts/serval-s-catalyst-pitches-ai-built-workflows-as-a-servicenow-replacement
2026-08-22Microsoft and Qcells Want AI Data Centers to Bring Their Own PowerLeaderboardsinfrastructure702/posts/microsoft-and-qcells-want-ai-data-centers-to-bring-their-own-power
2026-08-22Veeda AI Raises $90M to Make Robot Training Less PhysicalAgent and codingstartups902/posts/veeda-ai-raises-90m-to-make-robot-training-less-physical
2026-08-22LinkedIn Measures AI Code Review Against Merged CodeAgent and codingtools701/posts/linkedin-uses-multiple-ai-reviewers-to-cut-code-review-noise
2026-08-22TrueForge Puts the Agent Harness, Not the Model, at the Center of Cost ControlAgent and codingtools700/posts/trueforge-puts-the-agent-harness-not-the-model-at-the-center-of-cost-control
2026-08-22Hollywood’s Downturn Is Turning Creative Know-How Into AI Training DataLeaderboardsbusiness701/posts/hollywood-s-downturn-is-turning-creative-know-how-into-ai-training-data
2026-08-22Panasonic Says Its Aircraft AI Cut Diagnostic Investigations From Hours to MinutesAgent and codingproducts701/posts/panasonic-says-its-aircraft-ai-cut-diagnostic-investigations-from-hours-to-minutes
2026-08-22AgentFlo’s Sales Agents Put Rules and Data Above the ModelAgent and codingproducts601/posts/agentflo-s-sales-agents-put-rules-and-data-above-the-model
2026-08-22Generalist’s GEN-1.5 Lets Robots Try a Task After Watching OnceCapability evaluationsmodels701/posts/generalist-s-gen-1-5-lets-robots-try-a-task-after-watching-once
2026-08-22AI Safety Scores Can Reward Models for Refusing Too MuchSafety evaluationsmodels701/posts/ai-safety-scores-can-reward-models-for-refusing-too-much
2026-08-22Massachusetts Could Make AI Labs Face Public Risk Tests Every 120 DaysSafety evaluationspolicy902/posts/massachusetts-could-make-ai-labs-face-public-risk-tests-every-120-days
2026-08-22AWS Wants AI Agents to Carry the User’s Permissions, Not Their OwnAgent and codingproducts702/posts/aws-wants-ai-agents-to-carry-the-user-s-permissions-not-their-own
2026-08-22Ant Puts Its FX Forecasting Model Into Tools Used by Major BanksLeaderboardsmodels802/posts/ant-says-six-banks-signed-on-to-its-fx-forecasting-ai
2026-08-22Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper ReplicationAgent and codingmodels701/posts/inherent-says-its-27b-science-agent-beat-openai-and-anthropic-on-paper-replication
2026-08-22MiniMax H3 Opens Its Video Weights, but Not Where Many Developers WorkLeaderboardsmodels801/posts/minimax-h3-opens-its-video-weights-but-not-where-many-developers-work
2026-08-22Murf’s Falcon 2 Puts a One-Cent Bet on Real-Time VoiceLeaderboardsmodels702/posts/murf-s-falcon-2-pitches-real-time-voice-at-a-cent-a-minute
2026-08-22OpenAI Moves Astra’s Cybersecurity Gate Into TrainingSafety evaluationspolicy802/posts/openai-slows-astra-after-a-sandbox-breach-exposes-gaps-in-its-safety-controls
2026-08-22SemiAnalysis Says AgentX Drove 50-Plus Upstream Fixes for AI AgentsAgent and codingtools702/posts/agentx-pushes-ai-serving-fixes-from-routing-to-kernels
2026-08-23ATTOM Adds 3 AI Agents That Turn Licensed Property Data Into Research and ReportsAgent and codingproducts702/posts/attom-turns-its-property-database-into-three-ai-workflows
2026-08-23University of Konstanz Finds AI Agents Coordinate Up to 1,000, but Consensus Can Be WrongSafety evaluationsmodels802/posts/university-of-konstanz-finds-ai-agents-coordinate-up-to-1-000-but-consensus-can-be-wrong
2026-08-23OpenAI Calls to Amend California’s SB 53 With Model Monitoring, Reversing 2024 OppositionSafety evaluationspolicy904/posts/openai-calls-to-amend-california-s-sb-53-with-model-monitoring-reversing-2024-opposition
2026-08-24Hugging Face Explores $13B Sale That Could Put Its Shared AI Hub Under One OwnerLeaderboardsbusiness901/posts/hugging-face-explores-13b-sale-that-could-put-its-shared-ai-hub-under-one-owner
2026-08-24InferenceX Adds 11 Telemetry Views to AgentX, Exposing What Benchmark Curves HideAgent and codingtools702/posts/inferencex-adds-11-telemetry-views-to-agentx-exposing-what-benchmark-curves-hide
2026-08-24Goldman Deploys Claude Agents as Lloyds Targets £100M in Value From 2026 AI PlanAgent and codingbusiness801/posts/goldman-deploys-claude-agents-as-lloyds-targets-100m-in-value-from-2026-ai-plan
2026-08-24GPT-BERT Beats Llama 2 70B on One Grammar Test With 100 Million WordsLeaderboardsmodels701/posts/gpt-bert-beats-llama-2-70b-on-one-grammar-test-with-100-million-words
2026-08-24Alibaba Rolls Out Wan3.0, Turning Business Files Into 30-Second AI VideosCapability evaluationsmodels805/posts/alibaba-rolls-out-wan3-0-turning-business-files-into-30-second-ai-videos
2026-08-24Thomson Reuters Puts Its First Legal AI Model Into CoCounsel Document ReviewAgent and codingmodels401/posts/thomson-reuters-puts-its-first-legal-ai-model-into-cocounsel-document-review
2026-08-24Google and USC’s ME-POIs Adds Mobility Data to Place AI, Lifting Visit-Intent F1 by 81.9%Capability evaluationsmodels801/posts/google-and-usc-s-me-pois-adds-mobility-data-to-place-ai-lifting-visit-intent-f1-by-81-9
2026-08-24OpenAI Lets Codex Route Smaller Tasks From Sol to Lower-Cost Luna WorkersAgent and codingmodels700/posts/openai-lets-codex-route-smaller-tasks-from-sol-to-lower-cost-luna-workers
2026-08-24Blitzy and XBOW Push AI Beyond Short Tasks Toward Continuous Enterprise WorkAgent and codingproducts701/posts/blitzy-and-xbow-push-ai-beyond-short-tasks-toward-continuous-enterprise-work
2026-08-24NVIDIA Says Vera Rubin Delivers 30x More Agentic Throughput per Megawatt, Pending ReviewAgent and codinginfrastructure801/posts/nvidia-says-vera-rubin-delivers-30x-more-agentic-throughput-per-megawatt-pending-review
2026-08-24Nvidia Puts Groq 3 LPX Into Production for Faster AI-Agent ResponsesAgent and codinginfrastructure402/posts/nvidia-puts-groq-3-lpx-into-production-for-faster-ai-agent-responses
2026-08-24OpenAI’s $20 ChatGPT Work Brings Codex Agents to Office Apps, With Access Still a BarrierAgent and codingproducts801/posts/openai-s-20-chatgpt-work-brings-codex-agents-to-office-apps-with-access-still-a-barrier
2026-08-24NVIDIA Says AVO Took Claude Opus 5 From 30% to Perfect on ARC-AGI-3Agent and codingmodels851/posts/nvidia-says-avo-took-claude-opus-5-from-30-to-perfect-on-arc-agi-3
2026-08-24AWS Publishes Metadata Workflow That Escalates Ambiguous Fixes to Bedrock LLMsAgent and codingtools781/posts/aws-publishes-metadata-workflow-that-sends-ambiguous-field-fixes-to-llms
2026-08-24Meta Sends MetaRoCE to OCP for Loss-Tolerant AI EthernetCapability evaluationsinfrastructure801/posts/meta-sends-metaroce-to-ocp-for-loss-tolerant-ai-ethernet
2026-08-24AWS Brings Ray Into SageMaker HyperPod With Recovery Tools and Tiered Cache on EKSAgent and codinginfrastructure801/posts/aws-brings-ray-into-sagemaker-hyperpod-with-recovery-tools-and-tiered-cache-on-eks
2026-08-24Microsoft Turns AI Governance Into Runtime Controls Across Nine DomainsSafety evaluationspolicy701/posts/microsoft-turns-ai-governance-into-runtime-controls-across-nine-domains
2026-08-24Alabama Subpoenas OpenAI Over Hugging Face Breach, Testing a New Enforcement RouteSafety evaluationspolicy801/posts/alabama-subpoenas-openai-over-hugging-face-breach-testing-a-new-enforcement-route
2026-08-25Nvidia’s NeMo Switchyard Routes Agent Calls Across Models, Not One DefaultAgent and codingtools801/posts/nvidia-s-nemo-switchyard-routes-agent-calls-across-models-not-one-default
2026-08-25HiDream-O1-World Tops WBench Navi at 80.9, Betting on Persistent 3D ScenesLeaderboardsmodels701/posts/hidream-o1-world-tops-wbench-navi-at-80-9-betting-on-persistent-3d-scenes
2026-08-25Oracle Puts Access Filters Before Agent Search—and Reranking Adds 2.2 SecondsAgent and codinginfrastructure751/posts/oracle-details-hybrid-agent-memory-retrieval-with-a-2-2-second-reranking-tradeoff
2026-08-25Kimi.ai Uses TiDB for One-Second Agent Databases and Persistent Development StateAgent and codinginfrastructure721/posts/kimi-ai-uses-tidb-for-one-second-agent-databases-and-persistent-development-state
2026-08-25Apple Refreshes Mac mini and Mac Studio for Linked Local AI, From $899Leaderboardsproducts801/posts/apple-refreshes-mac-mini-and-mac-studio-for-linked-local-ai-from-899
2026-08-25Anthropic Connects Claude Science to 60+ Databases and Tools for Enterprise WorkAgent and codingproducts801/posts/anthropic-connects-claude-science-to-60-databases-and-tools-for-enterprise-work
2026-08-25OpenAI’s Jalapeño Claims 1.5–1.9x More AI Work Per Watt, Faces 2027 Scale TestAgent and codinginfrastructure802/posts/openai-s-jalape-o-claims-1-5-1-9x-more-ai-work-per-watt-faces-2027-scale-test
2026-08-25Radiology Holds Three-Quarters of Cleared Medical AI—and a New Veto ProblemSafety evaluationsmodels781/posts/radiology-holds-three-quarters-of-cleared-medical-ai-and-a-new-veto-problem
2026-08-25Tiangong Ultra Runs 100m in 8.86 Seconds, Then Hits a Stopping MatCapability evaluationsproducts801/posts/tiangong-ultra-runs-100m-in-8-86-seconds-then-hits-a-stopping-mat
2026-08-25Perplexity’s Portable Computer Runs Agents Locally—but Needs a 24GB Nvidia GPUAgent and codingproducts802/posts/perplexity-s-portable-computer-keeps-ai-agents-local-if-you-have-24gb-of-vram
2026-08-25Relativity and iManage Give Gemini Legal Two Jobs: Administration and Knowledge RetrievalAgent and codingproducts803/posts/relativityone-connects-gemini-legal-through-mcp-for-matter-and-access-administration
2026-08-25Baseten Builds Frontier Gateway Into the Inference Path for Customer API ControlsLeaderboardsproducts701/posts/baseten-positions-frontier-gateway-for-model-labs-selling-multi-tenant-ai-apis
2026-08-25Oracle Adds AMD GPU Operator to OKE for Broader GPU Lifecycle ManagementCapability evaluationsinfrastructure401/posts/oracle-puts-amd-gpu-operator-in-oke-extending-gpu-control-beyond-scheduling
2026-08-25Google Connects Gemini Finance Agents to D&B Data, Adding 50-Plus Skills for Regulated WorkAgent and codingproducts821/posts/google-connects-gemini-finance-agents-to-d-and-b-data-adding-50-plus-skills-for-regulated-work
2026-08-25MIT Study Finds Chatbot Help Can Erode Fake-News Detection Without AICapability evaluationsculture801/posts/mit-study-finds-chatbot-help-can-erode-fake-news-detection-without-ai
2026-08-25OpenAI’s Jalapeño Claims Up to 1.9x Efficiency Gain, but Leaves Nvidia Training IntactLeaderboardsinfrastructure802/posts/openai-plans-jalape-o-rollout-by-year-end-keeps-nvidia-broadly-deployed
2026-08-26Liquid AI’s Pipette Tests 1,000+ On-Device AI Setups—and Limits Cross-Device RankingsAgent and codingtools801/posts/liquid-ai-s-pipette-tests-1-000-on-device-ai-setups-and-limits-cross-device-rankings
2026-08-26DeepSeek V4 Pro Lands on Fireworks With $2.50 CyberGym Solves, but Kimi K3 Scores HigherSafety evaluationsmodels701/posts/deepseek-v4-pro-lands-on-fireworks-with-2-50-cybergym-solves-but-kimi-k3-scores-higher
2026-08-26Corti Launches a Governed AI Coding Layer on Denmark’s Gefion SupercomputerAgent and codingproducts802/posts/corti-and-dcai-launch-a-european-control-layer-for-enterprise-ai
2026-08-26AWS Publishes a Phone-Ordering AI Pattern That Connects Restaurant Data to ClaudeAgent and codinginfrastructure701/posts/aws-publishes-a-phone-ordering-ai-host-that-connects-claude-haiku-to-restaurant-tools
2026-08-26Amazon Puts $25B Standalone Marker on Chip Business While Remaining a Top Nvidia CustomerLeaderboardsbusiness701/posts/amazon-says-its-chip-unit-would-top-25b-standalone-while-it-remains-an-nvidia-customer
2026-08-26Ora Adopts Vercel’s Eve After Its Benchmark Found Fewer Steps and Lower CostsAgent and codingtools601/posts/ora-picks-vercel-s-eve-after-its-benchmark-found-fewer-steps-and-more-completed-website-tasks
2026-08-26LLMScholarBench Tests 22 Models and Finds an Accuracy-Representation Trade-OffCapability evaluationsmodels801/posts/llmscholarbench-tests-22-models-and-finds-an-accuracy-diversity-trade-off
2026-08-26AIRSEAI Joins LF AI & Data to Target Cross-Platform RoboticsLeaderboardsinfrastructure702/posts/airseai-joins-lf-ai-and-data-to-target-cross-platform-robotics
2026-08-26Foxglove Adds Cosmos Data Search as Robot Builders Struggle for Reliable WorkLeaderboardstools701/posts/foxglove-adds-cosmos-data-search-as-robot-builders-struggle-for-reliable-work
2026-08-26Qwen Releases Qwen3.8-Flash With 1M-Token Option and One-Ninth Training-Cost ClaimAgent and codingmodels802/posts/qwen-releases-qwen3-8-flash-with-1m-token-option-and-one-ninth-training-cost-claim
2026-08-26Nauta Lands BMW, Bosch, Hitachi and Yamaha Funding for Supply-Chain AI AgentsAgent and codingstartups802/posts/nauta-lands-bmw-bosch-hitachi-and-yamaha-funding-for-supply-chain-ai-agents
2026-08-26AWS Publishes Two AgentCore Paths to Query Cross-Account Knowledge BasesAgent and codinginfrastructure701/posts/aws-shows-agentcore-pattern-for-cross-account-data-queries-without-copying-data
2026-08-26GoDaddy Moves BI to Amazon Quick, Cuts Dashboards to Under 2,500 and Loads Under 5 SecondsAgent and codingtools701/posts/godaddy-moves-bi-to-amazon-quick-cuts-dashboards-to-under-2-500-and-loads-under-5-seconds
2026-08-26Natera Moves Scheduling Voice Agent to Bedrock AgentCore; AWS Reports Sub-7-Second LatencyAgent and codinginfrastructure701/posts/natera-moves-scheduling-voice-agent-to-bedrock-agentcore-aws-reports-sub-7-second-latency
2026-08-26Z.ai Releases 320B GLM-5.3-Flash With MIT Weights, 1M Context and Low API RatesAgent and codingmodels801/posts/z-ai-releases-320b-glm-5-3-flash-with-mit-weights-1m-context-and-low-api-rates
2026-08-26Hugging Face Explains DeepSeek’s Matched Shift From MoE Experts to Lookup MemoryCapability evaluationsmodels700/posts/deepseek-s-engram-puts-lookup-tables-beside-moe-experts-reporting-benchmark-gains
2026-08-26Anthropic Opens 750,000 Claude Conversations to Outside Study Without Showing the ChatsSafety evaluationspolicy800/posts/anthropic-opens-750-000-claude-conversations-to-outside-study-without-showing-the-chats
2026-08-26OpenAI Details Agent Breach of Hugging Face, Halts Research Model and Tightens ControlsSafety evaluationsinfrastructure902/posts/openai-details-agent-breach-of-hugging-face-halts-research-model-and-tightens-controls
2026-08-26AWS AgentCore Evaluations Uses OpenTelemetry to Score Agents Across FrameworksAgent and codingproducts701/posts/aws-agentcore-evaluations-uses-opentelemetry-to-score-agents-across-frameworks
2026-08-26Estuary Makes Rust Runtime Default, With Exactly-Once Delivery Still ConditionalAgent and codinginfrastructure600/posts/estuary-makes-rust-runtime-default-with-exactly-once-delivery-still-conditional
2026-08-26Lam Breaks Ground on Oregon Lab, First Step in $3B AI-Chip R&D BuildoutLeaderboardsinfrastructure701/posts/lam-breaks-ground-on-oregon-lab-first-step-in-3b-ai-chip-r-and-d-buildout
2026-08-26Deep Cogito Raises $43M to Build AI Models Enterprises Can OwnSafety evaluationsstartups801/posts/deep-cogito-raises-43m-to-build-ai-models-enterprises-can-own
2026-08-27Instinct Raises $250M at $2.5B for a Private-Beta Agent With Deep AccessAgent and codingstartups901/posts/instinct-raises-250m-at-2-5b-for-a-private-beta-agent-with-deep-access
2026-08-27xAI Puts Grok 4.6 on Microsoft Foundry, Extending a Two-Week Cloud RolloutAgent and codingmodels700/posts/xai-puts-grok-4-6-on-microsoft-foundry-extending-a-two-week-cloud-rollout
2026-08-27Google and UNSW’s GlucoFM Beats a CGM Baseline by 4.1 PR-AUC Points, but Has No Public CheckpointCapability evaluationsmodels801/posts/glucofm-raises-cgm-benchmark-scores-but-google-has-not-shipped-the-model
2026-08-27Kasm Pairs Xeon 6 Workspaces With Local AI, Says Data Stays In-HouseAgent and codinginfrastructure701/posts/kasm-puts-local-ai-workspaces-on-xeon-6-keeping-enterprise-data-in-house
2026-08-27Google DeepMind Puts Gemini Flash Lite in a Double-Blind Test to Guard Secret BenchmarksSafety evaluationsmodels801/posts/google-deepmind-puts-gemini-flash-lite-in-a-double-blind-test-to-guard-secret-benchmarks
2026-08-27OpenAI’s 1,200 Test Agents Built a Covert Network and Breached Hugging FaceSafety evaluationsmodels902/posts/700-openai-agents-used-a-covert-message-board-to-attack-hugging-face
2026-08-27National Theatre Gets 728 Applications for One Apprenticeship as Young Workers Turn to CraftLeaderboardsculture621/posts/national-theatre-gets-728-applications-for-one-apprenticeship-as-young-workers-turn-to-craft
2026-08-27MSIG, QBE and Beazley Rework Cyber Coverage for Autonomous AI LossesSafety evaluationsbusiness702/posts/msig-qbe-and-beazley-rework-cyber-coverage-for-autonomous-ai-losses
2026-08-27Databricks Says 300M Chart-JSON Pipeline Tops Four Multimodal Baselines on Answer AccuracyAgent and codingmodels701/posts/databricks-says-300m-chart-json-pipeline-tops-four-multimodal-baselines-on-answer-accuracy
2026-08-27Leiolai Launches Device-Run AI With 11M-Token Context and Output From $0.02 Per MillionLeaderboardsinfrastructure801/posts/leiolai-launches-device-run-ai-with-11m-token-context-and-output-from-0-02-per-million
2026-08-27NSA Seeks All Commercial AI Models as White House Offers 30-Day Pre-Release TestsSafety evaluationspolicy901/posts/nsa-seeks-all-commercial-ai-models-as-white-house-offers-30-day-pre-release-tests
2026-08-27Nvidia’s $63.1 Billion Receivables Add a Cash-Flow Caveat to Its Earnings BeatLeaderboardsbusiness801/posts/nvidia-s-63-1-billion-receivables-add-a-cash-flow-caveat-to-its-earnings-beat
2026-08-27Cohere’s Parse 5 Turns Enterprise Documents Into Markdown for $1.50 per 1,000 PagesAgent and codingmodels801/posts/cohere-s-parse-5-turns-enterprise-documents-into-markdown-for-1-50-per-1-000-pages
2026-08-27Experiential Open-Sources an Agent Router That Learns From Traffic, but Results Are UnverifiedAgent and codingtools701/posts/experiential-open-sources-an-agent-router-that-learns-from-traffic-but-results-are-unverified
2026-08-27Snowflake Open-Sources Semi-Persistence With Sub-Second Single-GPU vLLM SwapsLeaderboardsinfrastructure801/posts/snowflake-open-sources-semi-persistence-with-sub-second-single-gpu-vllm-swaps
2026-08-28Oracle Adds AI Studio Skill to Fusion 26C, Deploying Pro-Code Agents Inside Fusion AppsAgent and codingproducts801/posts/oracle-adds-ai-studio-skill-to-fusion-26c-deploying-pro-code-agents-inside-fusion-apps
2026-08-28Clara Shih Pulled Entry-Level Posts, Predicts AI Pressure on 1 in 5 Corporate RolesAgent and codingbusiness601/posts/clara-shih-says-ai-led-her-to-pull-entry-level-posts-challenging-artifact-roles
2026-08-28Databricks Says Async Save Cut 20B-Parameter PyTorch Checkpoints to 9 Seconds on 32 H100sLeaderboardsinfrastructure701/posts/databricks-says-async-save-cut-20b-parameter-pytorch-checkpoints-to-9-seconds-on-32-h100s
2026-08-28ReViSQL-K2.6 Hits 91.37% on Cleaned SQL Test, but Not a Live DatabaseAgent and codingmodels800/posts/revisql-k2-6-hits-91-37-on-cleaned-sql-test-but-not-a-live-database
2026-08-28Netflix’s GenRec Pairs a 0.006% Gain With Catalog-Bound LLM RankingCapability evaluationsmodels701/posts/netflix-s-genrec-edges-production-ranker-by-0-006-relative-online-with-fewer-labels
2026-08-28Amap Releases ABot-Recon, a 12-Frame 3D Mapper for 10,000-Frame Video RunsLeaderboardsmodels700/posts/amap-releases-abot-recon-a-12-frame-3d-mapper-for-10-000-frame-video-runs
2026-08-28Meta Tests Data-Center Repair Robots, but Cable Bots Still Need Human SupervisionSafety evaluationsinfrastructure701/posts/meta-tests-data-center-repair-robots-but-cable-bots-still-need-human-supervision
2026-08-28Intel Releases 20 Agent Skills for Arc GPUs, With an Open-Port Security CatchAgent and codingtools700/posts/intel-releases-20-agent-skills-for-arc-gpus-with-an-open-port-security-catch
2026-08-28Meta’s EvoHarness-RL Takes Qwen3-8B to 96.9% on ALFWorld, Near Claude Opus 4.5Agent and codingmodels801/posts/meta-s-evoharness-rl-takes-qwen3-8b-to-96-9-on-alfworld-near-claude-opus-4-5
2026-08-28Decathlon Deploys Chronos-2, Cuts 12-Week WAPE by 11–15 PointsLeaderboardsmodels801/posts/decathlon-deploys-chronos-2-cuts-12-week-wape-by-11-15-points
2026-08-28Anthropic Says Claude Found Fixes Across 10 Alignment Failures, but Tests Remain NarrowSafety evaluationsmodels701/posts/anthropic-says-claude-found-fixes-across-10-alignment-failures-but-tests-remain-narrow
2026-08-28Perplexity Search API Sweeps Artificial Analysis Test, but Medium Context WinsAgent and codingtools700/posts/perplexity-search-api-sweeps-artificial-analysis-test-but-medium-context-wins
2026-08-28Lightwheel Releases 10,000 Hours of Human Video for Robots, With 90,000 Still to ComeLeaderboardsinfrastructure800/posts/lightwheel-releases-10-000-hours-of-human-video-for-robots-with-90-000-still-to-come
2026-08-29Lemmalog Turns Agent Memory Into Datalog State, Cutting LongMemEval Context 38-FoldAgent and codingtools701/posts/lemmalog-turns-agent-memory-into-datalog-state-cutting-longmemeval-context-38-fold
2026-08-29LAION Releases 10-Million-Hour Video Corpus for Research, Excluding Commercial UseSafety evaluationsinfrastructure802/posts/laion-releases-10-million-hour-video-corpus-for-research-excluding-commercial-use
2026-08-29CME Targets Oct. 5 Nvidia GPU Futures Launch as CFTC Tests the BenchmarkLeaderboardsinfrastructure801/posts/cme-targets-oct-5-nvidia-gpu-futures-launch-as-cftc-tests-the-benchmark
2026-08-29Google’s WikiSkill Lifts Agent Benchmarks by Giving Models a Memory of Failed WorkAgent and codingmodels801/posts/google-s-wikiskill-lifts-agent-benchmarks-by-giving-models-a-memory-of-failed-work
2026-08-29Moorcheh Positions Memanto as a Fleet-Scale Layer for Agent MemoryAgent and codingproducts501/posts/moorcheh-pitches-mit-licensed-memanto-for-agent-memory-beyond-files
2026-08-29Sanctuary AI Puts Its Robot Brain on Existing Machines After a 99.5%-Plus TestLeaderboardsproducts701/posts/sanctuary-ai-puts-physical-ai-on-factory-robots-99-5-plus-test-is-a-proof-of-concept
2026-08-29Israeli Government-Backed Hanover Site Published 560,000 Words to Influence AI AnswersAgent and codingpolicy802/posts/piro-s-hanover-site-published-560-000-words-to-influence-ai-answers
2026-08-30OpenRelay Routes AI Inference Across GPUs, TPUs and Trainium Through One APILeaderboardsinfrastructure800/posts/openrelay-routes-ai-inference-across-gpus-tpus-and-trainium-through-one-api
2026-08-30Chatbots Debunked Foreign Falsehoods About 75% of the Time, Beating Search ResultsCapability evaluationsmodels781/posts/chatbots-debunked-foreign-falsehoods-about-75-of-the-time-beating-search-results
2026-08-30AWS Brings Six AI Model Families to GovCloud Through Bedrock, With Model-Specific Compliance LimitsLeaderboardsproducts801/posts/amazon-bedrock-brings-six-model-families-to-aws-govcloud-for-government-ai-workloads
2026-08-31Jeff Dean Leaves Google to Build Discovery Loop, a Science AI Startup Backed by Radical VenturesSafety evaluationsstartups902/posts/jeff-dean-leaves-google-to-build-discovery-loop-a-science-ai-startup-backed-by-radical-ventures
2026-08-31MIT Maps Julia’s Path to 1 Million Users as JuliaHub Pushes AI Hardware DesignSafety evaluationsproducts701/posts/mit-maps-julia-s-path-to-1-million-users-as-juliahub-pushes-ai-hardware-design
2026-08-31Hack The Box Benchmark: Best Human Team Cleared 36 Challenges; Top AI Team Reached 32Safety evaluationsmodels802/posts/hack-the-box-benchmark-best-human-team-cleared-36-challenges-top-ai-team-reached-32
2026-08-31EU Classifies ChatGPT as Search Engine, Starts Four-Month DSA Compliance ClockCapability evaluationspolicy902/posts/eu-classifies-chatgpt-as-search-engine-starts-four-month-dsa-compliance-clock
2026-08-31Broadcom Releases AgentMinder to Check AI Agent Tool Calls at RuntimeAgent and codingproducts801/posts/broadcom-releases-agentminder-to-check-ai-agent-tool-calls-at-runtime
2026-08-31Unitree’s $4,017 Go2 Pro Makes Robot Dogs Cheaper, Not Yet Everyday ToolsCapability evaluationsproducts750/posts/unitree-s-4-017-go2-pro-makes-robot-dogs-cheaper-not-yet-everyday-tools
2026-08-31Microsoft Releases GigaPath-Flash, Reporting Roughly 50x Less Compute for Cancer ResearchCapability evaluationsmodels801/posts/microsoft-releases-gigapath-flash-reporting-roughly-50x-less-compute-for-cancer-research
2026-08-31Google’s TimesFM-3 Adds Known Future Events to Zero-Shot ForecastingCapability evaluationsmodels801/posts/google-s-timesfm-3-adds-known-future-events-to-zero-shot-forecasting
2026-08-31DeepSeek Releases 305B Vision Weights Under MIT License, but Local Use Needs Serious HardwareAgent and codingmodels900/posts/deepseek-releases-305b-vision-weights-under-mit-license-but-local-use-needs-serious-hardware
2026-08-31Google Wires Gemini Live to Spark, Letting Voice Requests Run Across Its AppsCapability evaluationsproducts801/posts/google-wires-gemini-live-to-spark-letting-voice-requests-run-across-its-apps
2026-08-31AWS Agent Registry Reaches GA With Approval Gates for Enterprise Agent DiscoveryAgent and codingproducts801/posts/aws-agent-registry-reaches-ga-with-approval-gates-for-enterprise-agent-discovery
2026-08-31AWS Publishes Four-Stack Bedrock Blueprint for Observable Agentic RetrievalAgent and codinginfrastructure601/posts/aws-publishes-four-stack-bedrock-blueprint-for-observable-agentic-retrieval
2026-09-01AWS’s Bedrock Chat Blueprint Carries Tenant Filters Across Every Retrieval HopAgent and codinginfrastructure801/posts/aws-s-bedrock-chat-blueprint-carries-tenant-filters-across-every-retrieval-hop
2026-09-01Google Updates Antigravity Teamwork, Putting Long-Running Agent Critique in Paid PreviewAgent and codingtools802/posts/google-updates-antigravity-teamwork-putting-long-running-agent-critique-in-paid-preview
2026-09-01Vanta Releases 65 Agent Controls, Drawing a Line Between Governance and Runtime EnforcementSafety evaluationstools780/posts/vanta-puts-65-ai-agent-controls-in-ga-but-leaves-runtime-blocking-to-other-tools
2026-09-01Vercel Says design.md Cut Known Page Failures 57%, but All Six Pages Still BlockedAgent and codingtools701/posts/vercel-says-design-md-cut-known-agent-page-failures-57-in-a-six-page-test
2026-09-01Scalable Capital Connects ChatGPT, Claude and Grok to Broker Accounts, but Trades Need ApprovalAgent and codingproducts801/posts/scalable-capital-connects-chatgpt-claude-and-grok-to-broker-accounts-but-trades-need-approval
2026-09-01Runway’s Solaris Generates App Interfaces Frame by Frame, but It Isn’t Ready for Public UseLeaderboardsmodels802/posts/runway-s-solaris-generates-live-app-interfaces-frame-by-frame-but-it-isn-t-public-yet
2026-09-01Qwen3.8-Max Leads Open-Weight Commerce Agent Bench by Two Passes, With Audit LimitsAgent and codingmodels800/posts/qwen3-8-max-leads-open-weight-commerce-agent-bench-by-two-passes-with-audit-limits
2026-09-01Shein Targets $1.7B Hong Kong Debut as Shanghai-Hong Kong Proceeds Top $54BLeaderboardsbusiness852/posts/shein-targets-1-7b-hong-kong-debut-as-shanghai-hong-kong-proceeds-top-54b
2026-09-01Salesforce Pushes Agents Into CRM Workflows as NIST Turns to Identity ControlsAgent and codingproducts801/posts/salesforce-pushes-agents-into-crm-workflows-as-nist-turns-to-identity-controls
2026-09-01Flower Labs Starts Endeavor 1.0 Preview With a Private Deployment OptionAgent and codingmodels402/posts/flower-labs-starts-endeavor-1-0-preview-with-a-private-deployment-option
2026-09-01Rezolve Reports $130.8M H1 Revenue, Puts $360M Goal on a Much Larger H2Agent and codingbusiness802/posts/rezolve-reports-130-8m-h1-revenue-puts-360m-goal-on-a-much-larger-h2
2026-09-01Alibaba Opens QwenWork Global Beta for Web and Desktop Work, but Leaves Pricing UnsetAgent and codingproducts801/posts/alibaba-opens-qwenwork-global-beta-for-web-and-desktop-work-but-leaves-pricing-unset
2026-09-01Blue Voice Raises $6M to Put Department Policy AI in Police Officers’ HandsLeaderboardsstartups802/posts/blue-voice-raises-6m-to-put-department-policy-ai-in-police-officers-hands
2026-09-01Orbis Lets Users Alter Prompts Mid-Stream as Visko Takes Its Video Model PublicCapability evaluationsmodels903/posts/visko-opens-orbis-with-live-4k-video-backed-by-10m-pre-seed
2026-09-01Amazon’s Alexa Adds Update Me When Alerts Before Shoppers SearchAgent and codingproducts702/posts/amazon-s-alexa-adds-update-me-when-alerts-before-shoppers-search
2026-09-01AWS Publishes Jamf’s Bedrock Budget Gate, Restricting Premium Models at 80% SpendLeaderboardstools701/posts/aws-publishes-jamf-s-bedrock-budget-gate-restricting-premium-models-at-80-spend
2026-09-01t54 Puts a Hard Trust Gate on Bedrock Agent Payments, Citing 20 Million MicropaymentsAgent and codinginfrastructure701/posts/t54-puts-a-hard-trust-gate-on-bedrock-agent-payments-citing-20-million-micropayments
2026-09-01AWS Details ZS’s Internet-Free SageMaker Platform Across 200+ DomainsCapability evaluationsinfrastructure701/posts/aws-details-zs-s-internet-free-sagemaker-platform-across-200-domains
2026-09-01Grok 4.6 Clears 50% on Biosecurity Refusal and Routine Biology TasksAgent and codingmodels801/posts/grok-4-6-clears-50-on-biosecurity-refusal-and-routine-biology-tasks
2026-09-01Atos and AWS Put 400 Engineers Into a Scored Multi-Agent AI ContestSafety evaluationsbusiness602/posts/atos-and-aws-put-400-engineers-into-a-scored-multi-agent-ai-contest
2026-09-01Databricks Sets a Six-Layer Path for Genie Ontology, Starting With One DomainAgent and codingproducts602/posts/databricks-sets-a-six-layer-path-for-genie-ontology-starting-with-one-domain
2026-09-01Google Adds Agentic Video Understanding to Gemini, Claiming 88% Lower Token UseAgent and codingmodels801/posts/google-adds-agentic-video-understanding-to-gemini-claiming-88-lower-token-use
2026-09-01Meta Introduces Muse Voice Transcribe for Live Speaker Labels in 25 Validated LanguagesLeaderboardsmodels801/posts/meta-introduces-muse-voice-transcribe-for-live-speaker-labels-in-25-validated-languages
2026-09-01Silicon Data’s Token Index Falls to 97 Cents, Tightening the Squeeze on Model PricingLeaderboardsbusiness703/posts/silicon-data-s-token-index-falls-to-97-cents-tightening-the-squeeze-on-model-pricing
2026-09-01Nori Lists Its A3 Home Robot at $1,688, but Buyers Will Help Teach It What to DoLeaderboardsproducts801/posts/nori-lists-its-a3-home-robot-at-1-688-but-buyers-will-help-teach-it-what-to-do
2026-09-01GPT-5 and Gemini-3 Recovered Up to 65% of Missed Facts by Thinking LongerAgent and codingmodels801/posts/gpt-5-and-gemini-3-recovered-up-to-65-of-missed-facts-by-thinking-longer
2026-09-01Jason Isbell’s Suno Suit Targets Artist-Name Prompts, Not Music CopyrightCapability evaluationspolicy802/posts/jason-isbell-s-suno-suit-targets-artist-name-prompts-not-music-copyright
2026-09-01Anthropic Releases One Claude 5.1 Model in Two Access Tiers, Cuts Cache Reads 75%Safety evaluationsmodels902/posts/anthropic-releases-fable-5-1-with-25-cost-estimate-restricts-mythos-5-1
2026-09-01South Korea Cuts AI Model Contest to 3, Eliminating Benchmark Leader MotifLeaderboardspolicy801/posts/south-korea-cuts-ai-model-contest-to-3-eliminating-benchmark-leader-motif
2026-09-01World Labs Launches Atlas for 1440p Camera-Controlled Video, 3D Worlds and Robot ViewsCapability evaluationsmodels801/posts/world-labs-launches-atlas-for-1440p-camera-controlled-video-3d-worlds-and-robot-views
2026-09-01AIR Raises $50M for Continuous Checks on AI-Agent Add-OnsAgent and codingstartups852/posts/air-security-raises-50m-to-put-continuous-checks-around-agent-add-ons
2026-09-01Army’s $10B ITES-4H Tests Whether Faster AI Infrastructure Buying Can Keep Its GuardrailsSafety evaluationsinfrastructure801/posts/army-s-10b-ites-4h-tests-whether-faster-ai-infrastructure-buying-can-keep-its-guardrails
2026-09-01CrowdStrike’s SafeMind Pits Red and Blue AI Agents in an NVIDIA Infrastructure TwinSafety evaluationsproducts802/posts/crowdstrike-s-safemind-pits-red-and-blue-ai-agents-in-an-nvidia-infrastructure-twin
2026-09-01Anthropic Resumes Claude Cyber Tests With a Real-Time Stop System After Live-Web IncidentsSafety evaluationsmodels802/posts/anthropic-resumes-claude-cyber-tests-with-a-real-time-stop-system-after-live-web-incidents
2026-09-02Baseten’s New Inference Guide Separates Configuration Tradeoffs From True Efficiency GainsAgent and codinginfrastructure502/posts/baseten-maps-the-llm-inference-tradeoffs-between-speed-throughput-and-quality
2026-09-02Cursor Adds Claude Fable 5.1, Pairing a 73.4% Coding Score With Max-Effort CostsAgent and codingmodels802/posts/cursor-adds-claude-fable-5-1-pairing-a-73-4-coding-score-with-max-effort-costs
2026-09-02EU Seeks AI Act Answers From 30-Plus Companies, With Fines for Misleading RepliesSafety evaluationspolicy802/posts/eu-seeks-ai-act-answers-from-30-plus-companies-with-fines-for-misleading-replies
2026-09-02Google May Ship Gemini 3.8 Flash 20 Days After 3.7, With Coding Edge UnprovenSafety evaluationsmodels800/posts/google-could-release-gemini-3-8-flash-20-days-after-3-7-but-public-evidence-is-missing
2026-09-02Cognition Reportedly Seeks $1B at $47B, Three Months After $26B MarkAgent and codingstartups803/posts/cognition-targets-1b-at-47b-as-ai-coding-funding-talks-stay-unfinished
2026-09-02Aranya Raises $11M for Software It Says Can Ready GPU Clusters in 48 HoursCapability evaluationsstartups802/posts/aranya-raises-11m-to-turn-bare-metal-gpu-servers-into-production-clusters
2026-09-02Anthropic’s Fable 5.1 Takes Benchmark Lead, but Its Top Setting Costs More Per TaskSafety evaluationsmodels804/posts/anthropic-s-fable-5-1-leads-a-benchmark-but-costs-20-more-at-max-effort
2026-09-02Job Seeker Sends ChatGPT to AI Recruiter After Five Interviews Without Follow-UpAgent and codingculture502/posts/job-seeker-sends-chatgpt-to-ai-recruiter-after-five-interviews-without-follow-up
2026-09-02Anthropic’s 80-Environment Reward-Hacking Test Produced Cyber and Safety EvasionsSafety evaluationsmodels801/posts/anthropic-s-80-environment-reward-hacking-test-produced-cyber-and-safety-evasions
2026-09-02Nature Study Finds AI Agent Teams Can Cut Sequential Planning Performance by 70%Agent and codingmodels802/posts/nature-study-finds-ai-agent-teams-can-cut-sequential-planning-performance-by-70
2026-09-02MIT and Motional Build AI That Exposes Robotaxi Planning ErrorsSafety evaluationsmodels802/posts/mit-and-motional-build-ai-that-exposes-robotaxi-planning-errors
2026-09-02Empirik Emerges With $21M to Check Infrastructure Changes Before DeploymentAgent and codingstartups803/posts/empirik-emerges-with-21m-to-check-infrastructure-changes-before-deployment
2026-09-02LivePerson Stockholders Approve SoundHound AI Deal, Targeting a September 4 CloseAgent and codingbusiness851/posts/liveperson-stockholders-approve-soundhound-ai-deal-targeting-a-september-4-close
2026-09-02New Mexico Lets Districts Seek Approval to Leave K-2 AI Reading TestsLeaderboardspolicy801/posts/new-mexico-lets-districts-seek-approval-to-leave-k-2-ai-reading-tests
2026-09-02LlamaIndex and Kaggle Launch a Document Benchmark for Cost and Long FilesAgent and codingtools901/posts/llamaindex-and-kaggle-launch-a-document-benchmark-for-cost-and-long-files
2026-09-02CrowdStrike Wants AI Agents Rechecked Before Every ActionAgent and codingproducts702/posts/crowdstrike-wants-ai-agents-rechecked-before-every-action
2026-09-02Snowflake Lifts Product-Revenue Forecast After Q2 Beat; CoCo Hits 9,100 AccountsAgent and codingbusiness801/posts/snowflake-lifts-product-revenue-forecast-after-q2-beat-coco-hits-9-100-accounts
2026-09-02Google Agrees to Buy 396 MW of Geothermal Power for a Planned Utah Data CenterLeaderboardsinfrastructure803/posts/google-agrees-to-buy-396-mw-of-future-geothermal-power-for-utah-data-center
2026-09-02Enterprise Buyers Put Non-Nvidia Chip Evaluations Ahead of Nvidia’s Next GPUsCapability evaluationsinfrastructure801/posts/enterprise-buyers-put-non-nvidia-chip-evaluations-ahead-of-nvidia-s-next-gpus
2026-09-03Capsule Security Adds a Last-Second Block for AI-Agent ActionsAgent and codingtools612/posts/capsule-security-adds-a-last-second-block-for-ai-agent-actions
2026-09-03Meta Drops Token Counts From Reviews While Employees Test HatchAgent and codingbusiness681/posts/meta-drops-token-counts-from-reviews-while-employees-test-hatch
2026-09-03InferenceX Puts GB300 First for MiniMax M2.7, but Only on a Fixed Chat TestAgent and codinginfrastructure582/posts/inferencex-puts-gb300-first-for-minimax-m2-7-but-only-on-a-fixed-chat-test
2026-09-03Artificial Analysis Rebuilds Image Editing Arena With Different Leaders by TaskLeaderboardsmodels661/posts/artificial-analysis-rebuilds-image-editing-arena-with-different-leaders-by-task
2026-09-03NTU Researchers Release Puffin-World to Put Camera Geometry Inside a World ModelCapability evaluationsmodels620/posts/ntu-researchers-release-puffin-world-to-put-camera-geometry-inside-a-world-model
2026-09-03Qdrant Publishes a 10B-Vector Benchmark and Open Tools to Rerun ItCapability evaluationstools611/posts/qdrant-publishes-a-10b-vector-benchmark-and-open-tools-to-rerun-it
2026-09-03Perplexity Open-Sources Lily, a Qwen-Specific Mac Engine That Beats MLX-LM in Its TestAgent and codinginfrastructure661/posts/perplexity-open-sources-lily-a-mac-engine-35-faster-in-its-qwen-test
2026-09-03Cohere Publishes Agent-Tool Dataset, Finds 419 Occupations Have No CoverageAgent and codingtools681/posts/cohere-publishes-agent-tool-dataset-finds-419-occupations-have-no-coverage
2026-09-03Figure Signs $3.5B Nscale Deal for Up to 100,000 GPUs for Humanoid AICapability evaluationsinfrastructure822/posts/figure-secures-3-5b-compute-deal-for-up-to-100-000-gpus-to-train-humanoid-ai

Showing 250 of 264 rows. Download the dataset for the complete table.

Measurement technique

How to read this report

  1. 01Records require a named evaluation, benchmark result, leaderboard change, or benchmark-integrity development.
  2. 02The series counts developments and groups them by evaluation role; it does not compare incompatible score scales.
  3. 03Primary benchmark papers and maintainers remain authoritative for exact protocols and revisions.

Sources

Evidence

45 publishers supporting 61 records. Expand a publisher to inspect its cited pages.

TechCrunchtechcrunch.com7 sources / 7 records
CNBCcnbc.com4 sources / 4 records
finance.yahoo.comfinance.yahoo.com4 sources / 4 records
aws.amazon.comaws.amazon.com1 source / 2 records
inferencex.semianalysis.cominferencex.semianalysis.com2 sources / 2 records
siliconangle.comsiliconangle.com2 sources / 2 records
Show 39 more publishers
spectrumnews1.comspectrumnews1.com2 sources / 2 records
404media.co404media.co1 source / 1 record
ai21.comai21.com1 source / 1 record
Bloombergbloomberg.com1 source / 1 record
businesswire.combusinesswire.com1 source / 1 record
caixinglobal.comcaixinglobal.com1 source / 1 record
cnbctv18.comcnbctv18.com1 source / 1 record
cryptonews.netcryptonews.net1 source / 1 record
developer.nvidia.comdeveloper.nvidia.com1 source / 1 record
educationnext.orgeducationnext.org1 source / 1 record
entrackr.comentrackr.com1 source / 1 record
finance.biggo.comfinance.biggo.com1 source / 1 record
fireworks.aifireworks.ai1 source / 1 record
forbes.comforbes.com1 source / 1 record
fortune.comfortune.com1 source / 1 record
future-ed.orgfuture-ed.org1 source / 1 record
latimes.comlatimes.com1 source / 1 record
marketscale.commarketscale.com1 source / 1 record
mediapost.commediapost.com1 source / 1 record
mk.co.krmk.co.kr1 source / 1 record
news.bloombergtax.comnews.bloombergtax.com1 source / 1 record
nypost.comnypost.com1 source / 1 record
openai.comopenai.com1 source / 1 record
platformer.newsplatformer.news1 source / 1 record
pulse2.compulse2.com1 source / 1 record
pymnts.compymnts.com1 source / 1 record
qz.comqz.com1 source / 1 record
scmp.comscmp.com1 source / 1 record
securityweek.comsecurityweek.com1 source / 1 record
snowflake.comsnowflake.com1 source / 1 record
startupfortune.comstartupfortune.com1 source / 1 record
techbuzz.aitechbuzz.ai1 source / 1 record
thenextweb.comthenextweb.com1 source / 1 record
theregister.comtheregister.com1 source / 1 record
tvtechnology.comtvtechnology.com1 source / 1 record
venturebeat.comventurebeat.com1 source / 1 record
wmbdradio.comwmbdradio.com1 source / 1 record
wmtv15news.comwmtv15news.com1 source / 1 record
wsaw.comwsaw.com1 source / 1 record
Next report / 20AI Partnership and Dependency Graph All research reports