2026-08-17Meta Launches Muse Code With a 92% Discount for Training Rights Agent and coding models 69 1 /posts/meta-launches-muse-code-with-a-92-discount-for-training-rights 2026-08-17Anthropic’s Dario Amodei Says AI Must Deliver, Not Advertise, Its Way Out of a Trust Crisis Safety evaluations culture 39 3 /posts/anthropic-s-dario-amodei-says-ai-must-deliver-not-advertise-its-way-out-of-a-trust-crisis 2026-08-17Zuckerberg’s superintelligence promise runs into AI’s trust test Safety evaluations culture 65 1 /posts/zuckerberg-s-superintelligence-promise-runs-into-ai-s-trust-test 2026-08-17A Naming Error Let Anthropic Models Reach a Real Production Database Safety evaluations models 46 2 /posts/a-naming-error-let-anthropic-models-reach-a-real-production-database 2026-08-17Alibaba Launches HappyShrimp for One-Prompt Songs, but Leaves Key Use Terms Unstated Leaderboards models 46 1 /posts/alibaba-launches-happyshrimp-for-one-prompt-songs-but-leaves-key-use-terms-unstated 2026-08-18Zhipu Says GLM-5.3 Can Find Bugs Across an Exploitation Chain Safety evaluations models 72 2 /posts/zhipu-says-glm-5-3-can-find-bugs-across-an-exploitation-chain 2026-08-18Alibaba Releases Downloadable Qwen3.8-Max Weights Alongside a Laptop Model Agent and coding models 63 3 /posts/alibaba-puts-qwen-on-laptops-and-opens-its-biggest-model-as-meta-courts-the-same-developers 2026-08-18Nvidia’s Reported Lancium Deal Would Put Grid Connections at the Center of Its AI Buildout Leaderboards business 35 3 /posts/nvidia-s-ai-infrastructure-bet-reaches-the-grid-not-just-the-data-center 2026-08-18ByteDance’s Hollywood Copyright Truce Leaves the Guardrails Private Leaderboards policy 53 4 /posts/hollywood-and-bytedance-trade-a-copyright-fight-for-an-unseen-ai-guardrail-deal 2026-08-18Snowflake Wants AI Apps to Stop Paying Frontier-Model Prices for Every Task Agent and coding products 77 2 /posts/snowflake-wants-ai-apps-to-stop-paying-frontier-model-prices-for-every-task 2026-08-18OpenAI Funds a Study of How AI Could Shift Tax Revenue Safety evaluations policy 41 1 /posts/openai-funds-a-study-of-how-ai-could-shift-tax-revenue 2026-08-18OpenAI Paused Deployment-Bound Training as It Tightened Frontier Security Safety evaluations policy 8 1 /posts/openai-paused-deployment-bound-training-as-it-tightened-frontier-security 2026-08-18Etched’s $700 Million Round Puts a $21 Billion Price on Inference Hardware Capability evaluations startups 54 3 /posts/etched-s-700-million-round-puts-a-21-billion-price-on-inference-hardware 2026-08-18Wispr’s $280 Million Bet Is That Voice Can Move Beyond Dictation Leaderboards startups 56 4 /posts/wispr-s-280-million-bet-is-that-voice-can-move-beyond-dictation 2026-08-19AgentX Replays Claude Code Sessions to Test AI Serving Systems Agent and coding tools 72 1 /posts/agentx-replays-claude-code-sessions-to-test-ai-serving-systems 2026-08-19TestMu’s New Agent Test Shows What It Can’t Verify Safety evaluations products 43 1 /posts/testmu-s-new-agent-test-shows-what-it-can-t-verify 2026-08-19LMCache Reworks the Cache Plumbing That Can Stall Long-Running Agents Agent and coding tools 65 1 /posts/lmcache-reworks-the-cache-plumbing-that-can-stall-long-running-agents 2026-08-19Dane County Routes Non-Emergency Calls Through AVA to Protect 911 Capacity Safety evaluations products 38 4 /posts/dane-county-puts-an-ai-gatekeeper-on-its-non-emergency-line 2026-08-19Oakley Buys Majority Stake in Graphwise as It Plans Expansion and Acquisitions Agent and coding business 44 2 /posts/oakley-takes-majority-control-of-graphwise-as-it-plans-a-global-ai-push 2026-08-19Jinho Jang Puts a 27B Refusal-Removed Model Into a Local Download Safety evaluations models 80 0 /posts/jinho-jang-puts-a-27b-refusal-removed-model-into-a-local-download 2026-08-19Nome Uses AI to Find Rare-Disease Treatment Paths Before the Hard Work Begins Agent and coding startups 65 1 /posts/nome-promises-rare-disease-families-a-treatment-path-the-clinical-work-still-lies-ahead 2026-08-19Anthropic Will Watermark Claude Text by Steering Its Word Choices Capability evaluations tools 58 1 /posts/anthropic-will-watermark-claude-text-by-steering-its-word-choices 2026-08-19A Reporter’s LLM Wiki Speeds Recall. He Still Checks the Original Sources. Agent and coding tools 55 1 /posts/casey-newton-built-an-ai-wiki-for-his-beat-now-he-has-to-keep-it-alive 2026-08-19Vivodyne Builds a Human-Tissue Data Factory for Drug AI Capability evaluations startups 68 1 /posts/vivodyne-built-a-robot-lab-for-the-human-biology-data-ai-still-lacks 2026-08-20SandboxAQ Opens a Drug-Screening Model That Does Not Need Protein Structures Capability evaluations models 78 0 /posts/sandboxaq-opens-a-drug-screening-model-that-does-not-need-protein-structures 2026-08-20OpenAI Tests Cross-Session Safety Checks Without Prompt Access Safety evaluations products 76 0 /posts/openai-tests-cross-session-safety-checks-without-prompt-access 2026-08-20OpenAI Takes Codex From Coding to Tax Returns Agent and coding tools 68 0 /posts/openai-takes-codex-from-coding-to-tax-returns 2026-08-20TrueFoundry Gives Away Its Agent Runtime to Sell the Layer Beneath It Agent and coding tools 65 0 /posts/truefoundry-gives-away-its-agent-runtime-to-sell-the-layer-beneath-it 2026-08-20Palomar Opens a Lean Registry That Checks Proofs and Their Descriptions Capability evaluations tools 72 0 /posts/palomar-opens-a-lean-proof-registry-with-an-llm-semantic-check 2026-08-20UK Cyber Tests Show AI Agents Going Beyond the Technical Task Safety evaluations models 84 0 /posts/ai-cyber-tests-are-reaching-real-targets-not-just-sandboxes 2026-08-20At This D.C. Charter, AI Permission Changes With the Assignment Leaderboards culture 65 2 /posts/this-d-c-charter-made-ai-a-schoolwide-skill-not-a-shortcut 2026-08-20Mistral Gives Enterprise AI Five Ways to Keep Digging Through Documents Agent and coding products 68 0 /posts/mistral-gives-enterprise-ai-five-ways-to-keep-digging-through-documents 2026-08-20Simple AI Opens 2,000 Hours of Robot Training Data Collected Without Robots Leaderboards tools 67 0 /posts/simple-ai-opens-2-000-hours-of-robot-training-data-collected-without-robots 2026-08-20Ramp Opens Its AI Router to U.S. Customers, With a One-Year Data Default Capability evaluations products 58 1 /posts/ramp-opens-its-ai-router-to-u-s-customers-with-a-one-year-data-default 2026-08-20Anthropic Opens Claude Academy as Free Training for Its AI Products Agent and coding products 68 0 /posts/anthropic-opens-claude-academy-as-free-training-for-its-ai-products 2026-08-21Micro1’s Reported $500M Run Rate Tests the Economics of AI Training Data Leaderboards startups 55 1 /posts/micro1-s-reported-500m-run-rate-tests-the-economics-of-ai-training-data 2026-08-21Callosum Raises $100 Million as It Routes AI Work Across Models and Chips Capability evaluations startups 66 2 /posts/callosum-raises-100-million-to-route-ai-work-across-models-and-chips 2026-08-21Starcloud Has $250 Million for Orbital AI. It Still Needs a Ride. Leaderboards startups 66 1 /posts/starcloud-has-250-million-for-orbital-ai-it-still-needs-a-ride 2026-08-21Open Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest. Agent and coding models 74 1 /posts/open-models-are-catching-the-frontier-faster-benchmark-scores-aren-t-the-whole-contest 2026-08-21Nvidia Maps AI Memory Between Models Instead of Making Them Start Over Capability evaluations tools 74 1 /posts/nvidia-maps-ai-memory-between-models-instead-of-making-them-start-over 2026-08-21Nvidia’s AVO Clears ARC-AGI-3’s Public Set. Withheld Tests Still Matter. Agent and coding models 72 1 /posts/nvidia-s-avo-clears-arc-agi-3-s-public-set-withheld-tests-still-matter 2026-08-21OpenAI Makes GPT-5.6 Sol Cheaper for Metered Use, Not Easier to Access Capability evaluations products 72 0 /posts/openai-cuts-sol-s-output-price-but-not-chatgpt-s-limits 2026-08-21Grok Reversed a China-Influence Finding After an Audit of Its Sources Agent and coding models 65 0 /posts/grok-reversed-its-china-campaign-finding-after-an-audit-exposed-the-citation-chain 2026-08-21Anthropic’s IPO Pitch Faces Two Tests: Compute Growth and Enterprise Data Control Safety evaluations business 77 2 /posts/anthropic-s-ipo-pitch-faces-two-tests-compute-growth-and-enterprise-data-control 2026-08-22DeepSeek Gives V4-Flash a Separate Vision API With a 384-Token Image Cap Agent and coding models 71 1 /posts/deepseek-adds-vision-to-an-experimental-v4-flash-endpoint-with-a-384-token-image-cap 2026-08-22Meta Is a Major Microsoft AI Customer—and a Potential Foundry Rival Leaderboards business 75 1 /posts/meta-is-a-major-microsoft-ai-customer-and-a-potential-foundry-rival 2026-08-22Meta’s 30B Muse Glimmer Tries to Make 131K Context Fit in 24 GB Agent and coding models 45 1 /posts/meta-s-30b-muse-glimmer-tries-to-make-131k-context-fit-in-24-gb 2026-08-22AI21’s 8B Verifier Challenges the Case for Bigger Search Models Agent and coding models 68 1 /posts/ai21-s-8b-verifier-challenges-the-case-for-bigger-search-models 2026-08-22AWS’s RAG Cost Cut Comes With a 19% Latency Bill Capability evaluations tools 62 1 /posts/aws-s-rag-cost-cut-comes-with-a-19-latency-bill 2026-08-22GLM-5.3’s Cheap Retries Put Fable 5’s Coding Premium Under Pressure Agent and coding models 64 1 /posts/glm-5-3-s-cheap-retries-put-fable-5-s-coding-premium-under-pressure 2026-08-22Oracle Wants AI Agents to Pick Trusted Reports, Not Write SQL Agent and coding tools 57 1 /posts/oracle-wants-ai-agents-to-pick-trusted-reports-not-write-sql 2026-08-22Thinking Machines Releases Inkling, but Argues Open Weights Need a Gate Safety evaluations policy 72 1 /posts/thinking-machines-releases-inkling-but-argues-open-weights-need-a-gate 2026-08-22FDA Opens Door to Clinician-Style Tests for Medical AI Safety evaluations policy 45 2 /posts/fda-opens-door-to-clinician-style-tests-for-medical-ai 2026-08-22Databricks Pushes Feature Stores From Batch Lag to 200ms Freshness Capability evaluations products 72 1 /posts/databricks-pushes-feature-stores-from-batch-lag-to-200ms-freshness 2026-08-22Stanford Won Databricks’ Agent Cup, but 18.8% of Questions Stumped Every Team Agent and coding models 71 2 /posts/stanford-won-databricks-agent-cup-but-18-8-of-questions-stumped-every-team 2026-08-22Ox Alpha Offers a Million Tokens, but Not a Name Agent and coding models 65 1 /posts/ox-alpha-offers-a-million-tokens-but-not-a-name 2026-08-22Roblox Opens Three AI Safety Models for Child-Protection Teams Safety evaluations tools 68 1 /posts/roblox-opens-three-ai-safety-models-for-child-protection-teams 2026-08-22Serval’s Catalyst Pitches AI-Built Workflows as a ServiceNow Replacement Agent and coding startups 58 1 /posts/serval-s-catalyst-pitches-ai-built-workflows-as-a-servicenow-replacement 2026-08-22Microsoft and Qcells Want AI Data Centers to Bring Their Own Power Leaderboards business 72 2 /posts/microsoft-and-qcells-want-ai-data-centers-to-bring-their-own-power 2026-08-22Veeda AI Raises $90M to Make Robot Training Less Physical Agent and coding startups 68 2 /posts/veeda-ai-raises-90m-to-make-robot-training-less-physical 2026-08-22LinkedIn Measures AI Code Review Against Merged Code Agent and coding business 67 1 /posts/linkedin-uses-multiple-ai-reviewers-to-cut-code-review-noise 2026-08-22TrueForge Puts the Agent Harness, Not the Model, at the Center of Cost Control Agent and coding models 72 0 /posts/trueforge-puts-the-agent-harness-not-the-model-at-the-center-of-cost-control 2026-08-22Hollywood’s Downturn Is Turning Creative Know-How Into AI Training Data Leaderboards culture 68 1 /posts/hollywood-s-downturn-is-turning-creative-know-how-into-ai-training-data 2026-08-22Panasonic Says Its Aircraft AI Cut Diagnostic Investigations From Hours to Minutes Agent and coding business 63 1 /posts/panasonic-says-its-aircraft-ai-cut-diagnostic-investigations-from-hours-to-minutes 2026-08-22AgentFlo’s Sales Agents Put Rules and Data Above the Model Agent and coding products 64 1 /posts/agentflo-s-sales-agents-put-rules-and-data-above-the-model 2026-08-22Generalist’s GEN-1.5 Lets Robots Try a Task After Watching Once Capability evaluations models 68 1 /posts/generalist-s-gen-1-5-lets-robots-try-a-task-after-watching-once 2026-08-22AI Safety Scores Can Reward Models for Refusing Too Much Safety evaluations models 70 1 /posts/ai-safety-scores-can-reward-models-for-refusing-too-much 2026-08-22Massachusetts Could Make AI Labs Face Public Risk Tests Every 120 Days Safety evaluations policy 78 2 /posts/massachusetts-could-make-ai-labs-face-public-risk-tests-every-120-days 2026-08-22AWS Wants AI Agents to Carry the User’s Permissions, Not Their Own Agent and coding tools 70 2 /posts/aws-wants-ai-agents-to-carry-the-user-s-permissions-not-their-own 2026-08-22Ant Puts Its FX Forecasting Model Into Tools Used by Major Banks Leaderboards models 68 2 /posts/ant-says-six-banks-signed-on-to-its-fx-forecasting-ai 2026-08-22Inherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper Replication Agent and coding models 55 1 /posts/inherent-says-its-27b-science-agent-beat-openai-and-anthropic-on-paper-replication 2026-08-22MiniMax H3 Opens Its Video Weights, but Not Where Many Developers Work Leaderboards models 72 1 /posts/minimax-h3-opens-its-video-weights-but-not-where-many-developers-work 2026-08-22Murf’s Falcon 2 Puts a One-Cent Bet on Real-Time Voice Leaderboards models 69 2 /posts/murf-s-falcon-2-pitches-real-time-voice-at-a-cent-a-minute 2026-08-22OpenAI Moves Astra’s Cybersecurity Gate Into Training Safety evaluations models 39 2 /posts/openai-slows-astra-after-a-sandbox-breach-exposes-gaps-in-its-safety-controls 2026-08-22SemiAnalysis Says AgentX Drove 50-Plus Upstream Fixes for AI Agents Agent and coding tools 75 2 /posts/agentx-pushes-ai-serving-fixes-from-routing-to-kernels 2026-08-23ATTOM Adds 3 AI Agents That Turn Licensed Property Data Into Research and Reports Agent and coding products 50 2 /posts/attom-turns-its-property-database-into-three-ai-workflows 2026-08-23University of Konstanz Finds AI Agents Coordinate Up to 1,000, but Consensus Can Be Wrong Safety evaluations models 63 2 /posts/university-of-konstanz-finds-ai-agents-coordinate-up-to-1-000-but-consensus-can-be-wrong 2026-08-23OpenAI Calls to Amend California’s SB 53 With Model Monitoring, Reversing 2024 Opposition Safety evaluations policy 72 4 /posts/openai-calls-to-amend-california-s-sb-53-with-model-monitoring-reversing-2024-opposition 2026-08-24Hugging Face Explores $13B Sale That Could Put Its Shared AI Hub Under One Owner Leaderboards business 68 1 /posts/hugging-face-explores-13b-sale-that-could-put-its-shared-ai-hub-under-one-owner 2026-08-24InferenceX Adds 11 Telemetry Views to AgentX, Exposing What Benchmark Curves Hide Agent and coding tools 63 2 /posts/inferencex-adds-11-telemetry-views-to-agentx-exposing-what-benchmark-curves-hide 2026-08-24Goldman Deploys Claude Agents as Lloyds Targets £100M in Value From 2026 AI Plan Agent and coding business 68 1 /posts/goldman-deploys-claude-agents-as-lloyds-targets-100m-in-value-from-2026-ai-plan 2026-08-24GPT-BERT Beats Llama 2 70B on One Grammar Test With 100 Million Words Leaderboards models 68 1 /posts/gpt-bert-beats-llama-2-70b-on-one-grammar-test-with-100-million-words 2026-08-24Alibaba Rolls Out Wan3.0, Turning Business Files Into 30-Second AI Videos Capability evaluations models 76 5 /posts/alibaba-rolls-out-wan3-0-turning-business-files-into-30-second-ai-videos 2026-08-24Thomson Reuters Puts Its First Legal AI Model Into CoCounsel Document Review Agent and coding models 72 1 /posts/thomson-reuters-puts-its-first-legal-ai-model-into-cocounsel-document-review 2026-08-24Google and USC’s ME-POIs Adds Mobility Data to Place AI, Lifting Visit-Intent F1 by 81.9% Capability evaluations models 58 1 /posts/google-and-usc-s-me-pois-adds-mobility-data-to-place-ai-lifting-visit-intent-f1-by-81-9 2026-08-24OpenAI Lets Codex Route Smaller Tasks From Sol to Lower-Cost Luna Workers Agent and coding models 65 0 /posts/openai-lets-codex-route-smaller-tasks-from-sol-to-lower-cost-luna-workers 2026-08-24Blitzy and XBOW Push AI Beyond Short Tasks Toward Continuous Enterprise Work Agent and coding business 68 1 /posts/blitzy-and-xbow-push-ai-beyond-short-tasks-toward-continuous-enterprise-work 2026-08-24NVIDIA Says Vera Rubin Delivers 30x More Agentic Throughput per Megawatt, Pending Review Agent and coding models 76 1 /posts/nvidia-says-vera-rubin-delivers-30x-more-agentic-throughput-per-megawatt-pending-review 2026-08-24Nvidia Puts Groq 3 LPX Into Production for Faster AI-Agent Responses Agent and coding products 85 2 /posts/nvidia-puts-groq-3-lpx-into-production-for-faster-ai-agent-responses 2026-08-24OpenAI’s $20 ChatGPT Work Brings Codex Agents to Office Apps, With Access Still a Barrier Agent and coding products 78 1 /posts/openai-s-20-chatgpt-work-brings-codex-agents-to-office-apps-with-access-still-a-barrier 2026-08-24NVIDIA Says AVO Took Claude Opus 5 From 30% to Perfect on ARC-AGI-3 Agent and coding tools 68 1 /posts/nvidia-says-avo-took-claude-opus-5-from-30-to-perfect-on-arc-agi-3 2026-08-24AWS Publishes Metadata Workflow That Escalates Ambiguous Fixes to Bedrock LLMs Agent and coding tools 56 1 /posts/aws-publishes-metadata-workflow-that-sends-ambiguous-field-fixes-to-llms 2026-08-24Meta Sends MetaRoCE to OCP for Loss-Tolerant AI Ethernet Capability evaluations tools 78 1 /posts/meta-sends-metaroce-to-ocp-for-loss-tolerant-ai-ethernet 2026-08-24AWS Brings Ray Into SageMaker HyperPod With Recovery Tools and Tiered Cache on EKS Agent and coding products 76 1 /posts/aws-brings-ray-into-sagemaker-hyperpod-with-recovery-tools-and-tiered-cache-on-eks 2026-08-24Microsoft Turns AI Governance Into Runtime Controls Across Nine Domains Safety evaluations policy 78 1 /posts/microsoft-turns-ai-governance-into-runtime-controls-across-nine-domains 2026-08-24Alabama Subpoenas OpenAI Over Hugging Face Breach, Testing a New Enforcement Route Safety evaluations policy 78 1 /posts/alabama-subpoenas-openai-over-hugging-face-breach-testing-a-new-enforcement-route 2026-08-25Nvidia’s NeMo Switchyard Routes Agent Calls Across Models, Not One Default Agent and coding tools 76 1 /posts/nvidia-s-nemo-switchyard-routes-agent-calls-across-models-not-one-default 2026-08-25HiDream-O1-World Tops WBench Navi at 80.9, Betting on Persistent 3D Scenes Leaderboards models 68 1 /posts/hidream-o1-world-tops-wbench-navi-at-80-9-betting-on-persistent-3d-scenes 2026-08-25Oracle Puts Access Filters Before Agent Search—and Reranking Adds 2.2 Seconds Agent and coding tools 58 1 /posts/oracle-details-hybrid-agent-memory-retrieval-with-a-2-2-second-reranking-tradeoff 2026-08-25Kimi.ai Uses TiDB for One-Second Agent Databases and Persistent Development State Agent and coding tools 58 1 /posts/kimi-ai-uses-tidb-for-one-second-agent-databases-and-persistent-development-state 2026-08-25Apple Refreshes Mac mini and Mac Studio for Linked Local AI, From $899 Leaderboards products 72 1 /posts/apple-refreshes-mac-mini-and-mac-studio-for-linked-local-ai-from-899 2026-08-25Everlaw Connects Gemini Legal to Governed Evidence as Weil Deploys It Leaderboards business 64 3 /posts/everlaw-puts-gemini-legal-on-governed-evidence-in-private-beta-weil-deploys-it 2026-08-25Anthropic Connects Claude Science to 60+ Databases and Tools for Enterprise Work Agent and coding business 78 1 /posts/anthropic-connects-claude-science-to-60-databases-and-tools-for-enterprise-work 2026-08-25OpenAI’s Jalapeño Claims 1.5–1.9x More AI Work Per Watt, Faces 2027 Scale Test Agent and coding products 85 2 /posts/openai-s-jalape-o-claims-1-5-1-9x-more-ai-work-per-watt-faces-2027-scale-test 2026-08-25Radiology Holds Three-Quarters of Cleared Medical AI—and a New Veto Problem Safety evaluations culture 68 1 /posts/radiology-holds-three-quarters-of-cleared-medical-ai-and-a-new-veto-problem 2026-08-25Tiangong Ultra Runs 100m in 8.86 Seconds, Then Hits a Stopping Mat Capability evaluations products 67 1 /posts/tiangong-ultra-runs-100m-in-8-86-seconds-then-hits-a-stopping-mat 2026-08-25Perplexity’s Portable Computer Runs Agents Locally—but Needs a 24GB Nvidia GPU Agent and coding products 68 2 /posts/perplexity-s-portable-computer-keeps-ai-agents-local-if-you-have-24gb-of-vram 2026-08-25Relativity and iManage Give Gemini Legal Two Jobs: Administration and Knowledge Retrieval Agent and coding products 62 3 /posts/relativityone-connects-gemini-legal-through-mcp-for-matter-and-access-administration 2026-08-25Baseten Builds Frontier Gateway Into the Inference Path for Customer API Controls Leaderboards products 55 1 /posts/baseten-positions-frontier-gateway-for-model-labs-selling-multi-tenant-ai-apis 2026-08-25Oracle Adds AMD GPU Operator to OKE for Broader GPU Lifecycle Management Capability evaluations products 56 1 /posts/oracle-puts-amd-gpu-operator-in-oke-extending-gpu-control-beyond-scheduling 2026-08-25Google Connects Gemini Finance Agents to D&B Data, Adding 50-Plus Skills for Regulated Work Agent and coding products 72 1 /posts/google-connects-gemini-finance-agents-to-d-and-b-data-adding-50-plus-skills-for-regulated-work 2026-08-25MIT Study Finds Chatbot Help Can Erode Fake-News Detection Without AI Capability evaluations culture 58 1 /posts/mit-study-finds-chatbot-help-can-erode-fake-news-detection-without-ai 2026-08-25OpenAI’s Jalapeño Claims Up to 1.9x Efficiency Gain, but Leaves Nvidia Training Intact Leaderboards products 81 2 /posts/openai-plans-jalape-o-rollout-by-year-end-keeps-nvidia-broadly-deployed 2026-08-26Liquid AI’s Pipette Tests 1,000+ On-Device AI Setups—and Limits Cross-Device Rankings Agent and coding tools 67 1 /posts/liquid-ai-s-pipette-tests-1-000-on-device-ai-setups-and-limits-cross-device-rankings 2026-08-26DeepSeek V4 Pro Lands on Fireworks With $2.50 CyberGym Solves, but Kimi K3 Scores Higher Safety evaluations models 68 1 /posts/deepseek-v4-pro-lands-on-fireworks-with-2-50-cybergym-solves-but-kimi-k3-scores-higher 2026-08-26Corti Launches a Governed AI Coding Layer on Denmark’s Gefion Supercomputer Agent and coding products 67 2 /posts/corti-and-dcai-launch-a-european-control-layer-for-enterprise-ai 2026-08-26AWS Publishes a Phone-Ordering AI Pattern That Connects Restaurant Data to Claude Agent and coding tools 58 1 /posts/aws-publishes-a-phone-ordering-ai-host-that-connects-claude-haiku-to-restaurant-tools 2026-08-26Amazon Puts $25B Standalone Marker on Chip Business While Remaining a Top Nvidia Customer Leaderboards business 72 1 /posts/amazon-says-its-chip-unit-would-top-25b-standalone-while-it-remains-an-nvidia-customer 2026-08-26Ora Adopts Vercel’s Eve After Its Benchmark Found Fewer Steps and Lower Costs Agent and coding products 57 1 /posts/ora-picks-vercel-s-eve-after-its-benchmark-found-fewer-steps-and-more-completed-website-tasks 2026-08-26LLMScholarBench Tests 22 Models and Finds an Accuracy-Representation Trade-Off Capability evaluations models 72 1 /posts/llmscholarbench-tests-22-models-and-finds-an-accuracy-diversity-trade-off 2026-08-26AIRSEAI Joins LF AI & Data to Target Cross-Platform Robotics Leaderboards tools 67 2 /posts/airseai-joins-lf-ai-and-data-to-target-cross-platform-robotics 2026-08-26Foxglove Adds Cosmos Data Search as Robot Builders Struggle for Reliable Work Leaderboards business 66 1 /posts/foxglove-adds-cosmos-data-search-as-robot-builders-struggle-for-reliable-work 2026-08-26Qwen Releases Qwen3.8-Flash With 1M-Token Option and One-Ninth Training-Cost Claim Agent and coding models 77 2 /posts/qwen-releases-qwen3-8-flash-with-1m-token-option-and-one-ninth-training-cost-claim 2026-08-26Nauta Lands BMW, Bosch, Hitachi and Yamaha Funding for Supply-Chain AI Agents Agent and coding startups 64 2 /posts/nauta-lands-bmw-bosch-hitachi-and-yamaha-funding-for-supply-chain-ai-agents 2026-08-26AWS Publishes Two AgentCore Paths to Query Cross-Account Knowledge Bases Agent and coding tools 62 1 /posts/aws-shows-agentcore-pattern-for-cross-account-data-queries-without-copying-data 2026-08-26GoDaddy Moves BI to Amazon Quick, Cuts Dashboards to Under 2,500 and Loads Under 5 Seconds Agent and coding products 64 1 /posts/godaddy-moves-bi-to-amazon-quick-cuts-dashboards-to-under-2-500-and-loads-under-5-seconds 2026-08-26Natera Moves Scheduling Voice Agent to Bedrock AgentCore; AWS Reports Sub-7-Second Latency Agent and coding tools 62 1 /posts/natera-moves-scheduling-voice-agent-to-bedrock-agentcore-aws-reports-sub-7-second-latency 2026-08-26Z.ai Releases 320B GLM-5.3-Flash With MIT Weights, 1M Context and Low API Rates Agent and coding models 85 1 /posts/z-ai-releases-320b-glm-5-3-flash-with-mit-weights-1m-context-and-low-api-rates 2026-08-26Hugging Face Explains DeepSeek’s Matched Shift From MoE Experts to Lookup Memory Capability evaluations models 68 0 /posts/deepseek-s-engram-puts-lookup-tables-beside-moe-experts-reporting-benchmark-gains 2026-08-26Anthropic Opens 750,000 Claude Conversations to Outside Study Without Showing the Chats Safety evaluations tools 75 0 /posts/anthropic-opens-750-000-claude-conversations-to-outside-study-without-showing-the-chats 2026-08-26OpenAI Details Agent Breach of Hugging Face, Halts Research Model and Tightens Controls Safety evaluations models 93 2 /posts/openai-details-agent-breach-of-hugging-face-halts-research-model-and-tightens-controls 2026-08-26AWS AgentCore Evaluations Uses OpenTelemetry to Score Agents Across Frameworks Agent and coding tools 72 1 /posts/aws-agentcore-evaluations-uses-opentelemetry-to-score-agents-across-frameworks 2026-08-26Estuary Makes Rust Runtime Default, With Exactly-Once Delivery Still Conditional Agent and coding tools 62 0 /posts/estuary-makes-rust-runtime-default-with-exactly-once-delivery-still-conditional 2026-08-26Lam Breaks Ground on Oregon Lab, First Step in $3B AI-Chip R&D Buildout Leaderboards business 68 1 /posts/lam-breaks-ground-on-oregon-lab-first-step-in-3b-ai-chip-r-and-d-buildout 2026-08-26Deep Cogito Raises $43M to Build AI Models Enterprises Can Own Safety evaluations startups 65 1 /posts/deep-cogito-raises-43m-to-build-ai-models-enterprises-can-own 2026-08-27Instinct Raises $250M at $2.5B for a Private-Beta Agent With Deep Access Agent and coding startups 69 1 /posts/instinct-raises-250m-at-2-5b-for-a-private-beta-agent-with-deep-access