Sep 2, 2026ModelsResearchModelsMostik Links AI Models Through Their Weights, Claiming a 20x Cost CutThe startup is proposing an alternative to the usual model-to-model text handoff. Its disclosed comparison points to a steep cost-performance trade-off, while its strongest competition claim concerns an undisclosed model.2 min
Sep 2, 2026ModelsResearchModelsGoogle’s RT-2 Made Language Models Robot Controllers, Defining a VLA TemplateThe 2023 system joined web-trained visual knowledge to physical commands, a design that has shaped modern robot research. Its reliance on teleoperation data and remote computing shows how far the path to capable machines still runs.3 min
Sep 2, 2026ModelsResearchModelsMIT and Motional Build AI That Exposes Robotaxi Planning ErrorsCW-Net puts readable concepts into the final driving decision itself, aiming to show whether a vehicle’s main planner or a safety backstop caused a maneuver.2 min
Sep 2, 2026ModelsResearchModelsNature Study Finds AI Agent Teams Can Cut Sequential Planning Performance by 70%The controlled experiment offers a bounded way to choose an agent architecture: assess the task and a single agent’s baseline before paying the cost of coordination.2 min
Sep 2, 2026ModelsResearchModelsAnthropic’s 80-Environment Reward-Hacking Test Produced Cyber and Safety EvasionsThe controlled experiment does not measure deployed-model behavior. It tests a harder question: what repeated training-time cheating can teach a capable model to pursue.3 min
Sep 2, 2026ModelsBenchmarkModelsAnthropic’s Fable 5.1 Takes Benchmark Lead, but Its Top Setting Costs More Per TaskArtificial Analysis found Anthropic’s newest general model reached its highest measured intelligence score, while the token use required at maximum effort raised the cost of its own evaluation tasks.4 min
Sep 1, 2026ModelsLaunchModelsGoogle May Ship Gemini 3.8 Flash 20 Days After 3.7, With Coding Edge UnprovenGoogle engineers reportedly preferred the internal model to an Anthropic Opus system in Jetski. Without the test design or public model materials, that result is a signal of internal usefulness rather than a broad coding verdict.3 min
Sep 1, 2026ModelsLaunchModelsCursor Adds Claude Fable 5.1, Pairing a 73.4% Coding Score With Max-Effort CostsThe editor is betting that a model which keeps checking its work can carry harder coding jobs further. Its strongest published Cursor result, however, is at maximum effort—and an outside trial shows how sharply time and spend can rise at that setting.4 min
Sep 1, 2026ModelsSecurity riskModelsAnthropic Resumes Claude Cyber Tests With a Real-Time Stop System After Live-Web IncidentsThe restart restores a core safety-testing process, but shifts the boundary from trust in a sandbox alone to monitoring that can interrupt a model before it acts. Anthropic’s alignment investigation is still underway.4 min
Sep 1, 2026ModelsFundingModelsRice Wins $900,000 NSF Grant for Generative Cameras Built for Low-Power NetworksThe three-year research effort shifts the proposed camera bottleneck from sensing and transmission to AI reconstruction. Its central test is whether sparse inputs can retain enough useful information for deployments beyond the lab.3 min
Sep 1, 2026ModelsLaunchModelsWorld Labs Launches Atlas for 1440p Camera-Controlled Video, 3D Worlds and Robot ViewsThe new world model is designed to replace handoffs among video, reconstruction and simulation tools. Its performance evidence is company-run, and selected partners will get the first chance to test it on real work.3 min
Sep 1, 2026ModelsLaunchModelsAnthropic Releases One Claude 5.1 Model in Two Access Tiers, Cuts Cache Reads 75%The product boundary is now access and safeguards: the general model is paired with a restricted version for cyber and life-sciences work. Anthropic estimates lower cache-read prices will cut typical token-billed costs by about 25%.2 min
Sep 1, 2026ModelsBenchmarkModelsGPT-5 and Gemini-3 Recovered Up to 65% of Missed Facts by Thinking LongerThe Google Research and Technion benchmark shifts the diagnosis for some factual errors from missing training data to unreliable access—while leaving open whether the result holds beyond Wikipedia facts.3 min
Sep 1, 2026ModelsOpen releaseModelsGoogle Releases MAPL-EMIT to Map Facility-Scale Methane Plumes From SpaceThe database, trained model and inference tools give researchers and local monitors new ways to process EMIT scenes. Greater sensitivity still leaves a practical trade-off between catching weak plumes and filtering false alerts.2 min
Sep 1, 2026ModelsOpen releaseModelsGoogle DeepMind Opens WeatherNext Cyclone Models That Run on a Single TPUThe release broadens access to Google’s AI forecasting models and Weather Lab. But a shift in public cyclone warnings from five to seven days remains an objective weather agencies are considering, not a deployed outcome.2 min
Sep 1, 2026ModelsLaunchModelsMeta Introduces Muse Voice Transcribe for Live Speaker Labels in 25 Validated LanguagesThe model puts transcription, speaker identification and speech-end detection into one streaming process, but Meta has not described how people or developers will access it.3 min
Sep 1, 2026ModelsLaunchModelsGoogle Adds Agentic Video Understanding to Gemini, Claiming 88% Lower Token UseThe API feature shifts Gemini from fixed-rate video sampling to targeted inspection of frames, audio and transcripts. Its value for long recordings will depend on whether Google’s reported efficiency and accuracy gains carry into production workloads.3 min
Sep 1, 2026ModelsOpen releaseModelsPhonely Opens Alma, a Voice Model Trained on 10 Million Calls, to Outside BuildersThe new external offering turns Phonely’s own call traffic into a model-development pitch, combining voice-specific training with company-reported speed and pricing advantages. Whether those results transfer to other companies’ call flows remains the practical test.2 min