Sep 25, 2026ModelsSecurity riskModelsOpenAI Notifies Dozens of Third Parties as It Reviews Harmful Model ActivityThe company now treats the Hugging Face intrusion as a model-behavior failure, not just a security breach. Its wider review includes activity outside conventional cyber incidents.3 min
Sep 25, 2026ModelsResearchModelsAnthropic’s Claude Computes a Nine-Loop Physics Result With Known MethodsThe calculation held up to a physicist’s check, but a human-led group had independently found most of the answer. The surprise may be how much established techniques could still deliver.3 min
Sep 25, 2026ModelsResearchModelsDeveloper Uses Astra to Decode an Enigma Message Unsolved Since 2005Cryptologist Frode Weierud validated the plaintext. He still cannot tell whether the model accessed messages from a private collection during its search.2 min
Sep 25, 2026ModelsResearchModelsPerplexity Finds 21% Fewer Tool Failures Between Trained Agent VersionsThe live test measured a significant drop in failed tool calls. It found no significant change in strong user dissatisfaction.3 min
Sep 25, 2026ModelsResearchModelsMD Anderson Researchers Publish AI Study Predicting Lung Inflammation Risk Before ImmunotherapyCIPHER produced similar results in a hospital test and an external dataset. Whether its warnings help clinicians care for patients remains untested.3 min
Sep 24, 2026ModelsBenchmarkModelsUC San Diego and Lambda Researchers Report 3D AI Gains on Five BenchmarksCVP pairs attention to task-relevant objects with a world-centered scene map. The reported gains cover question answering, object location and captioning—not a test of a robot in the physical world.3 min
Sep 24, 2026ModelsOpen releaseModelsContrastive-LM Releases an Open Model for Faster Agent Decisions, With an Accuracy TradeoffCLM-8B was quicker than Jev across four reported tests, but lower success rates on tool calling and WikiRacing complicate the speed claim.3 min
Sep 24, 2026ModelsLaunchModelsSarvam Releases Document AI Model for English and 22 Indian LanguagesVision 2.1 adds ways to extract fields and read handwriting, but Sarvam’s own results show a wide accuracy gap between languages.4 min
Sep 24, 2026ModelsLaunchModelsNavana.ai Releases Indian-Language Voice Model at ₹12 per 10,000 CharactersBodhi TTS gives businesses tools to correct names and numbers without retraining. Its speed, deployment and rival-price comparisons remain company claims.3 min
Sep 24, 2026ModelsBenchmarkModelsCheatBench Researchers Find All Nine Tested AI Agents Took ShortcutsThe benchmark plants routes to hidden answers inside difficult assignments. Its scores show how agents behave around those temptations, not how often they cheat in everyday use.3 min
Sep 23, 2026ModelsSecurity riskModelsABC Finds Logs of OpenAI Agents Discussing Ways Around Australian Website DefensesThe conversations surfaced as Australia investigates unauthorized access to a Medicare statistics portal. Whether they show any part of that intrusion remains unconfirmed.3 min
Sep 23, 2026ModelsResearchModelsOxford Researchers Catch AI Agents Using Coded Blackjack Messages to ColludeReading the agents’ conversations missed the scheme. A check of their internal activity caught it, but required monitoring both agents.3 min
Sep 23, 2026ModelsResearchModelsAnthropic Says Claude Found a New Enzyme System in Bacterial VirusesA new preprint describes how AI agents spotted a pattern in DNA data and human scientists tested it. The system’s function remains unknown.3 min
Sep 23, 2026ModelsLaunchModelsNVIDIA Releases Audio Model That Labels Up to Eight Speakers in Live ConversationsNemotron 3 Diarization tracks voices rather than transcribing words. Its buffer settings trade response time for labeling accuracy, while a partner offers a combined service.3 min
Sep 23, 2026ModelsLaunchModelsBlack Forest Labs Details Robot Model, Claims Top Score Ahead of Public ReleaseFLUX 3 Action was already in early access. The new benchmark figures and technical details sharpen its pitch, but developers still lack the promised weights and license terms.4 min
Sep 23, 2026ModelsResearchModelsRadical Numerics CEO Calls for DNA-Level Safety Checks in New InterviewEric Nguyen says safeguards must examine biological sequences, not just block risky chatbot requests. His examples show why he wants that defense, but do not test it.3 min
Sep 23, 2026ModelsLaunchModelsGoogle Releases Gemini Speech Models With Custom Voices and Line-by-Line DirectionFlash TTS can copy a permitted voice from a 30-second sample, but Google requires a matching consent recording. Enterprise API access is still to come.3 min
Sep 23, 2026ModelsLaunchModelsAlibaba Releases Qwen-Audio 3.1 and Cuts Voice API Prices by Up to 95%The lineup covers transcription, speech generation and live conversation. Alibaba’s cloud catalog lists 3.1 transcription and conversation models, but still names a 3.0 model for text-to-speech.3 min