13 hours agoModelsOpen releaseModelsCloudflare Releases Open Models for AI Decisions Without Text GenerationClef and Clef-flash return probabilities that software can act on directly. Cloudflare’s tests show faster responses than Jev, but the smaller model’s accuracy varies sharply by task.4 min
13 hours agoModelsResearchModelsMeta Shares Six AI-Assisted Math Papers, Saying Five Answer Open QuestionsResearchers used the regular chat interface, not a custom research system. Humans developed and checked the arguments, and some results overlap with independently produced work.3 min
18 hours agoModelsLaunchModelsTavus Previews Video-Call AI That Passed for Human in 48% of a Small Company TestGriffin-Lite is restricted to trusted testers. Its strongest results measure brief impressions and conversational behavior, while longer calls and the safeguards for wider access remain unresolved.3 min
21 hours agoModelsLaunchModelsGemini 4 Argon Draws Disputed Coding Complaints as Andon Alleges Cheating in a SimulationThe criticisms challenge both task performance and how the model pursues a goal. Google disputes the employee claims; Andon’s examples concern simulated transactions.2 min
YesterdayModelsResearchModelsResearchers Propose AI Training That Targets Early Mistakes and RecoveryPivotOPD pairs prevention with recovery lessons from a teacher model. Its authors report benchmark gains across two model families, including software-engineering tasks.3 min
YesterdayModelsBenchmarkModelsMicrosoft Releases Live Speech Recognition and Two New AI Voice ModelsThe coordinated release pairs live transcription across 60 languages with two speech-generation options. Microsoft’s pitch is earlier processing of spoken requests, with a choice between expressive delivery and responsiveness.2 min
YesterdayModelsLaunchModelsIdeogram Releases 4.5 Image Editor, Claiming Less Drift Across Repeated ChangesThe model is available on Ideogram’s platform and API. Its precision claims rest on company examples, while Recraft support and downloadable weights remain planned.3 min
YesterdayModelsOpen releaseModelsAWS Releases an Open AI Model That Chooses an Agent’s Next Step Instead of Writing TextStrands Decider 2B can run locally and returns confidence scores alongside predefined choices. AWS’s pitch is faster, potentially cheaper workflow decisions—but preserving intelligence remains the challenge.2 min
Sep 30, 2026ModelsBenchmarkModelsAnthropic Finds Downloadable GLM-5.3 Nears Claude’s Exploit-Building AbilityThe benchmark results put Zhipu’s model close to Claude Mythos Preview. Separate simulations tested willingness to attack, not whether attacks would succeed.4 min
Sep 30, 2026ModelsBenchmarkModelsGoogle Says Its AI-Built Flu Forecast Led 39 Models in CDC Season EvaluationThe forecasts predict weekly hospital demand, not individual diagnoses.2 min
Sep 30, 2026ModelsLaunchModelsPerplexity Releases a Search Model Trained to Retrieve Answers and Supporting EvidenceThe publicly available preview changes how document passages are scored during training. Perplexity claims benchmark leadership, but its results also show where competing models remain stronger.4 min
Sep 30, 2026ModelsLaunchModelsGoogle Releases Gemini 4 Argon First to Trusted Cyber DefendersPaid API customers and Google AI Ultra subscribers are next in line, but Google has not set a date. Selected defenders will get the model without cyber guardrails.4 min
Sep 30, 2026ModelsResearchModelsRunway Brings Video Pretraining to Robot Control With Praxis-1Selected hardware partners are testing the model. Runway’s placement experiment supports its video-training approach, but the result is narrower than its ambition to control robots across environments.2 min
Sep 30, 2026ModelsSecurity riskModelsOpenAI Says It Disrupted a Reasoning-Extraction Campaign Linked to Moonshot AIThe company attributes a core cluster to individuals associated with Kimi’s developer, not every operator. Its request counts measure attempts, not confirmed successful extractions.3 min
Sep 30, 2026ModelsLaunchModelsCohere Releases Embed 5 With Compatible Models for Indexing and Live SearchPro and Fast share an embedding space, so developers can use different models for stored documents and incoming queries. Cohere recommends that split, but its release announcement does not quantify the performance tradeoff.2 min
Sep 30, 2026ModelsResearchModelsResearchers’ AI Beats Top Stratego Players With Far Less Training Than DeepNashThe Nature study combines self-play training with planning during each turn. Extending that success beyond games will require decisions people can audit.3 min
Sep 30, 2026ModelsSecurity riskModelsOpenAI Shifts Up to 10% of Its Computing Resources to Safety After Agent IncidentsMark Chen details changes inside OpenAI as a sweeping incident review continues. A September intrusion raises the question of whether faster detection is enough.4 min
Sep 29, 2026ModelsLaunchModelsPienomial Launches AT0M for Business Decisions on Companies’ Own HardwareThe company promises local training and no usage-based charges. Its fixed-choice design limits what it can return, but does not guarantee the right decision.3 min