Sep 1, 2026ModelsBenchmarkModelsGrok 4.6 Clears 50% on Biosecurity Refusal and Routine Biology TasksThe reported result rewards a difficult balance: recognizing hazards hidden inside ordinary-looking research requests without shutting down legitimate biological work. Its surveillance score shows that refusal performance is not the whole biosecurity test.3 min
Sep 1, 2026ModelsOpen releaseModelsOrbis Lets Users Alter Prompts Mid-Stream as Visko Takes Its Video Model PublicVisko’s product bet is not simply faster generation. It is that a video session can keep its state while a user changes its direction—a claim public access can now put under practical scrutiny.2 min
Sep 1, 2026ModelsLaunchModelsFlower Labs Starts Endeavor 1.0 Preview With a Private Deployment OptionThe model’s differentiator is not just Flower’s frontier-performance claim. Customers can run it as a managed service or place selected workloads inside their own infrastructure, though the launch remains limited to early users.2 min
Aug 31, 2026ModelsBenchmarkModelsQwen3.8-Max Leads Open-Weight Commerce Agent Bench by Two Passes, With Audit LimitsThe benchmark checks whether agents leave commerce software and files in the required state, rather than simply narrate a workflow. But configuration choices and unavailable immutable result bundles limit what the narrow ranking can settle.2 min
Aug 31, 2026ModelsResearchModelsRunway’s Solaris Generates App Interfaces Frame by Frame, but It Isn’t Ready for Public UseThe experimental system shifts the interface from a prebuilt program to continuously generated pixels. That could create more adaptive software, but it also leaves text, accuracy, long sessions and accessibility unresolved.3 min
Aug 31, 2026ModelsLaunchModelsSkild AI’s S1 Uses One Human Video to Prompt Robots Through 10-Minute TasksThe model’s central bet is that a demonstration can become immediate instruction rather than another task-specific training run. The harder question is whether that capability can turn into dependable deployments.3 min
Aug 31, 2026ModelsPartnershipModelsGoogle DeepMind’s Farm AI Reaches Six African Countries as FAO Plans Data IntegrationThe expansion puts Google DeepMind’s agricultural insights in Kenya, Uganda, Ghana, Rwanda, Nigeria and Zambia. A planned FAO integration would test whether those outputs can help countries update crop maps and agricultural statistics faster.2 min
Aug 31, 2026ModelsOpen releaseModelsDeepSeek Releases 305B Vision Weights Under MIT License, but Local Use Needs Serious HardwareThe release gives developers the model files and serving references needed to run or adapt DeepSeek-V4-Flash-Vision-Exp independently. Its 168GB size means that freedom is aimed more at infrastructure-equipped teams than typical individual users.3 min
Aug 31, 2026ModelsLaunchModelsGoogle’s TimesFM-3 Adds Known Future Events to Zero-Shot ForecastingThe 330-million-parameter model can combine related data streams with inputs such as promotion calendars and weather forecasts. Google reports leading results on three public benchmarks; performance on an organization’s own data remains the practical open question.3 min
Aug 31, 2026ModelsOpen releaseModelsMicrosoft Releases GigaPath-Flash, Reporting Roughly 50x Less Compute for Cancer ResearchThe open-weight research release is designed for repeated analyses across large cancer cohorts, while its performance claims remain limited to initial benchmarks and cohorts.2 min
Aug 31, 2026ModelsResearchModelsAI ECG Tool Flags Heart Risk in Two Seconds, but Still Needs the Scan That Diagnoses ItThe reported results suggest routine electrical heart traces could help prioritize scarce ultrasound appointments. The system’s stated role, however, stops short of confirming disease or ruling it out.2 min
Aug 31, 2026ModelsBenchmarkModelsChatGPT o3 Handles Simulated Liquidity Choices, but Consistency Slips in Harder CasesThe simulated wholesale-payments exercise defines a narrow starting point for delegated cash management: routine payment priorities, with human oversight for potentially anomalous patterns.2 min
Aug 30, 2026ModelsBenchmarkModelsHack The Box Benchmark: Best Human Team Cleared 36 Challenges; Top AI Team Reached 32The latest benchmark points to a useful division of labor in cyber work: agents can accelerate solving, but elite performance still depends on people choosing paths and checking results.3 min
Aug 30, 2026ModelsBenchmarkModelsChatbots Debunked Foreign Falsehoods About 75% of the Time, Beating Search ResultsThe result makes chatbots a potentially stronger starting point for investigating state-spread claims than a search-results page. It does not make their answers self-validating: language, citations and product design still shape the outcome.4 min
Aug 29, 2026ModelsBenchmarkModelsGoogle’s WikiSkill Lifts Agent Benchmarks by Giving Models a Memory of Failed WorkThe framework turns task traces into reusable instructions instead of changing a model’s training, offering a practical route to more capable agents while leaving uneven task gains and cross-model transfer as constraints.3 min
Aug 28, 2026ModelsOpen releaseModelsGnani Launches 30B Evon Model and Self-Hosted Agent Stack for Indian InstitutionsThe open-weight release couples Indian-language support with tools intended to keep sensitive data inside a customer’s own infrastructure.3 min
Aug 28, 2026ModelsResearchModelsAnthropic Says Claude Found Fixes Across 10 Alignment Failures, but Tests Remain NarrowThe release turns safety post-training into a repeatable model-run search process. Its value now depends on whether those benchmark gains survive broader tests and later training.3 min
Aug 28, 2026ModelsEnterprise adoptionModelsDecathlon Deploys Chronos-2, Cuts 12-Week WAPE by 11–15 PointsThe retailer’s comparison with a weekly-retrained TFT system makes the operational case for fine-tuned forecasting models. The published production results, however, come from Southeast Asia and Latin America.4 min