Sep 20, 2026ModelsLaunchModelsAnthropic Weighs New AI Model as OpenAI’s Astra Gains GroundReuters-reported deliberations leave Anthropic balancing competitive urgency against safety review and the cost of turning another frontier model into a product.2 min
Sep 20, 2026ModelsBenchmarkModelsBrood War Bench Finds AI Agents Can Win Matches but Not Master StrategyCodex Astra / xhigh went 18–0 in the new agent benchmark, yet its creator says no tested system moved beyond beginner-level play or reliably coordinated the work needed to build a lasting advantage.3 min
Sep 20, 2026ModelsResearchModelsUC San Diego Publishes Virtual Cells That Pair AI With PhysicsThe paired Cell studies use time-resolved microscopy for two jobs: mapping patterns across drug-treated cells and testing whether a simulated cell can reproduce a drug response.2 min
Sep 19, 2026ModelsBenchmarkModelsLangChain Publishes Jev Test Matching Human Labels on 500 Agent DecisionsThe result points to a cheaper, more repeatable way to grade agent behavior, but it comes from five fixed weather-agent runs—not a production test.4 min
Sep 19, 2026ModelsLaunchModelsFigure Tests Helix 2.5 in 30 Unseen Homes, Reports 56% Zero-Shot SuccessThe humanoid test is a notable attempt to measure whether robot skills transfer between real homes. But its reported 56% end-to-end success rate also shows how far dependable household automation remains from settled.3 min
Sep 19, 2026ModelsLaunchModelsAlibaba Releases Qwen3.8-Omni-Flash for Tool-Using Media TasksThe new hosted model is meant to turn long recordings into inputs for AI workflows. Its low listed API rates are clear; its effectiveness on complex, multi-step media jobs remains to be tested.2 min
Sep 19, 2026ModelsBenchmarkModelsAgent Memory Challenge Will Open a Shared Test for Long-Running AI AgentsThe second cycle moves beyond storing past context, testing whether memory systems can surface useful, current evidence across conversations, code and images under one evaluation protocol.3 min
Sep 18, 2026ModelsSecurity riskModelsGoogle Says Gemini Reached Three Company Networks During a Cybersecurity TestA target-name collision and unintended internet access moved an offensive-security evaluation beyond its simulated setting. Google says Gemini stopped after recognizing the systems were real and caused no harm.2 min
Sep 18, 2026ModelsOpen releaseModelsJina AI Releases Document Parser, Claims 2.57 Pages Per SecondThe open-weight model combines speculative decoding with API access, while its reported results point to damaged scans as a practical limitation.3 min
Sep 18, 2026ModelsLaunchModelsZ.AI Adds GLM-5.3-Flash to Coding Plan but Leaves FlashX Off ItThe release pairs multimodal, long-context capabilities with two distinct access options: a higher-quota plan model and a faster API variant.2 min
Sep 18, 2026ModelsLaunchModelsMeta Adds SAM 3.1 to Its API for Object Detection, Segmentation and TrackingThe hosted model combines detection, pixel-level masks and identity-preserving video tracks in one request, aiming to remove the need to assemble and tune separate vision systems.3 min
Sep 18, 2026ModelsLaunchModelsSpaceXAI Releases Grok Voice Transcribe 2.0 at Its Existing API PricesThe new batch and live-transcription model is available now, while SpaceXAI’s strongest accuracy results come from its own testing and version 1.0 remains the current default.2 min
Sep 17, 2026ModelsLaunchModelsOpenRouter’s Union Alpha Goes Paid After Demand Overwhelms CapacityThe anonymous coding model’s sudden pricing shift turns a free stress test into a paid service, while its identity and a Cloudflare route give developers more context for deciding whether to use it.2 min
Sep 17, 2026ModelsModelsOpenAI May Delay a Reported Hodge Conjecture Result After Math BacklashEmployees reportedly expect a breakthrough soon, but the company is said to be weighing how to present it after its Navier–Stokes claim drew criticism over credit and verification.3 min
Sep 17, 2026ModelsBenchmarkModelsScale AI and Korea AI Safety Institute Release Benchmark Exposing a Translation GapThe study found lower harmful-response scores in Korean prompts, but also more refusals of harmless requests—complicating claims that a model is simply safer in another language.3 min
Sep 17, 2026ModelsLaunchModelsPrismML Shrinks Qwen3.8 27B to 5.9 GB for Local AIThe open-weight model aims to bring larger reasoning and multimodal workloads to local hardware, but PrismML’s own results show the compression tradeoff varies by task.2 min
Sep 17, 2026ModelsOpen releaseModelsAnthropic Opens Claude-Built Biology Tool Optimizations It Says Run 4x FasterThe release targets a stubborn bottleneck in protein and molecular-structure software. Anthropic’s reported gains are promising, but the largest experimental runs still failed to produce correct structures.3 min
Sep 17, 2026ModelsResearchModelsLasso Security Finds Text Watermarking Can Change AI Agent ActionsThe study does not test Claude’s planned implementation. But it finds that a provenance feature can alter tool use and prompt-injection resistance in open-weight models.3 min