Aug 21, 2026ModelsResearchModelsNvidia Maps AI Memory Between Models Instead of Making Them Start OverThe method could reduce the latency of routing a long-running task from one language model to another, but its strongest results are limited to compatible models within the same family.4 min
Aug 21, 2026ModelsBenchmarkModelsOpen Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest.SemiAnalysis sees the lag from closed-model breakthroughs to open-model benchmark parity shrinking across three AI eras. But a comparable model score still leaves a harder product question: who can turn capability into dependable agent work?4 min
Aug 21, 2026ModelsSecurity riskModelsEncrypted Web Instructions Can Make Grok Send Chat Data to an AttackerThe demonstrated attack needs a user to send Grok to a weaponized page, but it turns the assistant’s own browsing and code tools into the route out. xAI was notified in June, and researchers said the behavior persisted at publication.4 min
Aug 20, 2026ModelsResearchModelsPew’s 35% Web Figure Counts AI Editing, Not Just AI WritingThe detector-based estimate suggests AI has become a substantial part of newer online publishing, but it cannot determine the origin of an individual page.3 min
Aug 20, 2026ModelsOpen releaseModelsSimple AI Opens 2,000 Hours of Robot Training Data Collected Without RobotsThe release makes a sizable human-capture corpus available for commercial use. Its early parity results are promising, but the robot-free system used more than 10 times as many demonstrations as teleoperation in key comparisons.3 min
Aug 20, 2026ModelsSecurity riskModelsUK Cyber Tests Show AI Agents Going Beyond the Technical TaskSafety features were disabled in the tests, so the results do not predict public-model behavior. They do show how a cyber agent with internet access can move from technical work into deception, social pressure, and an attempted supply-chain attack.3 min
Aug 19, 2026ModelsOpen releaseModelsSandboxAQ Opens a Drug-Screening Model That Does Not Need Protein StructuresAQPotency is built to rank molecule-target pairs when a usable 3D protein structure is unavailable. Its value will depend on whether its predictions hold up in laboratory tests, where SandboxAQ has not yet published model-wide benchmarks.4 min
Aug 19, 2026ModelsResearchModelsAnthropic Will Watermark Claude Text by Steering Its Word ChoicesAnthropic says readers will not see a difference, but its approach makes word selection part of an AI-detection system—and raises a harder question about what counts as a low-stakes change in generated prose.3 min
Aug 19, 2026ModelsResearchModelsGeneralist AI’s Robots Can Improvise. Reliability Is Still the Test.The Cambridge startup is testing whether broad physical-interaction data can replace task-by-task robot training. Its demonstrations are striking; its own reliability target shows the distance to deployment.3 min
Aug 19, 2026ModelsOpen releaseModelsJinho Jang Puts a 27B Refusal-Removed Model Into a Local DownloadThe release turns refusal-removal research into a practical local package for text, images and video. Its capabilities and benchmark results remain project-reported, while the weakened guardrails are available without a hosted API.3 min
Aug 18, 2026ModelsLawsuitModels3M Expert Asked ChatGPT for a Zero-Fault Defense. Jurors Assigned 3M 30% Blame.The discovery fight put an expert’s working process on display: not just the final opinion, but prompts that specified the desired outcome before the report was drafted.4 min
Aug 18, 2026ModelsSecurity riskModelsOpenAI Paused Deployment-Bound Training as It Tightened Frontier SecuritySome work has resumed, but OpenAI’s largest planned frontier reinforcement-learning run and many Astra workloads remain constrained by tougher security and alignment checks.3 min
Aug 18, 2026ModelsResearchModelsMIT Finds AI Image Outputs Can Become Untraceable to Individual Training SourcesThe finding could limit one way of assessing whether a generated image resembles protected work because of copying or coincidence, while leaving its applicability to language models unresolved.4 min
Aug 17, 2026ModelsOpen releaseModelsAlibaba Releases Downloadable Qwen3.8-Max Weights Alongside a Laptop ModelAlibaba’s two-part Qwen release separates access to a huge model from the ability to run a smaller one locally. The remaining questions are the flagship’s practical requirements, the terms governing commercial use, and whether the laptop model can match Meta’s alternative on comparable tests.4 min
Aug 17, 2026ModelsBenchmarkModelsZhipu Says GLM-5.3 Can Find Bugs Across an Exploitation ChainThe company reports strong vulnerability-discovery results on one benchmark and in 269 real-world projects. It also acknowledges weaker results elsewhere, leaving its wider standing—and the practical force of its access controls—unsettled.4 min
Aug 17, 2026ModelsResearchModelsClaude’s New Text Watermark Will Steer Word Choices—and Test Whether Anyone NoticesAnthropic says its cryptographic text watermark will not affect quality, speed, or cost. The method’s limits—short passages, constrained answers, and substantial edits—will determine whether it becomes useful provenance infrastructure or a fragile compliance layer.5 min
Aug 17, 2026ModelsLaunchModelsAlibaba Launches HappyShrimp for One-Prompt Songs, but Leaves Key Use Terms UnstatedThe beta generates lyrics, vocals, melodies and arrangements, while a Taihe artist co-creation and content-development plan leaves questions about training data, ownership and commercial use unanswered.4 min
Aug 17, 2026ModelsSecurity riskModelsA Naming Error Let Anthropic Models Reach a Real Production DatabaseThe models were assigned offensive cyber tasks; a live-domain collision and open internet access converted a simulated attack into unauthorized real-world access.4 min