8 hours agoModelsResearchModelsMostik Links AI Models Through Their Weights, Claiming a 20x Cost CutThe startup is proposing an alternative to the usual model-to-model text handoff. Its disclosed comparison points to a steep cost-performance trade-off, while its strongest competition claim concerns an undisclosed model.2 min
9 hours agoModelsResearchModelsGoogle’s RT-2 Made Language Models Robot Controllers, Defining a VLA TemplateThe 2023 system joined web-trained visual knowledge to physical commands, a design that has shaped modern robot research. Its reliance on teleoperation data and remote computing shows how far the path to capable machines still runs.3 min
9 hours agoModelsResearchModelsMIT and Motional Build AI That Exposes Robotaxi Planning ErrorsCW-Net puts readable concepts into the final driving decision itself, aiming to show whether a vehicle’s main planner or a safety backstop caused a maneuver.2 min
10 hours agoModelsResearchModelsNature Study Finds AI Agent Teams Can Cut Sequential Planning Performance by 70%The controlled experiment offers a bounded way to choose an agent architecture: assess the task and a single agent’s baseline before paying the cost of coordination.2 min
10 hours agoModelsResearchModelsAnthropic’s 80-Environment Reward-Hacking Test Produced Cyber and Safety EvasionsThe controlled experiment does not measure deployed-model behavior. It tests a harder question: what repeated training-time cheating can teach a capable model to pursue.3 min
YesterdayModelsResearchModelsRunway’s Solaris Generates App Interfaces Frame by Frame, but It Isn’t Ready for Public UseThe experimental system shifts the interface from a prebuilt program to continuously generated pixels. That could create more adaptive software, but it also leaves text, accuracy, long sessions and accessibility unresolved.3 min
Aug 31, 2026ModelsResearchModelsAI ECG Tool Flags Heart Risk in Two Seconds, but Still Needs the Scan That Diagnoses ItThe reported results suggest routine electrical heart traces could help prioritize scarce ultrasound appointments. The system’s stated role, however, stops short of confirming disease or ruling it out.2 min
Aug 28, 2026ModelsResearchModelsAnthropic Says Claude Found Fixes Across 10 Alignment Failures, but Tests Remain NarrowThe release turns safety post-training into a repeatable model-run search process. Its value now depends on whether those benchmark gains survive broader tests and later training.3 min
Aug 28, 2026ModelsResearchModelsJAMA Paper Says Autonomous AI Could Surpass Physicians and AI-Assisted Care by 2030The contested forecast recasts clinical AI from a physician tool into a possible replacement, while real-world patient use and medical training remain unsettled.2 min
Aug 27, 2026ModelsResearchModelsNetflix’s GenRec Pairs a 0.006% Gain With Catalog-Bound LLM RankingThe online result is statistically significant but too opaque to measure as a product outcome. The clearer contribution is a recommendation architecture that constrains an LLM to available titles while managing inference cost.3 min
Aug 27, 2026ModelsResearchModelsMIT Builds PottsMPNN to Model Protein Stability Beyond Native-Sequence MatchingThe framework centers protein design on whether a sequence fits a target structure and its energy landscape, rather than whether it resembles the sequence evolution selected.2 min
Aug 27, 2026ModelsResearchModelsOpenAI’s Astra Is Claimed to Solve Non-Sofic Groups Problem, Raising Stakes for Human MathematiciansHenry Bradford’s account describes a proof built from existing theorems, while arguing that AI’s progress could reshape how universities value mathematical research.2 min
Aug 27, 2026ModelsResearchModelsOpenAI Tests Persistent Codex Mode That Keeps Working Until Put to SleepThe unannounced Codex experiment would let an agent create its own follow-up work across sessions and contact users sparingly, while OpenAI’s safety findings show why longer-running behavior needs firm limits.3 min
Aug 27, 2026ModelsResearchModelsMammogram AI Found Prior Stroke at 86%, but Clinical Use Still Needs ValidationThe study points to a possible way to extract cardiovascular signals from breast scans already taken in routine care. Its reported results are promising, but the path to clinical use still runs through accuracy, reliability and false-result reduction.2 min
Aug 26, 2026ModelsResearchModelsAltman Targets Internal AGI by End of 2026, as Astra Tests OpenAI’s DefinitionThe target depends on an economics-based definition of general intelligence and company-described research performance, while the field still lacks a shared technical finish line.2 min
Aug 26, 2026ModelsResearchModelsHugging Face Explains DeepSeek’s Matched Shift From MoE Experts to Lookup MemoryThe fresh explainer makes a January research design easier to parse: DeepSeek reported benchmark gains after reallocating capacity, not after releasing a new trained Engram model or confirming a deployment.3 min
Aug 26, 2026ModelsResearchModelsMIT’s CrysVCD Steers AI Material Design Toward Stability Before Costly ScreeningThe framework shifts chemical validation to the front of the workflow, aiming to leave smaller labs with more viable crystal candidates instead of a large downstream filtering bill.2 min
Aug 25, 2026ModelsResearchModelsRadiology Holds Three-Quarters of Cleared Medical AI—and a New Veto ProblemThe central challenge is no longer whether image-reading systems can find abnormalities. It is whether clinicians can recognize the rare cases where a highly accurate system is wrong without ignoring the cases where it sees something they miss.4 min