13 hours agoModelsBenchmarkModelsAnthropic’s Fable 5.1 Takes Benchmark Lead, but Its Top Setting Costs More Per TaskArtificial Analysis found Anthropic’s newest general model reached its highest measured intelligence score, while the token use required at maximum effort raised the cost of its own evaluation tasks.4 min
YesterdayModelsBenchmarkModelsGPT-5 and Gemini-3 Recovered Up to 65% of Missed Facts by Thinking LongerThe Google Research and Technion benchmark shifts the diagnosis for some factual errors from missing training data to unreliable access—while leaving open whether the result holds beyond Wikipedia facts.3 min
YesterdayModelsBenchmarkModelsGrok 4.6 Clears 50% on Biosecurity Refusal and Routine Biology TasksThe reported result rewards a difficult balance: recognizing hazards hidden inside ordinary-looking research requests without shutting down legitimate biological work. Its surveillance score shows that refusal performance is not the whole biosecurity test.3 min
YesterdayModelsBenchmarkModelsQwen3.8-Max Leads Open-Weight Commerce Agent Bench by Two Passes, With Audit LimitsThe benchmark checks whether agents leave commerce software and files in the required state, rather than simply narrate a workflow. But configuration choices and unavailable immutable result bundles limit what the narrow ranking can settle.2 min
Aug 31, 2026ModelsBenchmarkModelsChatGPT o3 Handles Simulated Liquidity Choices, but Consistency Slips in Harder CasesThe simulated wholesale-payments exercise defines a narrow starting point for delegated cash management: routine payment priorities, with human oversight for potentially anomalous patterns.2 min
Aug 30, 2026ModelsBenchmarkModelsHack The Box Benchmark: Best Human Team Cleared 36 Challenges; Top AI Team Reached 32The latest benchmark points to a useful division of labor in cyber work: agents can accelerate solving, but elite performance still depends on people choosing paths and checking results.3 min
Aug 30, 2026ModelsBenchmarkModelsChatbots Debunked Foreign Falsehoods About 75% of the Time, Beating Search ResultsThe result makes chatbots a potentially stronger starting point for investigating state-spread claims than a search-results page. It does not make their answers self-validating: language, citations and product design still shape the outcome.4 min
Aug 29, 2026ModelsBenchmarkModelsGoogle’s WikiSkill Lifts Agent Benchmarks by Giving Models a Memory of Failed WorkThe framework turns task traces into reusable instructions instead of changing a model’s training, offering a practical route to more capable agents while leaving uneven task gains and cross-model transfer as constraints.3 min
Aug 28, 2026ModelsBenchmarkModelsMeta’s EvoHarness-RL Takes Qwen3-8B to 96.9% on ALFWorld, Near Claude Opus 4.5The result is a strong benchmark showing for learned state management, not yet evidence that the framework improves the long-running enterprise workflows it is designed to target.3 min
Aug 27, 2026ModelsBenchmarkModelsReViSQL-K2.6 Hits 91.37% on Cleaned SQL Test, but Not a Live DatabaseThe result argues that a focused model, repaired examples, and domain-specific rewards can rival elaborate SQL-agent stacks. Its remaining test is whether that recipe transfers from curated academic data to corporate databases.3 min
Aug 27, 2026ModelsBenchmarkModelsDatabricks Says 300M Chart-JSON Pipeline Tops Four Multimodal Baselines on Answer AccuracyThe company’s test suggests chart values can become a useful retrieval layer instead of relying solely on page-image embeddings, though its comparison measures answer correctness and includes a synthetic benchmark it built.3 min
Aug 27, 2026ModelsBenchmarkModelsGoogle DeepMind Puts Gemini Flash Lite in a Double-Blind Test to Guard Secret BenchmarksThe pilot replaces a longstanding choice between exposing confidential tests and exposing proprietary model weights. Its value now rests on whether cryptographic separation can make external testing more credible for sensitive use cases.3 min
Aug 26, 2026ModelsBenchmarkModelsGoogle and UNSW’s GlucoFM Beats a CGM Baseline by 4.1 PR-AUC Points, but Has No Public CheckpointThe result suggests that separating slow glucose patterns from short-lived deviations can improve reusable wearable-data representations. For now, it is a retrospective research result rather than a clinical product or downloadable model.3 min
Aug 25, 2026ModelsBenchmarkModelsLLMScholarBench Tests 22 Models and Finds an Accuracy-Representation Trade-OffThe benchmark treats AI-generated expert lists as recommender systems to audit, showing that factual correctness and balanced representation cannot yet be improved together reliably.3 min
Aug 25, 2026ModelsBenchmarkModelsHiDream-O1-World Tops WBench Navi at 80.9, Betting on Persistent 3D ScenesHiDream.ai’s new model is designed to make generated environments hold together through movement and edits. Its strongest disclosed evidence is a navigation benchmark; its larger film, robotics and production ambitions remain proposed uses.3 min
Aug 24, 2026ModelsBenchmarkModelsNVIDIA Says AVO Took Claude Opus 5 From 30% to Perfect on ARC-AGI-3The result shifts attention from training larger models to the software that gives them memory, tools and recovery loops. Its clearest limitation is equally important: the perfect score came on ARC-AGI-3’s public environments, while private sets remain the harder test.3 min
Aug 24, 2026ModelsBenchmarkModelsGPT-BERT Beats Llama 2 70B on One Grammar Test With 100 Million WordsThe narrow result does not make child-scale models competitive with frontier chatbots. It does show that massive pretraining corpora do not settle every language-learning test.3 min
Aug 22, 2026ModelsBenchmarkModelsInherent Says Its 27B Science Agent Beat OpenAI and Anthropic on Paper ReplicationThe company’s result centers on reproducing known findings, not making discoveries. Its next test is whether that training approach can move from replication to reliable new science.2 min