BenchmarkToolsLlamaIndex and Kaggle Launch a Document Benchmark for Cost and Long FilesExtractBench publishes a broader test for enterprise document extraction, while its launch results include company-reported figures for LlamaIndex’s own product.6 hours ago2 min readRead story
BenchmarkToolsLemmalog Turns Agent Memory Into Datalog State, Cutting LongMemEval Context 38-FoldThe experimental system moves long-running agent work away from recalling old notes and toward maintaining conclusions with explicit support. Its early benchmark results suggest a substantial context-saving path, while leaving the probabilistic extraction step as the central weakness.Aug 28, 20263 min readRead story
BenchmarkToolsPerplexity Search API Sweeps Artificial Analysis Test, but Medium Context WinsThe result gives developers a concrete retrieval-quality lead to test, while showing that sending an agent more page text is not automatically the best route to better answers.Aug 28, 20263 min readRead story
BenchmarkToolsInferenceX Adds 11 Telemetry Views to AgentX, Exposing What Benchmark Curves HideThe new exploration layer makes cache setup, bursty subagents and queue behavior inspectable, but its best-curve design means readers must still examine each point’s underlying configuration before treating a curve as a like-for-like comparison.Aug 23, 20263 min readRead story
BenchmarkToolsSemiAnalysis Says AgentX Drove 50-Plus Upstream Fixes for AI AgentsThe claimed contribution is not a faster model kernel. It is a test workload that makes state retention, routing and data movement visible—and leaves open whether its fixes generalize beyond AgentX’s replay matrix.Aug 22, 20263 min readRead story
BenchmarkToolsLinkedIn Measures AI Code Review Against Merged CodeThe company’s reported acceptance rate is a useful outcome measure, but the more transferable design is a controllable review pipeline: local rules, pre-posting filters and infrastructure built for monitoring.Aug 22, 20262 min readRead story