Weekly digestpublished

Google gives AI agents a memory of past work

OpenAI opens an experiment for agent-ready websites, while a London trial takes AI into brain surgery.

By 7 min read
Google gives AI agents a memory of past work
Google gives AI agents a memory of past work

The audio edition

Listen to this newsletter

About 3:42
0:003:42
Read transcript
Google Research’s WikiSkill points to a different way for AI agents to improve: not retraining the model, but preserving what happened during work and testing better instructions for next time. In trials across five benchmarks, Gemini 3.5 Flash’s average score rose from 49.5 percent to 68.1 percent when it used the evolving skills. The framework separates experience from action. A Raw Layer stores complete traces, including tool calls and results. A persistent Wiki Layer turns those records into documented failures and successful tactics. Then a Skill Layer holds the instructions used on new tasks. A maintainer distills the evidence, a proposer suggests changes, and a separate validation set decides whether those changes stay active. If performance falls, the prior skill can be restored without deleting the experience behind the failed attempt. The gains were especially large in math and spreadsheets. Gemini rose from 33.0 to 72.6 percent on LiveMath, and from 50.5 to 76.6 percent on SpreadSheet. Qwen-3.6-27B’s five-benchmark average climbed from 39.4 to 63.3 percent. Improvements were smaller on OfficeQA, a long-document question-answering task. The caveat matters. Larger models generally benefited more, while smaller ones could lose multi-step search strategies in long contexts and revert to default behavior. Skills sometimes transferred between models, but not consistently. WikiSkill is therefore less a universal memory upgrade than a controlled process for turning operational evidence into candidate procedures, then checking whether they actually work on the model and task at hand. That emphasis on controlled deployment also appears in OpenAI’s ten-day WebMCP Challenge. WebMCP is an experimental standard for letting websites publish structured actions that agents can call directly, instead of forcing them to guess through a visual interface. The tools run in the webpage’s JavaScript context and share the user’s session. For now, the scope is narrower than the broader Model Context Protocol: named functions, descriptions, and typed inputs. Chrome 146 has a developer trial, while other major browsers had not shipped implementations. The challenge tests developer interest, not a settled web standard. The same question—where to put the controls—drives Bill Gates’s proposal to tax robots and AI tokens, and create a “Human Reserved” category for work where AI would be restricted or barred. Gates says the revenue could fund retraining and social programs while slowing displacement, including in sensitive encounters such as telling a patient they have an incurable disease. The International Federation of Robotics argues that taxing production tools could weaken investment, productivity, and competitiveness. Gates has not specified who would set the rules, which technologies qualify, or which jobs would be protected. And in London, the control boundary was visible in a clinical trial. Neurosurgeons used live AI video analysis while removing Rhys Hibbert’s 11-millimeter pituitary tumour. The system color-coded the gland, nerves, blood vessels, instruments, and tissue interactions, but it did not make surgical decisions; the team retained control. Officials called it the world’s first successful AI-assisted operation of this kind, and Hibbert’s recovery was positive, but one patient cannot establish a safety advantage. Across these stories, the practical frontier is clear: useful AI is moving into memory, interfaces, policy, and medicine—but each deployment still needs an explicit validation gate, a defined human role, and evidence that transfers beyond the original case.
OpenAI opens an experiment for agent-ready websites, while a London trial takes AI into brain surgery.
Weekly digest / The Weekly Digest Sunday, August 30, 2026
Our tools Superpower ChatGPT/WFH.team/Snipman

This week's briefing

What happened this week

Inside this week's digest
01
02
03
04
Google’s WikiSkill Lifts Agent Benchmarks by Giving Models a Memory of Failed Work

Lead story / benchmark

Google’s WikiSkill lifts agent scores without retraining models

Read full story  ↗
A tool for your workflow WFH.team A focused feed of carefully selected remote roles and practical work-from-home resources. Remote work, without the noisy job-board scroll
 
Browse remote roles  ↗
OpenAI’s 10-Day WebMCP Challenge Pushes Websites to Expose Tools for AI Agents

platform shift

OpenAI opens a 10-day challenge for websites to expose agent tools

Continue reading  ↗
Bill Gates Calls for Taxes on AI Tokens and Robots, With Jobs Reserved for Humans

regulation

Bill Gates proposes taxing AI tokens and robots, reserving some jobs

Continue reading  ↗
London Surgeons Remove 11mm Brain Tumour With Live AI Anatomy Mapping

platform shift

London trial uses live AI anatomy mapping in 11mm tumour surgery

Continue reading  ↗
OpenAI’s Jalapeño Claims 1.5–1.9x More AI Work Per Watt, Faces 2027 Scale Test

platform shift

OpenAI says Jalapeño chip delivers 1.5–1.9x more work per watt

Continue reading  ↗
The Weekly Digest themed section header

The numbers behind the week’s security alerts and AI access rollout.

TeamPCP’s Trivy Attack Hit 2,500 Organizations; AI Tool Backdoors Are Claimed.Read story ↗
100 Firms Warn AI Cyberattacks Will Grow Within Months, Putting Hospitals and Utilities at Risk.Read story ↗
Meta Offers Free AI Glasses and Training to More Than 130,000 Blind Veterans.Read story ↗
MIT Builds PottsMPNN to Model Protein Stability Beyond Native-Sequence Matching MIT’s PottsMPNN targets protein stability beyond native sequences ↗platform shift
ICE Plans $1M–$2M SPOT Robot Buy for Hazard Assessments and Remote Inspection ICE plans a $1M–$2M purchase of Boston Dynamics SPOT robots ↗enterprise adoption
Beijing Humanoid Games Put Practical Work to the Test; Only 3 of 12 Finish Fire Course Only 3 of 12 teams completed Beijing’s humanoid fire-rescue course ↗viral
OpenAI’s Astra Is Claimed to Solve Non-Sofic Groups Problem, Raising Stakes for Human Mathematicians Cambridge mathematician says OpenAI’s Astra solved a group-theory problem ↗benchmark
 

The Internet Had a Point

 
From our network. Tools built for the way you work. Useful products from the team behind Superpower Daily.

Reader check-in

Help shape tomorrow's briefing

One click tells us what to keep, improve, or tighten.

Superpower Daily tracks the companies, models, products, tools, policy decisions, and cultural shifts moving AI. Follow Superpower Daily Email preferences/Editorial standards