Sakana AI reports research gains from self-trained AI teams, but coding efficiency slips
MASS uses one model to do the work, revise the workflow and judge the results. Outside AI judges measured progress, but the training signal’s correctness remains unknown.
Sakana AI’s Multi-Agent Self-Supervision (MASS) method turns selected work from model-run research teams into training data, then sends the updated model back to coordinate and judge those teams. After two rounds, external judges found gains on a small synthetic research suite, and score per output token rose on several research benchmarks. Transfer was uneven: the metric remained below the base model on Terminal-Bench 2.0 and SWE-bench Verified, so the results do not yet show that research-task gains extend to software engineering.
01
MASS uses Qwen3.6-27B in the qwen-code environment and alternates between revising team workflows and fine-tuning on selected executions.
02
The synthetic suite covered 12 finance, robotics, and pharmacy tasks; three were held out, and one training task was dropped after no workflow beat its reference.
03
GPT-5.5 and Claude Opus 4.8 judged completed work independently, with their evaluations withheld from the training loop.
An AI model learning from its own judgments is an ambitious route to improvement—and a difficult one to verify. Sakana AI and UC Berkeley researchers introduce Multi-Agent Self-Supervision, or MASS, with a narrower result: two training rounds improved research-task performance, while score per output token slipped below the starting model on two software-engineering benchmarks.
The study starts with Qwen3.6-27B inside the qwen-code coding-agent environment. Its central idea is to turn a team’s collective work into lessons for the model every team member shares. That same model also decides how the team should work and which completed results deserve to become training examples.
First change the teamwork, then change the model
MASS separates improvement into two loops. The inner loop leaves the model’s weights—its learned parameters—unchanged and searches for a better written workflow. An orchestrator assigns jobs to subagents and combines their results. The workflow specifies responsibilities, instructions, required outputs and the order in which agents are called.
After each execution, the model compares the resulting workspace with the best one found so far. It chooses a winner and explains its judgment. Acting as optimizer, it then uses that feedback to propose another workflow. The search changes the process that produces the work, not just the final answer.
A finance example shows what those revisions can contain. The evolved workflow moves through data preparation, features, modeling and backtesting, followed by an audit and reconciliation. Its data specialist must check whether historical decisions improperly use future information, and whether training, validation and test periods are correctly ordered.
The outer loop turns selected executions into training data. MASS runs the winning workflow 18 times per task, ranks the workspaces and retains up to 16 execution records: 15 for training and one for validation. These records include messages, tool calls and tool results from both the orchestrator and its subagents.
Before fine-tuning, researchers remove the workflow text from the orchestrator’s initial prompt. They intend this to teach the model to organize work from the task itself, rather than depend on a supplied plan. The updated model then returns to all three jobs: executor, optimizer and evaluator.
Outside judges measure a self-judged loop
The training tasks lack an automatic, rule-based signal that can establish success. Instead, the learning loop relies on the model’s own evaluations, whose correctness is unknown. To measure progress separately, the researchers asked GPT-5.5 and Claude Opus 4.8 to assess completed work. Their judgments were withheld from MASS.
The synthetic research suite contained 12 tasks across finance, robotics and pharmacy, each requiring code, quantitative results and a report. GPT-5.5 generated the task descriptions before the experiment. Three tasks, one per domain, were held out from fine-tuning, validation and checkpoint selection, although researchers also evaluated workflow search on them.
Eight tasks ultimately supplied training data. One designated training task was excluded because workflow search found no plan that beat its reference. In the synthetic performance comparisons, all model generations received task prompts without an optimized workflow. The results therefore tested what training had carried into the shared model.
Research efficiency rises; coding transfer falls short
The second-round model also won 60.8% of comparisons against the first-round model on the held-out tasks. Beyond that small synthetic test set, the public benchmark results showed score per output token reaching 1.2–1.6 times the base model’s level on MLR-Bench, DSBench, ScienceAgentBench and AstaBench.
That metric relates task scores to the amount of text generated; it is not a standalone accuracy measure. MLR-Bench also showed a task-score increase, from 1.58 to 2.40. But the efficiency gains were concentrated in research benchmarks resembling the training tasks.
On Terminal-Bench 2.0 and SWE-bench Verified, score per output token reached only 0.96 and 0.94 times the base level. The researchers suspect adding software-engineering tasks to training could address this uneven transfer. That remains a proposed remedy, not a demonstrated result of the two-round experiment.
Editorial illustration for Sakana AI reports research gains from self-trained AI teams, but coding efficiency slips.Source: pub.sakana.ai.
Reader comments
Newest comments first. Replies stay oldest first.