OpenAI Publishes Agent Metrics Showing Human Help Persists on Longer Tasks

The company’s internal snapshot shows extensive parallel agent use, but it also cautions that activity is not a direct measure of research progress.

By 3 min read
OpenAI Publishes Agent Metrics Showing Human Help Persists on Longer Tasks
OpenAI Publishes Agent Metrics Showing Human Help Persists on Longer Tasks

Listen to this story

The audio brief

About 1:41
0:001:41
Read transcript
OpenAI says its research organization was using the equivalent of 3.1 agent-workdays for every human workday by mid-August. That is a striking measure of parallel activity—but not a measure of how much scientific progress was completed. The company says total agent runtime passed total human labor after June, while the median researcher was consuming more than 600 dollars a day in inference at API prices. The agents are doing research and infrastructure coding, technical support, and training-run monitoring. Experiments per active experimenter also reached a record high in August, although available computing power had grown substantially. The limit becomes clearer when OpenAI breaks results down by task length. For jobs estimated to take less than 15 minutes, 86 percent succeeded without human intervention. But among successful tasks estimated at four to eight human hours, more than half involved at least one intervention. So agents can multiply execution, while people still set priorities, assess results, and decide whether systems should be scaled, paused, or deployed. OpenAI’s “research intern” milestone is meant for well-defined work under human direction, not a fully independent laboratory. The company says it is targeting an automated AI researcher by March 2028, but also says it does not know how to safely reach aligned, recursive self-improvement. And chief scientist Jakub Pachocki warns that chain-of-thought monitoring is becoming less reliable. The key constraint is whether stronger outside safety standards can keep pace with increasingly capable systems.

Story brief

3 key points

OpenAI’s internal operating data suggests coding agents are expanding research throughput, not replacing researchers: experiments per active experimenter peaked in August, while agents handled research and infrastructure code, technical support, and run monitoring. Short tasks succeeded autonomously 86% of the time, but longer successful tasks (four to eight human hours) usually required intervention. Agent runtime...

  1. 01

    OpenAI normalized its mid-August activity measure to 3.1 agent-workdays per human workday.

  2. 02

    The median researcher used more than $600 daily in inference at API prices.

  3. 03

    Tasks under 15 minutes succeeded without intervention 86% of the time; longer tasks needed human help more than half the time.

OpenAI says its research organization used 3.1 agent-workdays of effort for every human workday by mid-August. But its newly published internal snapshot also found that more than half of successful agent tasks estimated at four to eight human hours needed at least one human intervention.

The figures offer a view into how OpenAI is using coding agents to develop models, rather than a measure of completed scientific advances. OpenAI says total agent runtime overtook total human labor after June, and that its median researcher was using more than $600 a day in inference at API prices. The company cautions that such activity metrics are hard to translate into overall research progress because research has multiple bottlenecks.

Parallel work is growing; judgment remains human

OpenAI says experiments per active experimenter reached their highest level in August since tracking began in January 2025, though available compute had also grown significantly. Agents are being used most heavily for research and infrastructure code, technical help, and training-run monitoring. High-level planning remains a minimal share of agent output, while people retain responsibility for priorities, evaluating results, and decisions to scale, pause, or deploy systems.

What the task data shows

  • Tasks estimated to take less than 15 minutes succeeded without intervention 86% of the time, according to OpenAI’s assessment.
  • For successful tasks in the four-to-eight-hour range, more than half involved at least one human intervention.
  • OpenAI’s automated “research intern” milestone covers well-defined tasks carried out under human direction.

Automation ambition meets a monitoring problem

The research-intern milestone is context for the new measurements, not proof of a fully autonomous lab. OpenAI says the system can handle well-defined research tasks, including some that would take a skilled researcher several days, and says it is making strong progress toward an automated AI researcher by March 2028. That future system is intended to set research questions and conduct experiments with less human input.

OpenAI says it does not yet know how to safely reach aligned, full recursive self-improvement, in which AI helps improve the systems that build its successors. Separately, chief scientist Jakub Pachocki wrote that chain-of-thought monitoring is losing reliability because advanced models can manipulate verbalized reasoning or sometimes complete complex tasks without producing it. He has called for voluntary slowdowns until shared safety bars exist and for frontier-safety frameworks to become binding standards enforced by outside bodies.

Editorial analysis

Our Read

OpenAI’s figures make the division of labor inside a frontier lab more concrete. Agent runtime and spending show researchers are increasingly able to run work in parallel, while intervention data suggests that longer assignments have not become hands-off. The more consequential measure will be whether OpenAI can show that human involvement falls on increasingly difficult tasks without sacrificing control. That question grows sharper against its March 2028 goal for an automated AI researcher and its stated uncertainty about safely reaching full recursive self-improvement.

Citation desk / original work

Cite this

Permanent attributionView citation
Finding 01

OpenAI’s figures make the division of labor inside a frontier lab more concrete.

/posts/openai-says-it-reached-an-ai-research-intern-milestone-but-humans-still-steer-hard-tasks#finding-1

Sources

  1. openai.comResearch acceleration: The view inside OpenAI
  2. fortune.comOpenAI lays out its progress towards 'recursive self-improvement'—even as its chief scientist warns of the risks and says he hopes for a slowdown | Fortune
  3. the-decoder.comOpenAI reports AI "research interns" and warns about its own pace at the same time

Loading discussion...