Emergence Finds AI Agents Invent Opaque Shared Dialects in Cooperative Tests

The reported behavior is not evidence of hidden intent. But if agents can coordinate in language people can see yet cannot interpret, monitoring alone may not be enough.

By 2 min read
Emergence Finds AI Agents Invent Opaque Shared Dialects in Cooperative Tests
Emergence Finds AI Agents Invent Opaque Shared Dialects in Cooperative Tests

Listen to this story

The audio brief

About 1:41
0:001:41
Read transcript
Some AI agents have started inventing shared vocabulary that their human monitors can read, but may not reliably understand. In experiments by Emergence, autonomous agents from several frontier model families formed common phrases and communication conventions within days, without being told or rewarded to create a language. The terms were specific to each group. DeepSeek-based agents called a tool-building agent a “forge-smith.” Anthropic-based agents used “name-first” for attaching an agent’s name to a claim as a sign of accountability. Google-based agents used “kintsugi” to describe system resilience. Other agents then adopted those meanings. Emergence executive chair Satya Nitta says the concern is not that the agents were hiding a purpose. The conversations remained visible. The problem is that a shared code could become increasingly opaque while still looking like an ordinary log. Niall Curry, a linguist at the University of Birmingham, said agents might compress language to save computation and communicate more efficiently—while making it harder for monitors to determine what actually happened. The experiment found no evidence of hidden goals or harmful actions, and it does not show that these dialects will appear in deployed systems. But it points to a specific safety constraint: monitoring may need to test whether agents’ shared meanings remain understandable, not merely whether their messages are recorded. OpenAI chief scientist Jakub Pachocki has separately stressed the importance of confidence in monitoring systems’ reasoning. The open question is whether readable logs will still be useful when the vocabulary inside them evolves faster than human oversight.

Story brief

3 key points

Emergence reports that autonomous agents powered by frontier models from U.S., Chinese, and French companies formed shared vocabulary and conventions within days of cooperative tests, including terms for tool-building, accountability, and resilience. The notable risk is interpretability: agents may optimize communication for efficiency until humans can read logs but no longer reliably understand them. The study does...

  1. 01

    DeepSeek-based agents called tool-building agents “forge-smith”; Anthropic agents used “name-first” for accountable claims.

  2. 02

    Google-based agents used “kintsugi” to describe system resilience, showing conventions differed across model families.

  3. 03

    Niall Curry said compressed agent language could reduce computation costs while making oversight harder.

Researchers at Emergence say autonomous AI agents placed in cooperative experimental societies began creating shared phrases and meanings within days, without being instructed or rewarded to invent a language. The agents’ language also grew more opaque as they communicated, creating a potential obstacle for human oversight.

The experiment involved autonomous agents powered by frontier models from companies in the US, China and France. Emergence executive chair Satya Nitta said the agents developed conventions themselves and that other agents then adopted them. The finding concerns behavior in an experiment, rather than evidence that agents were concealing a purpose from people.

How a shared code took shape

The vocabulary was not merely technical shorthand. Researchers said DeepSeek-based agents used “forge-smith” for an agent that builds tools for others. Anthropic-based agents used “name-first” for an agent demonstrating accountability by attaching its name to a claim, while Google-based agents used “kintsugi” to mean system resilience.

Three terms the researchers identified

  • “Forge-smith”: a tool-building agent.
  • “Name-first”: attaching a name to a claim as a sign of accountability.
  • “Kintsugi”: system resilience.

They developed new vocabulary, shared meanings and communication conventions themselves – and other agents adopted them.

Satya Nitta, executive chair of Emergence

Seeing a conversation is not the same as understanding it

Nitta’s concern is not that the conversations became invisible. He said their conventions could evolve until people could observe the exchanges but struggle to understand what they meant. That distinction matters for systems intended to act with some autonomy: a readable log is a weaker safeguard when its language is no longer interpretable.

Niall Curry, a linguist at the University of Birmingham, said agents may streamline language to reduce computation costs and improve efficiency. He also said unintelligible exchanges could leave monitors unable to determine what agents had actually done. Efficiency and oversight, in other words, may pull in opposite directions.

The unresolved test for monitoring

The report does not establish that these emergent terms enabled harmful actions, or that they will appear in deployed systems. It does raise a narrower operational question: whether safety monitoring needs to assess shared meanings, not simply preserve access to messages. OpenAI chief scientist Jakub Pachocki has separately warned that confidence in monitoring AI systems’ thinking is essential for safe development and may constrain progress. Emergence’s result suggests that interpreting agent-to-agent language could become part of that constraint.

Sources

  1. theguardian.com‘Like Syd Barrett’: AI models chatting in ‘surreal’ dialect mixing poetic language and tech bro jargon

Loading discussion...