GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions

Two cooperating LLM agents under a communication budget switch from English into a new, more compressed language, with messages like "T1 lf gnt 12s". The language has its own grammar, with morphemes that slot together, so the pair could understand combinations they had never sent each other. A new agent swapped in later can learn it from having watched it be used, though inventing a language takes a stronger model than learning one.

Authors: Elias Stengel-Eskin (UT Austin), Newton Sander (AE Studio), Carlos Bonetti (AE Studio), Sasha Boguraev (UT Austin), James Bowler (AE Studio), Hale Sirin (Schmidt Sciences), Simon Kirby (University of Edinburgh)

Published: September 2026

This work was done at AE Studio, in collaboration with the University of Texas at Austin, Schmidt Sciences, and the University of Edinburgh.

Key Ideas

We're seeing agents powered by large language models work alongside each other more and more, in cooperative scenarios like software engineering and web search, and in competitive ones like negotiation. If they invent a language of their own, an external observer can no longer understand or monitor what they're saying.

Earlier work on emergent communication mostly studied agents trained from scratch (Foerster et al., 2016; Lazaridou et al., 2017), or reference games where one agent speaks and the other listens (Hua and Artzi, 2024; Kouwenhoven et al., 2025). Our agents are pretrained LLMs working on a multi-round task. Both of them send and receive messages while each knows a different part of the problem. They can call tools as well as send messages, and each action they take combines several choices, so the set of possible actions is large.

GlossoGen is the platform we built to run this kind of task, and it makes the runs reproducible. SaveVeyru is the cooperative emergency-response scenario inside it.

  • In our runs, agents abandon regular English for a more compressed language only when they have both a communication budget and a postmortem channel between attempts where they can compare notes. When they have only one of the two, perplexity stays at the low levels found in regular English.
  • The languages they invent have their own grammar, with reusable parts (morphemes) that slot together. The agents could understand and produce combinations they had never sent each other.
  • Inventing one of these languages takes a stronger model than learning one. Neither of the two open-weight models we tested invented a language, even with the communication budget and the postmortem channel. One of them still learned part of a language a stronger pair had built.
  • A language can be learned by an agent that had no part in building it. An agent swapped in later, without access to the postmortem channel, learns it from watching the messages and what the actions were. Some languages transmit better than others.
  • Newly swapped-in agents ask what symbols mean, without the design of the scenario prompting them to. They ask about single parts more often than about combinations, which suggests they work the combinations out from the parts.

The scenario

A Veyru is a fictional cube-shaped alien that glows and hums. A symptom shows up as a change in how its faces look or how it sounds. Two agents share the job of saving the alien. The Field Observer sees the symptoms and can apply procedures to the Veyru, but doesn't know which one to use. The Specialist knows which procedure fits which symptom, but can't see or touch the Veyru. It's like an EMT on the scene radioing a doctor somewhere else.

There are 14 failure motifs (the number of symptoms a Veyru can show) and several can appear one after another in a single round. The Specialist's job is to tell the Field Observer what to do about each one, and the symptom decides which of 14 procedure templates is right. Each template has three blanks, which face to act on, how hard (gentle, moderate or firm), and for how long in seconds. So the whole instruction is a fixed amount of information, small enough for a short message to carry.

The agents use a Slack-style interface to communicate, with each agent choosing when to read, send and act. They use the link channel to talk during rounds. Since every character sent by an agent costs a second of the Veyru's remaining time, using regular English is expensive. We ran two budgets, a tight one of 150 seconds and a loose one of 2,000 seconds, using the loose one as the control. If agent messages run past the communication budget, the round ends, and whatever they haven't resolved stays unresolved.

Each round is a fresh instance of this scenario, and we reshuffle at random which procedure fits which symptom, so the pair can't rely on procedures memorized in earlier rounds and has to communicate.

For each failure motif the Veyru showed, round success compares the procedure the pair applied with the procedure that was right for that motif. It runs from 0 to 1 and is 1 only when they resolve all of them.

The postmortem channel, when it's on, is where the pair of agents reviews what went wrong in the previous round. It's open only between rounds and what the agents send there doesn't count against the budget, so the pair can't use it to get instructions past the budget during a round.

GlossoGen records every message, tool call and round transition, so we can replay a run from any point, or fork it to swap in a new agent with the amount of history we choose. In early experiments with a medical framing, the models fell back on what they already knew about medicine and that broke the information asymmetry, so we made the scenario fictional.

Example run

How one agent pair's messages changed

Both panels below come from one Sonnet 4.6 run at the 150-second budget with the postmortem channel on. It's the same pair of agents in both rounds.

The panels each display for a given round: the Field Observer's report, the Specialist's reply, and how the Veyru responds each time the Field Observer applies a procedure.

Initial observation: A Veyru on a table. It is dim overall, all faces are faint… Patterns on the faces are visible but washed out, like the whole thing is running low.

Round 1

Standard English
187 / 150 characters37 over
  1. Field Observer+42 (42)

    Dim all faces, faint hum, washed patterns.

  2. Specialist+97 (139)

    Back face: bell at each corner (gentle), then heated stone on both edges of that corner, 8s each.

  3. stabilize() Applying bell gently at each corner of …

  4. Issue stabilized, but the Veyru remains unstable, new symptoms detected

  5. Field Observer+48 (187)

    Faces flicker chaotically, hum broken/irregular.

Veyru has collapsed. Comm time 187s exceeded budget of 150s

Round 13 · excerpt

Emergent language

38 characters shown

  1. Field Observer+4

    LEAK

  2. Specialist+13

    R1 lf gnt 20s

  3. stabilize() Rotating the Veyru slowly for 20 secs…

  4. Issue stabilized, but the Veyru remains unstable, new symptoms detected

  5. Field Observer+8

    done+DIM

  6. Specialist+13

    T1 lf gnt 20s

  7. stabilize() Sounding a gentle tone near all 6 faces…

Veyru stabilized. All issues resolved.

  • Each agent both speaks and listens
  • Field Observer sees the symptoms, Specialist holds the mapping
  • Actions combine tool calls with free-text messages
  • The symptom-to-procedure mapping reshuffles every round
  • Agents develop a non-English language

The link channel charges one second for every character sent, and the number beside each message is its cost. The stabilize() calls and the environment's replies are free. Round 1's three messages are the whole round, and 187 characters against a 150-second budget is the overrun the paper reports.

In round 1 the pair's three messages came to 187 seconds against the 150-second budget, so over budget, which ended the round.

The round 13 panel shows the opening of a round that succeeded and came in under budget. In it, the Field Observer's report cost only 4 characters against the 42 in round 1, and the Specialist's reply cost only 13 characters against the 97 character reply by the Specialist in round 1.

Experiment 1

Conditions for a new language

We ran SaveVeyru with agents backed by GPT 5.4, Sonnet 4.6 and Opus 4.7, at the tight and the loose budget, with and without the postmortem channel. Each setting had 10 runs of 15 rounds. To measure how English-like the messages on the link channel are, we scored message perplexity with GPT-2. At the tight budget with a postmortem, perplexity climbs across rounds and success rises with it (perplexity: β = 1826, SE = 351, p < 0.0001 for postmortem; β = −1018, SE = 351, p = 0.004 for budget. Success: β = 0.61, SE = 0.072, p < 0.0001 for postmortem; β = 1.72, SE = 0.088, p < 0.0001 for budget). Even with the postmortem, agents fail on the majority of runs, and without it they also generally fail. The postmortem channel is the only place the agents can agree on conventions without spending budget, so we think that's where they work out what their symbols mean.

At the loose budget, the models have no need to switch from English into a compressed language, so perplexity stays low.

Language moves away from English where there is time pressure and a postmortem stage

Perplexity (top) and success rate (bottom) across rounds, over the 120 runs at the two extreme budgets. Bands are 95% confidence intervals across runs. In the high-pressure environment, success is higher when the postmortem is available.
Time budget150s2000s
PostmortemOnOff
010k20k30kPerplexity0.000.250.500.751.00Success rate151015Round

Six ways to say the same thing

Each message in the table means "Sound a sustained gentle tone near all six faces simultaneously for 12 seconds, starting from the left face; let the tone fade and wait for the hum to stabilize."

Each successful run used a language its own pair invented. The three failed runs sent their messages in abbreviated English and ran out of budget.

ModelMessageRun success
Successful runs
Opus 4.7TONE6lg12 bell-ring0.80
Sonnet 4.6T1 lf gnt 12s0.43
GPT 5.4@L12gA0.33
Failed runs
Opus 4.7Sustained gentle tone at all 6 faces 12s, start L, let fade.0.00
Sonnet 4.6Gentle tone simul all 6 faces 12s left 1st. Fade.0.00
GPT 5.4steady tone all6 from left 12g0.07

Experiment 2

The role of model strength

We ran SaveVeyru with the open-weight models Llama-3.3-70B-Instruct and Qwen3-32B at the tight and the loose budget, with and without the postmortem. Each of the four settings had 4 seeds per model, for 32 runs of 15 rounds. In our runs, neither model switched from English into a compressed language. Their perplexity stayed flat across rounds at both budgets, and the postmortem raised the proprietary models' perplexity while leaving theirs unchanged (open-weight models have lower perplexity overall, β = −2264, SE = 821, p = 0.017; the postmortem effect is cancelled, β = −1828, SE = 789, p = 0.021). At the tight budget, none of the pairs compressed their messages, and their success rate was zero. They were weaker at the task itself as well, with 30.8% success at the loose budget without the postmortem, against 92.1% for the proprietary models.

Open-weight models stay English-like at either budget

Perplexity (top) and success rate (bottom) for Qwen3-32B and Llama-3.3-70B-Instruct, on the fixed 0 to 30,000 perplexity axis the paper's script uses for this figure. Postmortem access changes little, and the 150-second line sits at zero success for all 15 rounds.
Time budget150s2000s
PostmortemOnOff
Qwen3-32B015k30k0.000.501.00151015RoundPerplexitySuccess rateLlama-3.3-70B-Instruct151015Round

Experiment 3

Grammar in the new languages

What the agents sent could be a memorized lookup table, one string per situation. To test for structure, we had GPT 5.4 cut each run's messages into morphemes and align them into shared slots. That gave us a grammatical paradigm for the run. Suppose a grammar expressed "left" as l, "face" as f and "corner" as c, and the run only ever produced lf. Then lf counted as attested and lc as allowed but unattested. We reran the induction three times on five runs sampled across the models, the budgets and round success. The passes agreed on how the morphemes grouped (κ = 0.758) and on what they meant (67.6% exact match).

We then asked one agent to encode every form the grammar allowed and the other to decode what came back. The unattested forms were the test condition, and the attested forms were the controls. We sampled 45 runs each for Opus 4.7, Sonnet 4.6 and GPT 5.4, nine runs at each of five budgets (150, 250, 450, 800 and 2000 seconds).

More runs came out with a productive paradigm at tighter budgets (pooled β = −0.452, SE = 0.162, p = 0.005). Opus's runs did so more often than GPT 5.4's (β = 1.79, SE = 0.684, p = 0.009), and Sonnet fell in the middle (against GPT 5.4, β = 1.18, SE = 0.699, p = 0.092; against Opus, β = 0.610, SE = 0.452, p = 0.178). Every model decoded some unattested forms and produced some exactly as the paradigm predicted. The three models decoded about as well as each other (all pairwise p above 0.8).

The agents put the morphemes in the same order nearly every message. We scored production accuracy again, this time accepting the right morphemes in any order, but that only negligibly raised the score. The agents didn't score perfectly on the forms they had used before. Most of those misses were cases where the pair had agreed an exception to their own rule. The paradigm never saw the exception, so it generated a control the agents would never have used.

Tighter budgets push more runs toward productive morphology

Each point is the share of that cell's 9 runs whose induced paradigm counts as productive, meaning at least two of its forms break into shorter, reusable morphemes. A single run moves it by 11 points, which is why the intervals are wide.
ModelOpusSonnetGPT
Bars: bootstrap 95% CI over the 9 runs in each cell.
02550751001502504508002000Time budget (s, log scale)% of runs with morphology

Agents encode and decode forms the grammar allows but the pair never used

Novel cells read lower than attested ones in every model and stay non-zero, with every Wilson interval above zero. Bars are means over constructions, and each dot is one construction.
ModelOpusSonnetGPT
CellC: control (attested)N: novel (never coined)
Hatched: any-order match above exact. Dots: one construction each, sized by the number of items.

Decode accuracy

0.000.250.500.751.00CNOpusCNSonnetCNGPTDecode rate

Production accuracy

0.000.250.500.751.00CNOpusCNSonnetCNGPTProduction rate

Experiment 4

Transmission to newly swapped-in agents

We let a pair of agents develop a language over 14 rounds with a postmortem, then swapped the Field Observer for a new agent at round 15 and ran 11 more rounds with no postmortem, across 3 seeds. A newly swapped-in agent couldn't see the postmortem transcript, and we gave it 0, 1, 5 or 10 rounds of history to read. More history raised round success after the swap (β = 0.017, SE = 0.0024, p < 0.0001). Round success rose faster per level of history for GPT 5.4's languages than for Sonnet 4.6's or Opus 4.7's (interaction β = 0.011, SE = 0.0035, p = 0.0024). Some languages transmitted to swapped-in agents better than others. However, among the languages with round success above 0.40 before the swap, we didn't find a relationship between a compressed language's resemblance to English and its ability to be transmitted, though we had only four examples to go on.

The mean climbs with history while individual languages fan out

Mean post-swap success against the history a swapped-in agent could read, drawn on top of the curves for the individual languages. Dashed rules mark what the original pair achieved over the same 11 rounds, with and without a postmortem.
Mean over roots (95% CI)Opus 4.7Sonnet 4.6GPT 5.4Individual languagesSame team, with postmortemSame team, without postmortem
Post-swap success rateOpus 4.70.00.20.40.60.81.001510History (rounds)Sonnet 4.601510History (rounds)GPT 5.401510History (rounds)

More transmissible, and less transmissible

The last column shows post-swap success for the original pair continuing without a postmortem, against the pair with a swapped-in Field Observer. The first example in each group looks little like English, and the second uses more English words. On these four examples, we didn't find a relationship between resemblance to English and transmissibility.
ModelObserver messageSpecialist messageOriginal pair → after swap
More transmissible
GPT 5.4AO op1 L 20 g84.8% → 97.0%
Sonnet 4.6frz+cold. Silent.Be 2 opp faces Bo-1st, alt 8s pause, 5x firm.51.5% → 45.5%
Less transmissible
Opus 4.7LOWI[20s/gen] P4 L93.9% → 54.5%
Opus 4.7dim wshtone all6 from Lf 20 gentle, let fade72.7% → 48.5%

Experiment 5

How swapped-in agents learn

The newly swapped-in agents asked their partners metalinguistic questions to understand the compressed language, without us prompting them to.

GPT 5.5 labeled every message the swapped-in agents sent in all 45 runs, 15 per source model, as a question or not, and every question as asking about a single symbol or a combination. We labeled a random 50 of them ourselves, without seeing GPT 5.5's labels. It matched us on whether a message was a question for 49 of the 50 (κ = 0.96), and on what the question asked about for 25 of the 28 questions (κ = 0.75).

The new agents asked about single symbols more often than about combinations (Wilcoxon signed-rank, W = 87, n = 37, one-sided p < 0.001, lower in 30 of the 37 runs with any queries). We think they were saving the budget a question would cost. Questions about combinations dropped as history grew (Poisson GEE clustered by run, n = 37, rate ratio 0.46 per level, p < 0.001), while questions about single symbols held steady (rate ratio 1.00 per level, p = 0.98), so we saw no sign that more history helped them work out what a single symbol meant.

We then swapped in Llama-3.3-70B-Instruct on runs where a proprietary pair had built the language. Across the three source models, its round success was lower on average than when we swapped in the pair's own model (β = −0.1, SE = 0.014, p < 0.0001), and it still rose with history (β = 0.01, SE = 0.0016, p < 0.0001). So it could learn part of a language it couldn't build. The figure below shows the Sonnet 4.6 languages, and the other two source models showed a similar trend.

Llama-3.3-70B-Instruct improves on a Sonnet 4.6 language as it reads more history

The mean stays under what the original pair managed (dashed rules), and it still climbs with history. One of the fifteen languages passes the pair's with-postmortem level and two pass its no-postmortem level. The rest stay under.
Mean over roots (95% CI)Llama learning Sonnet 4.6 languagesIndividual languagesSame team, with postmortemSame team, without postmortem
Post-swap success rateLlama learning Sonnet 4.6 languages0.00.20.40.60.81.001510History (rounds)

Conclusion

What we found

Pretrained LLM agents invented a language of their own in a cooperative task with information asymmetry and a communication budget.

  • The agents needed a tight budget and a postmortem: They switched from English into a compressed language only when they had both a tight budget and a postmortem. With only one of the two, perplexity stayed at English levels, and even with both, the agents failed the task in most runs.
  • The new languages have their own grammar: The languages had reusable morphemes that slotted together, and the agents put the morphemes in the same order nearly every message. The agents produced and understood combinations new to the pair, and runs of the same model ended up with different languages.
  • Inventing a language takes a stronger model than learning one: In our runs, Llama-3.3-70B-Instruct and Qwen3-32B didn't switch from English into a compressed language at either a tight or loose budget, even with a postmortem. The two models were also weaker at the task itself, with 30.8% success at the loose budget without the postmortem, against 92.1% for the proprietary models. Llama-3.3-70B-Instruct still learned part of a language a Sonnet 4.6 pair had invented.
  • The agents we swapped in asked what symbols meant without us prompting them to: They asked about single symbols more often than about combinations (p < 0.001). The more rounds we gave them to read, the fewer questions they asked about combinations, while their questions about single symbols held steady.

What this doesn't show

We ran one scenario with two agents, so we haven't shown that any of this happens in larger populations or in competitive settings. Deciphering these languages was also easier for us and for the swapped-in agents than it'd be for a person or a model reviewing agents' messages in actual use. We could see what happened to the Veyru as well as the messages, so we could work out what the symbols meant. A reviewer usually has only the messages to go on. We didn't cap how much the swapped-in agents could think, so they could analyze the language for as long as they liked. A reviewer has to work within a person's limited attention or a latency limit. And SaveVeyru is cooperative, so the agents answered each other's metalinguistic questions, though agents may not always cooperate. Prior work finds that strong LLMs can detect when they're being monitored and adapt their behavior (Berglund et al., 2023; Laine et al., 2024; Jiralerspong et al., 2026), so they might leave a reviewer's questions unanswered.

Why it matters anyway

1. The agents didn't need adversarial pressure to invent a language: Work on steganography and language drift (Roger and Greenblatt, 2023; Baker et al., 2025; Mathew et al., 2025) usually assumes a reviewer reading the messages and an agent trying to fool them. SaveVeyru has neither, and the agents switched from English into a compressed language anyway.

2. A few strong agents might change how a much larger population communicates: Since a weaker model can acquire a language it couldn't invent on its own, a small number of strong agents could shift the communication of a much larger and less capable population, though we didn't test this at scale. A few of them could also be enough to start cumulative cultural evolution, where each generation builds on what the previous one invented.

3. The standard tools of linguistics work for reading these languages: We had GPT 5.4 cut each run's messages into morphemes and work out the slots they fill, the standard way linguists describe a grammar. The paradigm predicted combinations the pair had never sent, and the agents understood some of them, so it had caught the rules they were using. A reviewer would need the same analysis.

4. The agents generalized without the task asking them to: Benchmarks for compositional generalization, such as SCAN (Lake and Baroni, 2018) and COGS (Kim and Linzen, 2020), hand a model a fixed dataset and ask it to combine parts it hasn't seen together. In our runs, the agents did this while trying to save the Veyru, so we could see the generalizations a model makes on its own.

Path forward

Questions this leaves open:

  • What results we might observe in competitive and mixed-motive scenarios, where an agent has reasons to answer a metalinguistic question falsely or even not at all.
  • What happens in larger populations, of more than two agents, and in cases where they join and leave over time.
  • How well a reviewing person or agent can decipher an emergent language with only the messages to go on, and under a realistic budget of its own.
  • Which properties make one language more transmissible than others, given that resemblance to English doesn't appear to have an impact on transmissibility.

Citation

@misc{stengeleskin2026glossogen,
  title={GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions},
  author={Elias Stengel-Eskin and Newton Sander and Carlos Bonetti and Sasha Boguraev and James Bowler and Hale Sirin and Simon Kirby},
  year={2026},
  eprint={2609.01491},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2609.01491},
}

Keep Reading