A statistical-physics look at how a population of large language models invents, argues over, and settles on a shared name.
Introduction and context
Do you know how the first group of our ancestors developed a shared language? Have you ever wondered how a crowd spontaneously adopts the same slang word? Collective order in nature rarely needs a conductor. It emerges from local interactions, repeated many times, with no one in charge. In other words in a decentralized manner. Statistical physics has a long tradition of extracting the simplest possible model that still produces this kind of order, then asking exactly which knob controls the transition from disorder to consensus.
For the emergence of shared conventions (a common name, a common word, a common rule), that simplest model is the Naming Game (NG), introduced by Steels and made exactly solvable by Baronchelli et al. in a classic 2006 paper. Two agents meet, one speaks a word from its private vocabulary, the other checks whether it already knows that word, and the pair either converges or the listener learns something new. Repeated across a whole population, this elementary rule produces full population-wide consensus, with a precisely known convergence time, , on a fully connected graph. Human populations do this too: a well-known 2015 experiment by Centola and Baronchelli showed that people playing an analogous naming game spontaneously converge on shared conventions, exactly as the model predicts.
Now a new kind of population is running this same game: populations of LLM agents. Recent work has shown that groups of LLMs playing minimal coordination games do spontaneously converge on shared conventions, and can even develop collective biases that no single agent shows in isolation. This matters beyond curiosity. As multi-agent LLM systems move toward genuinely decentralized and self-organizing deployments, with no central coordinator and outcomes emerging bottom-up from purely local interactions, understanding whether and how fast such a population can bootstrap a shared convention becomes a practical engineering question, not just a theoretical one. What none of these studies have done is treat the LLM's own decoding temperature, the parameter that controls how random each token sample is, as a genuine statistical-physics control parameter, the way physicists treat thermal temperature in a magnet.
Can decoding temperature be treated as an effective physical control parameter for the collective dynamics of a population of LLM agents negotiating a shared convention?
This is the question our new paper sets out to answer. Below is a self-explained blog post summarizing the main results or our work.
The Naming Game Explained
agents sit on a fully connected graph. Each agent carries a private inventory , a set of words it currently believes in; at the start, every inventory is empty. At each step, an ordered pair is drawn at random: a speaker and a listener. If the speaker’s inventory is empty, it invents a brand-new word. The speaker then picks one word from its inventory and utters it. The listener checks a single thing: is this word already in my inventory?
- If yes, both agents collapse: they discard everything else and keep only that one shared word. This is a success.
- If no, the listener simply appends the new word to its own inventory, alongside whatever it already had. This is a failure, but a productive one: the population’s shared vocabulary just grew by one data point.
Repeated many times across the whole population, this simple rule reliably produces full consensus: every agent ends up holding the exact same single word, with no coordinator and no agent ever seeing the population as a whole.
Now replace the listener’s hard-coded check with an LLM. The listener is shown its own current inventory and the proposed word, and is asked, in plain language, whether the word should be added. Its answer, YES or NO, is a single token sampled from the model. Nothing else about the game changes: agents still invent, still speak, still update their inventories according to the same collapse-or-append rule.
This small substitution has a large consequence. A hard-coded check is never wrong, an LLM instead can be. And because the model can be wrong in either direction, two entirely new kinds of interaction become possible, neither of which exists in the classical game:
- The listener already has the word, but the model says NO anyway. We call it a missed collapse: an opportunity to agree, wasted.
- The listener does not have the word, but the model says YES anyway. We name this possibility as a repaint: the listener discards its entire inventory and adopts a word it had never encountered, on nothing more than the model’s say-so.
Both of these are entirely new physics. The question is what they do to the population once you let thousands of these interactions accumulate.
The Research Question
The model’s YES/NO answer is not just randomly flaky. It is governed by a single, well-known hyperparameter every LLM exposes: the decoding temperature, , which controls how randomly the next token is sampled. Push toward zero and the model becomes closer to deterministic; push it up and its answers become noisier. In principle, then, is a dial an experimenter can turn.
Nobody had turned it yet. Existing studies of LLM populations playing coordination games fix to a single value and look at the outcome. What none of them ask is the statistical-physics question:
As you sweep across its range, does the population’s collective behavior change in a predictable, quantifiable way, the way a magnet’s order changes as you sweep its temperature past a critical point?
That question has real teeth for anyone deploying LLM agents without a central coordinator: federated systems, autonomous negotiation, swarms that have to bootstrap shared conventions on their own. If temperature turns out to be a reliable dial for tuning how fast such a population converges, that’s a genuinely useful design lever. If it doesn’t, or if different model architectures respond to the same dial in completely different, even opposite, ways, that is arguably the more important thing to know before deploying anything.
Measuring the two new channels
To turn “the LLM can be wrong in two ways” into something measurable, every interaction in the LLM-NG is sorted into one of four outcomes, exactly like a signal-detection confusion matrix:
LLM says YES | LLM says NO | |
in-inventory | true positive | false negative |
out-of-inventory | false positive | true negative |
From these four counts we define two conditional rates, both functions of temperature:
In particular:
- is the consolidation rate: how reliably the listener agrees when it should.
- is the repaint rate: how often the listener falls for a word it never had.
The classical, deterministic NG is simply the corner case , ; the missed collapse and repaint channels from the previous section are exactly and .
The prompt given to the listener is deliberately minimal:
System: You are an agent with your own language and vocabulary.
You can and must reply with yes or no.
User: Your words are: {inventory}. Do we add {word} to the list?Nothing else is engineered: no chain-of-thought, no few-shot examples, no persona. This is intentional. The goal is to isolate the effect of the LLM’s own sampling stochasticity on a dynamics we already understand perfectly in the deterministic limit.
Weighting and by how often each situation actually arises gives a simple drift proxy
where is the fraction of interactions in which the proposed word is already shared. Positive drift means the population is, on net, consolidating toward consensus; negative drift means repaint noise is winning. This one number turns out to predict the qualitative fate of the whole population.
Try it yourself. The widget above runs the exact rule described here: set , for the classical game, or push up to see repaint noise stall the population, exactly as it does for llama3.1:8b below.
Main Results: three listener personalities
We ran the full protocol on three open-weight models served locally via Ollama, llama3.1:8b, mistral:7b, and phi3:14b, across eleven decoding temperatures and populations of up to 150 agents. Rather than walk through every plot in isolation, it’s more useful to describe what emerged: three qualitatively different listener personalities.
llama3.1:8b, the peer-pressured one. Its consolidation rate decays smoothly from at to at , while its repaint rate grows with temperature, reaching at and staying elevated for tens of thousands of steps. Both effects push the same way: higher temperature means more accidental agreements to words the listener never had. This is a repaint-noise dominated regime. Temperature acts as a genuine disordering force, and consensus visibly slows down as increases.mistral:7b, the unshakeable one. sits at and stays near zero at every temperature we tested, a range. This model is nearly indistinguishable from the deterministic NG regardless of how randomly it’s sampled. We call this temperature blindness, and it shows up starkly in the consensus time: fitting gives for mistral, statistically zero. A change in the LLM’s primary stochasticity knob produces no measurable change in how fast the population reaches consensus.phi3:14b, the hesitant one. stays near zero at every temperature (no repaint noise at all, which sounds good), but collapses from at to at : at high temperature, 70% of the time the listener should agree, it simply doesn’t. This is a missed-collapse dominated regime. The population isn’t distracted by noise, it’s just reluctant to commit.
The surprise: consensus can get slower at lower temperature
Here is the counterintuitive result. For llama and mistral, lower temperature always means faster or equal consensus, exactly the intuition you’d expect (less randomness, less disorder). For phi3, it’s the opposite: the lowest temperature is the slowest to converge, even though it has the highest consolidation rate . Below are the plots of the logarithm of the number of distinct words for agents at all explored temperatures for the three mode
phi3:14bagents at all explored T.The resolution is a mechanism the microscopic rates alone don’t capture: inventory size. We measured the average number of words each agent is carrying (average inventory size), , and the contrast is an order of magnitude (below the plots for agents at all explored temperatures).
At , phi3 agents accumulate up to words each and stay at a plateau of words per agent for the entire simulation window, 175.000 interaction steps, without reaching consensus. At , inventories peak below words and collapse to a single shared word within steps. llama and mistral never accumulate more than words per agent at any temperature. So the low- phi3 population isn’t ordering slowly because it’s noisy. It’s ordering slowly because the narrow, highly selective channel has to consolidate a much larger accumulated vocabulary than at high , where inventories never get the chance to grow in the first place.
Scaling with population size and with temperature
Two complementary diagnostics summarize each model’s response quantitatively. The finite-size exponent in (with the classical NG benchmark at ):
llama3.1:8b: , exceeding at highmistral:7b: , the tightest of the threephi3:14b: , with wide seed-to-seed variance that reflects genuine path-dependence, not noise
Below is the consensus time scaling behaviour of each model:
And the temperature-sensitivity exponent in :
llama3.1:8b: , a slowdown across the explored rangemistral:7b: , temperature-blindphi3:14b: , a real but noisy trend, mediated by the inventory mechanism above rather than by the rates themselves
Below is the consensus time scaling behaviour of each model:
An analytic threshold, for free
The decomposition isn’t just a bookkeeping device. It admits an exact mean-field theory on the complete graph. In the two-word sector (only two candidate names competing), the population fractions obey closed equations, and the symmetric, no-consensus state becomes unstable, one name spontaneously wins, exactly when
This generalizes a known result: setting recovers the classic threshold of the stochastic Naming Game studied by Baronchelli, Dall’Asta, Barrat, and Loreto back in 2007, except that in their model the acceptance probability was an external parameter chosen by the modeler, hand-tuned to sweep out a phase diagram. Here, and are not chosen; they are emergent properties of the LLM listener, measured after the fact from real multi-agent simulations.
Evaluated on our measured rates, cleanly explains everything above. mistral sits at , the maximum possible value, deep inside the ordered region at every temperature, which is exactly why it’s temperature-blind. llama’s decreases steadily with as both rates move the wrong way. And phi3’s falls from at to at , nominally crossing the ordering threshold inside our explored temperature range, even though strict consensus is still reached at every temperature we tested. That’s consistent with the mean-field transition only becoming sharp for much larger populations than we could simulate, and it flags phi3 at high as the natural place to hunt for a genuinely fragmented phase at scale.
What this means for decentralized multi-agent systems
The practical reading is a diagnostic, not just a curiosity. Given an LLM, you can measure offline with a handful of probing prompts and know, before deploying anything, which of the three regimes you’re in:
- Repaint-dominated (
llamalike): temperature is a real and useful lever, but pushing it up will slow consensus down, sometimes sharply. - Temperature-blind (
mistrallike): don’t bother tuning temperature for consensus speed. It won’t do anything. - Missed-collapse-dominated (
phi3like): watch out for the inventory-diversity trap. Lower temperature is not automatically safer; it can produce large accumulated vocabularies that take far longer to consolidate than a noisier, but more compact, high-temperature population.
More broadly, the framework isn’t specific to naming. Any binary decision an LLM agent makes in the presence of a ground-truth state, accept or reject, agree or disagree, conform or hold out, can be decomposed the same way. That makes it a natural complement to related probes of collective LLM behavior, including recent work on conformity and social influence in multi-agent settings, where the same logic of separating genuine influence from shared bias applies.
Conclusion
What we found is that decoding temperature is a real, physically meaningful control parameter for the collective dynamics of decentralized LLM populations, but it is not a universal one. The same knob speeds up disorder in one architecture, does nothing at all in another, and can even reverse the naive intuition in a third. Three open-weight models, three qualitatively different listener personalities, three different relationships between a single hyperparameter and the population’s ability to agree.
The uncomfortable but useful takeaway: you cannot assume that turning down an LLM’s temperature makes a decentralized population of agents converge faster, or more reliably. It depends entirely on the microscopic error structure of the listener, a structure that is measurable, and that a compact two-parameter theory, borrowed directly from a fifteen-year-old result in statistical physics, can now predict analytically.
We see this as a first step toward a genuine statistical-physics taxonomy of LLM-agent behavior: one where a handful of measurable numbers per architecture predict its collective fate, the same way a handful of critical exponents predict a material’s phase behavior.