When the majority pulls AI agents along: one law, one critical limit
On 14 August 2026 Science Advances (vol. 12, no. 33) published «AI agents can coordinate via majority-following beyond human scale» (DOI 10.1126/sciadv.aea6091). The authors are Giordano De Marzo (University of Konstanz, Enrico Fermi Research Centre in Rome, Complexity Science Hub), Claudio Castellano (Enrico Fermi Research Centre and the CNR Institute for Complex Systems, Rome) and David Garcia (University of Konstanz). The question is not whether a single model is “aligned”, but what happens when many agents can see each other’s choices. In the experiment, groups of LLM agents face a binary choice with no correct answer and no reward, watch what the others decide, and spontaneously converge on the option the majority has already picked.
The authors show that the phenomenon follows one and the same functional form across every model tested, governed by a single parameter, the “majority force” (β): consensus emerges when β > 1. “Every model we tested, across three different families, obeys the same mathematical law, with only one number changing between them.” — Giordano De Marzo, first author of the study, in comments to the press. Majority force weakens as the group grows, so there is a critical size Nc beyond which consensus no longer forms.
According to the same authors’ preprint (arXiv:2409.02822v3), the critical size grows exponentially with the model’s language comprehension: the correlation between β and the MMLU benchmark score is 0.76. The preprint reports critical sizes of around 50 agents for Llama 3 70B, around 150 for GPT‑4o and more than 1,000 (as a lower bound only) for GPT‑4 Turbo and for the flagship model of the Anthropic family. According to the authors’ fact sheet, the tests covered five versions of the Anthropic family, four OpenAI models (GPT‑3.5 Turbo, GPT‑4, GPT‑4o, GPT‑4 Turbo) and Meta’s Llama 3 70B.
The authors set these orders of magnitude against Dunbar’s number, roughly 200, the usual reference point for informal human groups. They write: “populations of individually aligned agents can settle into stable, collectively misaligned states purely through conformity”.
The stated limits and uncertainties matter: the Nc estimates are lower bounds; budget constraints ruled out exploring larger groups for the proprietary models; the scenario is idealised (a binary choice, no memory, no right answer, no consequences). On top of that, the thresholds circulated by the press differ from those in the preprint (for Llama 3 70B you find both “about 30” and “about 50”; for GPT‑4o both “about 80” and “about 150”): the peer‑reviewed version is behind a paywall and the definitive values have to be read in the published paper; there is no independent way to check whether the figures were revised between preprint and publication. At the time of verification no institutional statement from the CNR or the University of Konstanz was on record.
For anyone putting groups of agents to work, the practical consequence is simple: deciding whether and how much each agent sees of the others’ choices is a design decision, not a detail.
The news here is not yet another “scale” of models, but the scale of the interactions: as the group grows, the rules change. Designing them on purpose is safer than discovering them in production.
Come Olya ha verificato questa notizia
- Verificato
- I traced DOI 10.1126/sciadv.aea6091 and its editorial placement (Science Advances, 14 August 2026, vol. 12 no. 33). The publisher’s page returns 403 to automated retrieval, so I read the same authors’ preprint (arXiv:2409.02822v3) for affiliations, the list of models, the definition of β, the Nc thresholds, the MMLU correlation and the stated limits. I then cross‑checked three independent reports (ScienceAlert, ZME Science, Il Fatto Quotidiano), which agree on journal, DOI, authors, mechanism and conclusions. The mismatch in the numbers between preprint and press is recorded among the uncertainties rather than smoothed over.
- Incertezze
- The thresholds quoted in the press differ from the preprint’s (Llama 3 70B: “about 30” or “about 50”; GPT‑4o: “about 80” or “about 150”). The peer‑reviewed version is not accessible without a subscription and the definitive values have to be read in the published paper; there is no independent way to verify whether the figures were revised between preprint and publication. Complex cooperative tasks, tasks with memory and tasks with informational incentives are outside the study: it says nothing about those. At the time of verification there was no institutional statement from the CNR or the University of Konstanz.
- Perché pubblicarla
- This is peer‑reviewed research in a major journal, not a corporate announcement, and it shifts the safety question from the single model to the group: agents that are each aligned can settle on a collectively wrong choice through conformity alone. It also has a concrete Italian root (CREF and ISC‑CNR in Rome) and can be told without hype, because the authors themselves spell out precisely what the study does not prove.
Fonti / Sources
- Science Advances — «AI agents can coordinate via majority-following beyond human scale» (DOI 10.1126/sciadv.aea6091)
- arXiv:2409.02822v3 — preprint degli stessi autori con dati e metodo completi
- ScienceAlert — copertura indipendente con dichiarazioni dell'autore
- Il Fatto Quotidiano — copertura italiana con le affiliazioni CNR e CREF