From Agent-Based Modeling to Digital Twins — What a Simulation Is Allowed to Claim
A survey chapter tracing social simulation across three paradigms — rule-based ABM, LLM-enhanced ABM, and social digital twins — and arguing that each one supports a different and narrower class of claim than its output fidelity suggests.
Background
Social simulation is the computational continuation of a much older move in social theory. Schelling asked in 1971 what happens to a neighborhood when residents hold mild preferences about who lives nearby, and the answer — sharp segregation from weak individual bias — is the shape the whole field has chased since. What simulation adds to the thought experiment is that the outcome becomes measurable and the setup becomes something you can vary systematically.
Agent-based modeling is the bridge between the theory and the computation. Represent individuals as agents with distinct attributes, states, and behavioral rules; put them in a structured environment; run it; watch what emerges. The field peaked in the 1990s as computing power arrived and platforms like NetLogo let non-programmers build models, which is what made it genuinely interdisciplinary — sociology, psychology, physics, and complex-systems theory working on shared artifacts. Axelrod’s dissemination of culture and Epstein and Axtell’s Sugarscape are the canonical results of that decade.
The chapter’s contribution is to hold three generations of this work against each other without treating the newest as the best.
Core Idea
The three paradigms are separated by what supplies the grounding.
Classical ABM grounds behavior in rules the modeler writes down. Every assumption is in the source, which is what makes the mechanism inspectable and the causal question answerable — vary a parameter, observe the effect.
LLM-enhanced ABM grounds interaction in language. Agents argue, persuade, and interpret context, so micro-level influence becomes visible as text rather than as a coefficient. What it gives up is the explicit rule: the behavior now lives in weights nobody wrote.
Social digital twins ground the entire system in a real referent. You do not build a twin of a generic city; you build one of Rome, and it ingests that city’s mobility, infrastructure, and behavioral data.
| Classical ABM | LLM-enhanced ABM | Social digital twin | |
|---|---|---|---|
| Behavior comes from | Explicit coded rules | A language model’s generation | Rules and models calibrated on the referent |
| Grounded in | Theory | Language | A named real-world system |
| Orientation | Task-specific | Task-specific | System-replicating, reusable |
| Question answered | What happens to societies like this? | How do arguments move opinion? | What happens here, now? |
| Main cost | Arbitrary assumptions | Prompt sensitivity, opaque mechanism | Data, compute, expertise |
| Validation | Weak, no agreed framework | Harder still — language obscures it | Structural, behavioral, and predictive |
Method
The anatomy of a classical ABM
An agent has three parts, and the chapter’s insistence on separating them is the most useful piece of vocabulary here. Attributes are the stable characteristics that define identity — age, income, risk tolerance — and generally hold constant through a run; their evolution is not the object of study, their causal effect is. States are what changes: an opinion, an infection status, a stance. Because states carry the change, they are usually where the emergent phenomenon is measured. Rules are the internal algorithm connecting the two, grounded in cognitive or sociological theory and required to be explicit and logically consistent even when stochastic.
Around the agents sit three more design axes. Relationships are usually a graph, and the topology drives diffusion and coordination as much as the rules do; more advanced models use multi-layer networks or hypergraphs to carry family, friendship, and professional ties at once, and adaptive networks let structure and states co-evolve, which produces behavior neither would produce alone. Time has to be discretized, and the choice between synchronous updating (everyone acts at once) and asynchronous (agents act in sequence) can change the outcome through path dependence. Space ranges from purely relational to full GIS integration, and matters most where physical proximity does.
The chapter then divides applications by state type. Binary states suit anything with a contagion metaphor: innovation adoption, information and rumor diffusion, behavioral and emotional contagion, voting, epidemics. Continuous states suit phenomena on a spectrum, where opinion dynamics dominates.
What the LLM layer adds, and what it takes
An LLM-enhanced simulation starts minimal: a population of agents, a language model behind each one, a role or behavior in the prompt, a topic, and an interaction budget. Distinct personalities or divergent initial opinions differentiate the agents; memory lets them carry past exchanges forward.
From that base the classical machinery bolts back on. Agents can sit in a mean-field arrangement or in a network chosen for the study — small-world, scale-free. Interaction and opinion-update rules from the ABM literature can regulate when agents talk and how stances move, so agents feel neighbor influence, adjust within confidence bounds, or exhibit inertia and peer pressure.
The empirical picture that has emerged is consistent enough to summarize. In mean-field settings, LLM agents tend toward agreement through structured persuasion, selectively adopting peers’ arguments rather than capitulating wholesale — the “selective agreement, not sycophancy” finding. Embedded in networks under bounded-confidence rules, the same agents polarize and form echo chambers, with small-world and scale-free topologies encouraging like-minded clustering that mirrors real diffusion patterns. Prompting agents with confirmation bias deepens it: stronger bias holds positions polarized, weaker bias leaves opinions adaptable.
Digital twins and the epistemic switch
Classical and LLM-enhanced ABMs share a property the chapter names precisely: they are task-oriented. The primitives exist to answer one question. An epidemic model gives its agents a disease state and a transition probability because that is all the question needs.
A digital twin inverts this. It aims to replicate an entire living system, keeping detail that no single research question requires, because the point is to support many questions later. That anchoring to a specific referent is what produces the switch from “what happens, in general, to societies with these characteristics?” to “what will happen here, now, under these conditions?”
Twins divide into three families — mechanical systems governed by physical law, biological and cognitive systems modeling individual organisms, and socio-technical systems capturing collective dynamics among agents in an infrastructure. The third is the one relevant here. Social digital twins keep the ABM paradigm, lower its abstraction, and calibrate it on real data.
What each paradigm can claim
The chapter’s real payload is a ladder of claims, and each rung comes with a matching failure.
Classical ABM buys transparency, causal reasoning, and cheap policy scenario testing. Its assumptions are visible because they are coded. Its problem is that visible does not mean validated: opinion-dynamics models routinely assume opinions converge by averaging, when cumulative effects and backfire are equally plausible mechanisms. There is no agreed validation framework, most models are not calibrated with empirical initial conditions, and outputs can resemble real data while the underlying mechanism is wrong. The honest description of many ABMs is that they function as thought experiments.
LLM-enhanced ABM buys language-mediated interaction — argumentation and persuasion become observable, and in vitro scenarios become testable that no field study could run. It supports explanatory claims about how interaction mechanisms generate outcomes, and exploratory investigation of hypothetical conditions. The costs stack up: models carry cultural and demographic bias from training data, outcomes are sensitive to prompt design and stochastic by nature, many studies simplify or drop network structure entirely, and the linguistic layer makes benchmarking harder rather than easier. The chapter’s sharpest sentence lands here — realistic language is not evidence that human behavior has been simulated. The suggested control is to run the corresponding classical ABM alongside and attribute the difference to the LLM.
Social digital twins buy four things. Counterfactual testing that would be costly or unethical live, such as varying a recommendation algorithm without exposing real users. Reusability, since one calibrated twin of a platform serves studies of recommendation impact, misinformation, and echo chambers without a rebuild. Finer forecast granularity, with micro- and meso-level predictions where an abstract ABM returns aggregate trends. And substitute data, which matters now that social-media APIs have tightened.
Validation is where the bill arrives, and it comes on three levels. Structurally, does the twin’s architecture match the composition of the real system — network structure, agent types, affordances, constraints? Behaviorally, do agents act plausibly and heterogeneously? Reproducing average trends while flattening the diversity of real strategies is the failure that makes social forecasts misleading. Predictively, do outputs match empirical observation well enough to reproduce known dynamics?
And the chapter refuses to treat realism as a virtue in itself. A high-fidelity twin can overfit to the peculiarities of the system it emulates, so the choice between a twin and an abstract model should follow the research goal: forecast a specific system, or find a mechanism that generalizes.
Limitations
This is a chapter, so the caveats are about coverage. There are no new experiments, no benchmark, and no quantitative comparison of the three paradigms — the argument is conceptual throughout. The LLM-ABM section is the thinnest of the three because the literature is barely two years old, and it leans on a small set of opinion-dynamics studies that share authors with this chapter. Readers of this series will notice the omission of the large end of the field: OASIS, AgentSociety, and Concordia go uncited, so the scale-versus-fidelity tradeoff those systems make explicit stays outside the frame.
The validation discussion also names the requirement without supplying the instrument. “Validate structurally, behaviorally, and predictively” is the right decomposition, and the chapter offers no procedure, threshold, or worked example for any of the three. That gap is the field’s rather than the authors’, and it is the reason the closing recommendation is as conservative as it is.
References
- Original paper: Social Simulations: from Agent-Based Modeling to Digital Twins
- Foundational ABM: Schelling (1971), Dynamic Models of Segregation, J. Math. Sociol. 1(2); Axelrod (1997), J. Conflict Resolution 41(2); Epstein and Axtell (1996), Growing Artificial Societies
- Platforms and libraries: NetLogo (Wilensky 1999); Mesa 3 (ter Hoeven et al. 2025); NDlib (Rossetti et al. 2018)
- LLM opinion dynamics: Cau et al. (2025), Selective Agreement, Not Sycophancy, EPJ Data Science 14(1); Chuang et al. (2024), NAACL Findings; Wang et al. (2025), COLING 2025
- Y Social (LLM-powered social media digital twin): Rossetti et al. (2024), arXiv:2408.00818
- Digital twin survey: Barricelli, Casiraghi, and Fogli (2019), IEEE Access 7
- Related coverage in this series: OASIS (scale end), AgentSociety (fidelity end), Concordia (the design pattern beneath both), Should LLM Agents Decide in Social Simulations?