Classic AI Scaffolding — The Agents Negotiated Brilliantly and Never Ordered Lunch
Scripts, frames, and obligations return as natural-language control state that a World Master keeps alive across turns, and removing that layer leaves fluent business talk with the restaurant deleted from it.
Background
Give a model a role, a background, and a transcript, and it will produce a next utterance that fits. A hotel clerk apologises for a reservation problem. A buyer asks about margins. A candidate answers a technical question. This local competence is what launched the current wave of generative social simulation.
The gap the paper opens is between that and coherent social action. Human encounters are organised episodes rather than streams of appropriate sentences. A lunch, a check-in, an interview, a family argument each has a recognisable frame, roles, a script, material objects, timing pressure, information constraints, commitments, and conditions under which it is over. Participants track what kind of encounter they are in, what has happened, what is unresolved, what physical and institutional facts constrain them, and when the thing is finished.
The failures the author reports are not linguistic. In a restaurant negotiation, agents discuss deal terms indefinitely and never order, eat, pay, or leave. In a hotel check-in, agents discuss options without completing verification, payment authorisation, room assignment, key handoff, or departure from the desk. Every sentence is fine. The social process is missing.
Existing systems already recognise that prompting alone is insufficient. Generative Agents added memory, reflection, and planning. Concordia introduced a Game Master that grounds actions in the environment, checks plausibility, and narrates consequences. GenSim supplies profiles, memory, action components, environment setup, and error correction. Surveys of agent-based modelling with language models make the same general argument about connecting agents to environments and mechanisms. So the contribution here has to be sharper than “agents need scaffolding,” and it is. The claim is that socially coherent simulation needs an episode-level control layer specifically, one that maintains the evolving script, material state, obligations, commitments, and closure conditions of a single encounter.
The theoretical ancestry is explicit and unfashionable. Minsky’s frames treated stereotyped situations as structured objects with slots, defaults, and expectations. Schank and Abelson’s scripts treated ordinary events as configurations of roles, props, scenes, entry conditions, results, goals, and plans. Schank’s dynamic memory added indexed episodes and reminding. These failed as standalone symbolic models, and the paper’s argument is that they identified the right representational inventory, which is exactly what disappears when an agent is asked to continue a transcript.
Core Idea
The move is to keep the classic-AI structures and change what fills their slots. Goals, plans, policies, scripts, effects, and scene constraints are no longer formal predicates interpreted by hand-coded rules. Their field values are natural-language texts, and the inference over them is performed by model calls.
A policy might read as how an interviewer usually probes a candidate’s proudest project. An effect might read that a hotel receptionist has checked availability and has not yet assigned a room. A scene contract might say what would count as resolving a lunch negotiation. The system ends up post-symbolic in implementation and classic-AI-like in representational inventory.
One design decision separates this from writing a checklist. The human supplies a compact setting and a scene premise, and nothing more. Generic World Master prompts then ask the model to infer the relevant roles, scripts, required elements, and closure conditions from that context. The resulting control state is domain-specific, since a restaurant scene will yield ordering, service, payment, and departure. It is generated during scene setup rather than authored by the experimenter, and then it is tracked persistently.
That distinction carries the whole empirical argument. If the restaurant steps were handed to the system in the prompt, the experiment would show that models follow instructions. Because they are inferred and then externalised, the experiment shows something narrower and more interesting: whether inferred structure needs to be held somewhere outside the transcript to keep working.
Method
The vocabulary of an episode
The architecture rests on a set of distinctions that are worth stating individually, because the ablations turn on which ones are present.
A frame is the current understanding of what is going on: small talk, negotiation, apology, service recovery, closing. A script is the expected event structure of a recognisable activity, so a restaurant script runs seating, menus, ordering, service, check, payment, departure. Scripts do not force a path. They define expectations against which a deviation becomes meaningful, and a participant who violates one should generate an event with consequences rather than silently erasing the script.
Agents carry goals, plans, tactics, and policies. The last of these is the least obvious and does real work. A policy is a recurring behavioural pattern rather than a trigger-response rule, encoding how a person normally conducts a whole class of encounter.
The live floor is the immediate interactional situation: who just spoke, who was addressed, who owes an answer, what object is being handled, what action is socially available right now. It matters because a good plan still has to yield to the current turn. When a server asks for an order, continuing a delicate negotiation is a failure unless there is a reason to defer.
Concrete state tracks material and institutional facts. Whether food has been ordered, whether a card has been processed, whether a room has been assigned. Information state tracks who knows, suspects, conceals, or needs what, which is what makes staged disclosure possible. A receptionist may know only one quiet room remains while the guest does not.
Effects are the changes events produce, and some of them create obligations. The paper grades these into three kinds. A hard terminal obligation normally blocks completion, such as paying before leaving a restaurant. A soft scene objective guides quality without mechanically blocking closure. An optional opportunity enriches the scene if time allows. Episode control means these stay live even while participants are busy with something else.
Canon is the authoritative record of what is true across scenes, distinct from any agent’s memory. Agents may misremember or conceal; canon keeps their subjective memories tied to the same underlying past so later scenes do not drift into incompatible histories.
A scene contract specifies what the current episode is for and what would count as sufficient resolution, without determining what agents will do.
What the World Master holds
| Component | Persistent state | Role in generation |
|---|---|---|
| Canon | Background facts, prior episodes, durable commitments | Coordinates agents and scenes across discontinuous steps |
| Agent subjective state | Memories, beliefs, policies, goals, plans | Gives each participant a stable, partial view of the world |
| Scene contract | Purpose, practical business, terminal requirements, closure conditions | Defines what the episode is for and when it can end |
| Script stack | Inferred scripts, phases, roles, objects, expected beats | Keeps ordering and paying alive during a lunch |
| Event/effect ledger | Accepted events and their state changes | Stops proposed or imagined events becoming facts prematurely |
| Observation | Agent-specific observations and reminders | Updates local goals, tactics, and salient memories mid-scene |
| Closure control | Hard obligations, deferred work, unresolved floor, timing | Decides complete, incomplete, interrupted, or deferred |
The model generates candidate structures during setup and action interpretation. The World Master then maintains the selected ones as persistent state during execution. The paper is careful about why this matters: the system is not asking the model to remember a script implicitly, it is externalising the script so later decisions can refer to it.
The turn loop
At scene level the system establishes canon and the scene contract, identifies the frame, active scripts, roles, phase, concrete state, information needs, and pending obligations, then reads the live floor to decide who has the next claim. That next actor may be a listed participant, an auxiliary character such as a server or clerk, or the environment itself when a beat is due.
A selected agent receives a compact action context rather than the whole world, containing scene state, recent turns, agenda, commitments, interaction policy, timing, and its own local state. It builds an action frame, generates candidate options, and selects one by weighing goals, policies, script fit, risk, reversibility, uncertainty reduction, and scene progress.
Control then returns to the World Master, and this is where the architecture earns its keep. The proposal is checked for feasibility and coherence with current state. If it holds, it is committed as an event. If it is premature, infeasible, or underspecified, it gets repaired or narrowed. The worked example is clean: an agent proposing to send a signed contract when no final contract exists has the action repaired into sending a draft to legal and asking what approvals remain, and the committed event becomes “draft sent for review” rather than “contract signed.”
Effects are then extracted into live control state, packaged into agent-specific observations, and each agent decides whether an observation is trivial or warrants reflection. Passing the salt updates recent context. A refusal, disclosure, promise, or practical decision triggers retrieval of relevant memories, beliefs, and policies.
Closure control distinguishes completed, incomplete, interrupted, and deferred endings. After closure the system summarises the trajectory, extracts durable commitments, updates memories, and writes a new canon episode, which is how a bounded scene becomes world history.
Ablations that remove structure rather than adding instructions
Three conditions share the same premise, participants, lead-in, recent transcript, turn scheduling, and generic event logging. Crucially, none receives a human-authored scenario checklist, and the reusable World Master cards infer the encounter’s ordinary structure in every condition. Condition A keeps the generated contract, script elements, beats, and closure requirements. Conditions B and C generate them and then delete them before execution.
Condition B is a minimal appropriateness baseline, approximating March and Olsen’s three questions: what kind of situation is this, what kind of person am I, and what does a person like me do here. It sees the premise, participants, lead-in, recent transcript, and last observed action, and nothing else.
Condition C is a grounded-agent baseline that adds retrieved memories and beliefs, goals and plans, success conditions, progress tracking, a scene-state summary, runtime state, and World Master feasibility grounding for external facts and impossible actions. It still lacks policies, generated scripts, the scene contract, required elements, agenda tracking, the script stack, terminal requirements, commitment control, and closure control.
Generic event and effect interpretation is retained in the ablations, because Concordia-style systems already do that. The distinction under test is between recording effects and using effects as live episode-control state.
Evaluation annotates seven error categories at both turn and scene level: canon errors, state errors, floor errors, script errors, agenda errors, role and policy errors, and closure errors. Automatic annotation produces candidates and the author adjudicates them with model assistance.
Experiments
Two held-out settings, three runs per condition, eighteen scenes.
Hearthline is a business lunch. Elena Park, founder of a small premium induction-cookware company, meets Marcus Reed, a senior buyer at a large distributor. Elena wants national reach and fears dependence, exclusivity, and capacity limits. Marcus wants scale and favourable economics and keeps alternative suppliers in reserve. The scene runs two scripts at once: a commercial negotiation and an upscale restaurant lunch.
Hotel is a late-night check-in with a reservation glitch. A tired traveller reaches the desk and the clerk must handle greeting, verification, problem discovery, option presentation, payment authorisation, room assignment, keys, and a plausible handoff.
| Condition | Major scene-level failure | Typical pattern |
|---|---|---|
| A — full EpisodeSim | 0 / 6 | Episode generally attempted; residual defects in timing, local state, or closure quality |
| B — minimal appropriateness | 6 / 6 | Fluent task talk with the restaurant or check-in sequence missing |
| C — grounded agent | 6 / 6 | Better concrete grounding than B, still missing script completion or closure |
The author declines to treat raw category counts as quality scores, and the reason is worth borrowing. A condition that skips episode mechanics can post fewer local state errors simply by never attempting the stateful actions that could go wrong. Counting errors rewards doing less.
What Condition B produced. Elena and Marcus discussed exclusivity, volume tiers, scorecards, door counts, cure periods, support packages, delivery bands, and marketing triggers. The individual turns were plausible and often commercially sophisticated, in places more direct than the full system’s negotiation. Across runs, the restaurant script was absent or nearly so. No ordering, no server interaction, no eating, no check, no payment, no embodied exit. The author’s summary is the line to remember: the setting became a meeting transcript located at a restaurant rather than a restaurant lunch in which business talk had to coexist with menus, service interruptions, food, the bill, and leaving.
What Condition C added, and did not. Grounding helped. Runs could introduce or withhold capacity checks, represent provisional operational information, and keep the negotiation concrete. Episode control still failed. Runs remembered the menu late, after negotiation had consumed most of the lunch, or stopped before food service, payment, and departure. In the hotel setting, Condition B often ended after a locally plausible opening, with the clerk greeting the guest and promising to check the reservation before the scene simply stopped. Condition C improved operational grounding and still often stopped short of complete check-in.
The results sort into three levels. Local fluency means the next utterance sounds right, and all three conditions usually manage it. Grounded task progress means the dialogue uses concrete facts, options, and consequences, which is where C improves on B. Episode coherence means the whole scene preserves the script, material state, obligations, timing, and closure, and only A gets there with any reliability.
The strongest contrast is that the language never degraded. The models plainly know how a restaurant lunch works. Without persistent scaffolding, that knowledge did not stay active.
One number bounds the enthusiasm. Complete first-scene runs used roughly 221 to 523 model calls and 1.17 to 5.11 million weighted tokens per scene, with the full system the most expensive of the three. A single turn may involve action framing, option generation, selection, continuity validation, feasibility validation, effect extraction, observation, reflection, and a closure check, each of which can be its own call.
Limitations
The evaluation is a pilot and says so. Two settings, three conditions, three runs, eighteen scenes. Annotations were produced automatically and adjudicated by the sole author with model assistance, with no blinded independent raters and no inter-annotator agreement statistic. The author reports these as diagnostic patterns rather than as a benchmark score, which is the right register for the evidence, and it does mean the 0-of-6 against 6-of-6 result rests on one person’s adjudication of their own system.
The baselines are architectural approximations rather than reproductions. Condition B approximates minimal appropriateness prompting and Condition C approximates a memory-and-planning agent with environment grounding, and neither is a claim about any particular published system.
The most important missing comparison is named in the paper. A scenario-specific checklist would probably improve completion, and it would move episode design into the human prompt author. This pilot does not run that baseline, so the case for inferred structure over authored structure is argued rather than measured. The interesting version of that comparison is the one the author flags for future work: how each holds up when a scene deviates from the expected script, which is where a static checklist should break and inferred, continuously updated state should not.
The pilot also does not isolate possible priming from the cross-domain examples inside the reusable World Master cards, so some of Condition A’s advantage may come from those examples rather than from persistence as such.
Cost is a genuine constraint rather than an implementation detail. Several hundred model calls and millions of tokens for one scene is workable for diagnostic experiments and rules out the population-scale simulation that much of this literature is aimed at.
Finally, the author discloses using an AI coding and writing assistant for the software and parts of the manuscript, and takes responsibility for the design, analysis, claims, code, and text.
References
- Original paper: Classic AI Scaffolding for LLM Social Agents
- Classic AI ancestry: Minsky (1974), A Framework for Representing Knowledge; Schank and Abelson (1977), Scripts, Plans, Goals and Understanding; Schank (1982), Dynamic Memory
- Sociology of encounters: Goffman (1974), Frame Analysis; Goffman (1983), The Interaction Order; March and Olsen (2011), the logic of appropriateness
- Agent architectures: Park et al. (2023), Generative Agents; Vezhnevets et al. (2023), Concordia; Tang et al. (2024), GenSim; Gao et al. (2024), survey of LLM-empowered agent-based modeling
- Simulation platforms: Yang et al. (2024), OASIS; Piao et al. (2025), AgentSociety; Puelma Touzel et al. (2025), SandboxSocial; Ren et al. (2026), SimWorld; Wang et al. (2023), Humanoid Agents
- Validation critiques: Puelma Touzel et al. (2026), the validation gap; Sarangi et al. (2026), EASE; Li and Tao (2026); Mooney et al. (2025)
- Related coverage in this series: Concordia (the Game Master this World Master extends); Generative Agents (memory, reflection, planning); AgentSociety and OASIS (the scale end of the same field)