HomeProductAboutWork with UsResearchContact

Digital Pantheon — Auditing a Coalition Negotiation Clause by Clause

LLM agents tuned into faithfully partisan versions of Flemish parties negotiate a coalition agreement, and every clause in the result is traced back to the manifesto paragraph it came from — or flagged as invention.

SAPIENSQ13 min read

Background

Coalition formation is a good stress test for LLM agents because it resists the two standard modeling approaches. Game theory captures the structure of who can form a majority with whom, and agent-based models capture the bargaining dynamics. Neither reaches the level where the work actually happens: parties arguing in language, out of ideological traditions that do not reduce to a payoff matrix.

LLM agents can argue. What they cannot easily do is stay partisan. The chapter of failures the paper assembles is worth taking seriously as a set. Prompt-based personas degrade reasoning and lack domain grounding. Multi-agent debate produces silent agreement, where agents capitulate to the majority answer, and strong aggregate debate scores can sit on top of brittle internal dynamics. Benchmarked models drift toward dominant-party positions. All of these trace to the same RLHF-instilled pull toward being helpful and neutral, and all of them are fatal to a simulation whose whole point is that the participants disagree.

The political-agent literature has its own gap. The Political Actor Agent predicts roll-call votes at 91.8% accuracy but forecasts discrete choices rather than the bargaining that produces them. POLCA models coalition negotiation and benchmarks the outcome, leaving the negotiation itself a black box. PoliCon supplies a large consensus-drafting benchmark and diagnoses dominant-party bias, while adjudicating with an opaque LLM judge. Prior systems, in short, either predict outcomes without simulating deliberation or simulate deliberation without attributing the result to anything verifiable.

Core Idea

Party alignment pipeline: agentic chunking of manifestos feeds instruction data and preference data; a base Gemma model passes through supervised fine-tuning and DPO into an aligned model; a parallel embedding and vector-database branch supplies retrieval augmentation to the final agent
Fig. 1.The per-party alignment pipeline. Manifesto chunks populate both the tuning datasets and the retrieval index. SFT adapts register, DPO steers the persona, and a party-specific vector collection supplies the factual boundary at inference time.Source: Van Mulders et al. 2026, Fig. 1

The design separates two things that prompting conflates. Persona — how combatively and in whose voice the agent argues — is written into the weights. Policy content — what the agent is allowed to claim its party wants — comes from retrieval against that party’s own manifesto and nowhere else.

SFT alone does not achieve the first. Cross-entropy fine-tuning raises the likelihood of the target text and leaves the base model’s ingrained prior for encyclopedic neutrality intact underneath, so adversarial prompting pulls the model back to it. DPO’s contrastive objective is what pushes probability mass away from the neutral and opposing responses, penalizing the encyclopedic register directly rather than layering a style on top of it.

The result is a stance the authors describe as the “Method Actor”: the model speaks as the party — “we will…” — instead of describing it from outside as “the party proposes…”.

Method

Chunking the manifestos

Hybrid chunking pipeline: raw manifesto PDFs pass through pre-processing steps, then a Stage 2 hybrid procedure where a deterministic Hard-Gate detects structural boundaries and a Gemma-3-27B Soft-Gate judges topic transitions, before finalize and validate
Fig. 2.The Agentic–Structural Hybrid chunking procedure. A deterministic Hard-Gate splits on formal section markers; the Soft-Gate fires only on residual transitions, asking a 27B model whether the trailing context and the next block share a topic. Deterministic logic where the structure is unambiguous, a model call where it is not.Source: Van Mulders et al. 2026, Fig. 2

Chunk quality is load-bearing here, because the same chunks become both the retrieval index and the training data, and a clause’s provenance is only auditable if the chunk it came from is a well-defined node in the document tree.

PDFs parse through PyMuPDF at the dictionary level for per-span bounding boxes and font sizes, which lets margin and sub-minimum-font blocks go without a layout model. Two cascaded gates then decide boundaries. The Hard-Gate is regex over formal section openings — numeric subsections, Article and Chapter markers, capitalized colon-terminated headers — and forces a split, which compensates for models’ structural blindness and guarantees that two lexically similar subsections never merge. The Soft-Gate fires only where the Hard-Gate found nothing: Gemma-3 27B at temperature 0.1 sees the trailing 800 characters plus the candidate block and returns a same-topic judgment. Running a 27B model on every boundary would be the cost bottleneck; running it only on ambiguous ones is what makes the procedure affordable.

Preference triplets come out of the same chunks. Each chunk over fifteen tokens goes to Gemma-3 27B at temperature 0.7 with a schema enforcing (Prompt, Chosen, Rejected) triples in Dutch. The Prompt is a citizen question leading to the chunk’s policy resolution; the Chosen answer restates the resolution in first-person assistant register; the Rejected answer is well written and either argues from an opposing ideology or offers a sterile non-committal stance. One artifact serves both stages — SFT takes Prompt → Chosen, DPO takes the full triple.

Alignment and grounding

Gemma-3 27B, quantized to 4-bit via QLoRA on Apple MLX, takes both stages. SFT runs at batch 2 for 5 epochs; DPO at batch 1 for 3 epochs with β=0.1\beta = 0.1 and a learning rate of 5×10−65 \times 10^{-6} resuming the SFT adapter, with epoch counts chosen to avoid catastrophic forgetting. Adapter steps scale as I=NB×EI = \frac{N}{B} \times E so that manifestos of different lengths receive comparable alignment depth.

Retrieval supplies the factual boundary. Each party gets its own ChromaDB collection in a local SQLite store — isolated collections, so no party can retrieve another’s platform. Embeddings come from a multilingual MiniLM served on CPU, which gives each agent complete recall of its manifesto without the memory cost.

The negotiation arena

The arena is hub-and-spoke: party agents as spokes, a formateur at the hub. Following convention, the largest party takes the formateur role, and here that means the formateur is instantiated from the same fine-tuned model as N-VA, system-prompted to broker neutrally. It receives the Banzhaf power indices of all parties to keep leverage realistic — the authors report that an untuned base formateur favored moderate parties even when given those indices, while a party-tuned one respected them.

Each turn, the active agent runs semantic search against its own collection and receives up to five chunks under a cosine-distance threshold, injected as strict ideological boundaries so it cannot concede a core promise or invent an out-of-line compromise. Agents state red lines, identify shared objectives, and propose syntheses; the formateur drafts an interim agreement and refines it across four rounds. Adapter hot-swapping keeps the base weights resident while party adapters load per turn, which is what lets all major parties run on consumer hardware.

Four rounds is a design choice rather than a measured convergence criterion — the point in preliminary runs where priorities and common ground had surfaced and further rounds added little.

MILT and the influence score

Rather than parse the transcripts, the evaluation models the negotiation as a temporal directed graph and traces backward. Every measure in the final agreement (L3) goes back through the agents’ opening standpoints (L1) to the retrieved manifesto chunks (L0). A Qwen3.6 27B classifier ingests all three levels in one call, which lets it record not only whether an idea mutated but when — separating preemptive dilution (L0 → L1) from reactive concession (L1 → L3).

Five provenance states result:

  1. Direct Lineage — maps to a manifesto and survives intact to the final agreement.
  2. Diluted Lineage — originates in a manifesto and is measurably weakened, before or during negotiation.
  3. Synthesized Lineage — combines competing manifesto chunks into a hybrid produced by the interaction.
  4. Pipeline Artifact — present in L1 and L3 but absent from L0, exposing prompt-induced setup noise.
  5. Orphan — appears in the final agreement with no grounding anywhere, a debate-phase hallucination.

The Coalition Influence Score turns these into attributable power, weighting Direct at +3, Diluted at +2, Synthesized at +1, and the two error states at 0:

CIS(π)=∑p : π(p)=πw(τ(p))\text{CIS}(\pi) = \sum_{p\,:\,\pi(p)=\pi} w(\tau(p))

The gradient encodes degrees of ideological victory — undiluted dominance, agenda-setting despite concessions, anchoring a principle inside a compromise. Zero-weighting the error classes keeps hallucinated provisions from inflating anyone’s score. The authors are explicit that 3/2/1 is an ordinal encoding rather than a calibrated magnitude, and that any monotonically decreasing weighting preserves the ordering.

A separate grounding pass grades every L3 measure against the historically adopted 2019 agreement as Present, Partially Present, or Absent, summarized per dossier by G=(npresent+0.5 npartial)/NG = (n_{\text{present}} + 0.5\,n_{\text{partial}})/N. This pass grades all five classes, including the error ones, so the question of whether lineage predicts materialization stays open to the data rather than assumed.

Experiments

Three runs cover 22 policy dossiers, each evaluated twice, for 2,629 classified provenance points at roughly 876 per run.

MILT stateMean countShareGrounding G
Direct Lineage (τ1)383.3 ± 9.143.7%0.333
Diluted Lineage (τ2)119.7 ± 5.513.7%0.256
Synthesized Lineage (τ3)120.0 ± 9.513.7%0.365
Pipeline Artifact (τ4)59.7 ± 4.06.8%0.232
Orphan / hallucination (τ5)193.7 ± 18.522.1%0.143
Ideological retention (τ1+τ2)503.0 ± 3.657.4%0.315
Systemic error (τ4+τ5)253.3 ± 18.928.9%0.164

The central finding sits in the right-hand column. Provisions with genuine manifesto provenance realize at G = 0.32–0.37; the two error classes collapse to 0.164, and Orphans are Absent in 75.4% of cases. Internal lineage, computable the moment the simulation finishes, predicts external materialization. One wrinkle cuts against the obvious reading: Synthesized clauses realize slightly better than Direct ones (0.365 against 0.333), so genuine cross-party compromise survives contact with reality at least as well as unilaterally retained demands.

The winner. N-VA takes every run, and the ranking never changes: N-VA 40.0% ± 3.3% of attributable influence, CD&V 31.8% ± 1.8%, Open Vld 28.2% ± 1.8%, with non-overlapping means. That is more balanced than the historical intra-coalition seat split, where N-VA held 50.0% of the 70 coalition seats. Composition is as stable as magnitude and more informative: about three-quarters of N-VA’s score is Direct Lineage every run, Open Vld draws most of its weight from Synthesized compromises where it is most often named, and CD&V sits between with an even split. The simulation reproduces a recognizable coalition signature. The largest party sets the agenda, the liberal junior partner brokers consensus, the centrist sits in between. N-VA’s advantage also concentrates in the class most likely to materialize.

Against reality. Pooled across runs, provisions are Present in 9.9 ± 2.0%, Partial in 35.9 ± 0.9%, and Absent in 54.2 ± 1.5%, giving G = 0.278 ± 0.017. Realization varies sharply by dossier, from Education at 0.491 down to Finance and Budget at 0.147, and the pattern follows the level of abstraction rather than the subject. Full presence concentrates in programmatic dossiers, and even there Partial is the modal state, indicating thematic correspondence over literal adoption. The worst-realized dossiers are saturated with quantified fiscal commitments or competences shared with the federal level — Brussels and the Flemish Periphery records no fully grounded proposal at all.

The diagnosis of the Absent category is the most useful sentence in the results. What fails is fabricated operational detail — invented budgets, percentages, deadlines, institutions — rather than thematic reasoning. Most Partial grades identify the right commitment and then append a specification the real agreement never contained. The paper offers a clean illustration on both sides: a guarantee against “forced school mergers” traces verbatim to an N-VA manifesto breakpoint and survives to the final text, while an invented “1.15% of GNP for education” target is flagged Orphan and is absent from reality.

Reproducibility splits along the same seam. Grounding proportions barely move across runs, while raw proposal counts per dossier swing widely — Poverty Alleviation produced 53, 28, and 14 distinct proposals in the three runs. Those swings concentrate in the ungrounded topologies, which is why the authors base every claim on proportions rather than volumes.

Limitations

The largest gap is one the paper names as future work: there is no component ablation. The marginal contribution of SFT, DPO, and RAG against a prompt-only or non-tuned baseline goes unmeasured. That leaves the architecture’s central claim — parameter-level alignment succeeding where prompting fails — resting on the prior literature and on the authors’ preliminary observation about the base formateur, with no controlled comparison in this paper.

The formateur is a second open question. It runs on the largest party’s own fine-tuned adapter, and the largest party wins every run. The authors flag the consistency (“N-VA being both largest party and formateur”) without an architecture that separates the two, and alternative formateur instantiations appear in the future-work list. Until that ablation exists, the winner is confounded with the arbitration.

Smaller caveats accumulate. The CIS weights every clause equally regardless of budgetary or political stakes, so Education and Animal Welfare count the same. The simulation is closed, with no socio-economic, federal, media-salience, or electorate-feedback channel. Realization is conditioned on a single referent, the 2019 Flemish agreement, which leaves open whether the lineage typology is a property of proportional-representation coalition politics or of this particular party system. The three-state grounding scheme records genuine contradiction as absence, so a proposal that inverts the real agreement scores the same as one that never appears. And the NLI labels are automated, with human expert validation of a sample listed as future hardening.

References

  • Original paper: Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
  • Coalition agreement referent: Vlaamse Regering (2019), Regeerakkoord van de Vlaamse Regering 2019–2024
  • Political agents: Li et al. (2024), Political Actor Agent; Moghimifar et al. (2024), POLCA; Zhang et al. (2026), PoliCon; Liang et al. (2026), UNBench, AAAI 2026
  • Parameter-level ideological steering: Agiza, Mostagir, and Reda (2025), PoliTune, AAAI/ACM AIES
  • Debate failure modes: Cui et al. (2025), Free-MAD (silent agreement); Smit et al. (2024), Should we be going MAD?, ICML 2024
  • Process versus outcome evaluation: Gritta et al. (2026), Process Evaluation for Agentic Systems, EACL Findings 2026
  • DPO: Rafailov et al. (2023), arXiv:2305.18290
  • Banzhaf/Penrose power index: Penrose (1946), JRSS 109(1)
  • Related coverage in this series: Agentopia (DPO-adjacent rejection sampling on life trajectories); ARCANE (DPO on temporally displaced negatives)