AI YOU Town — A Digital Twin That Knows What It Doesn't Know
A personal-digital-twin framework that treats persona inference as sequential state estimation: 22 psychological dimensions updated by Bayesian posterior, wrapped in conformal prediction sets, and held steady over 100-turn dialogue by a periodically refreshed memory anchor.
Background
Three problems in personal-digital-twin work are usually solved apart from each other, and the paper’s argument is that they are the same problem seen from three sides.
Inferring personality from conversation needs uncertainty. A single-pass LLM prediction returns a number with no calibrated confidence, and a system downstream has no way to tell a well-evidenced trait from a guess. Evidence also arrives across turns, yet most pipelines score utterances independently and keep no persistent belief to update. And inference sits in a separate box from generation, so the understanding a system builds of a user rarely constrains how it speaks as that user.
The framing that ties them together is the distinction the paper draws between three kinds of persona. Demographic simulation assigns group-level attributes and asks whether outputs match a population’s distribution; it is cross-sectional by construction and its documented risk is caricature. Character simulation targets a fictional or public figure using stored knowledge, and its risk is hallucinating behavior when that knowledge runs out. Both treat traits as fixed. Individualized simulation — a personal digital twin — has to model longitudinal behavior, adaptive responses, and preferences that move. Inference and simulation close a loop: simulated interaction generates evidence, evidence updates the profile, the updated profile improves the next simulation.
Core Idea

The profile has three layers of its own. Demographics — gender, age, occupation, education, region — provide context and are treated as optional and uncertain throughout. The psychometric layer carries validated constructs: Big Five, adult attachment along anxiety and avoidance, self-efficacy, perceived loneliness, and positive and negative affect. These are the dimensions that get quantitative inference, evidence attribution, and uncertainty estimates. The third layer is MBTI, included because users recognize it, with the paper noting its contested psychometric standing in the same sentence.
Extending past the Big Five is a deliberate choice. The Big Five leaves out interpersonal structure like attachment style, wellbeing measures like self-efficacy and loneliness, and momentary emotional state — all of which matter for a companion system and none of which the five factors capture.
Method
Estimation, then updating
The estimator runs on prompting rather than a trained classifier, because the target schema is heterogeneous: some fields numeric, some categorical, some free text, and many missing outright in a short conversation. A Direct prompt instructs the model to emit null when evidence is insufficient, defines confidence 0 as no evidence and 1 as strong evidence, and bounds Big Five and interest scores to . A CoT variant inspects topics, style, values, goals, and trait cues before emitting the same JSON. When context allows, the estimator also runs a contrastive check — comparing an estimate from full recent context against one from minimal context, and down-weighting attributes that only survive under the richer prompt. This is what suppresses profile updates driven by a single unusual turn.
The updating is where the LLM stops deciding. For a numeric observation arriving with confidence , the observation variance is
and the posterior follows the standard precision-weighted Gaussian conjugate update:
A low-confidence observation carries high variance and therefore barely moves the mean. Categorical variables update discretely instead of averaging: on conflicting labels the system keeps whichever has stronger evidence. Conformal calibration at then turns point estimates into prediction sets, with nonconformity , calibration grouped by turn bucket and dimension, and a global fallback where a group has too little data.
Three monitors
Persona inference alone leaves a companion system unsafe over long horizons, so three auxiliary monitors run alongside. The affective monitor estimates current emotion, intensity, and short-term trajectory by combining a structured classifier with a history-aware transition model. The relationship monitor tracks sentiment, interaction status, and possible progression, holding uncertain readings as hypotheses until repeated evidence supports them. The risk monitor combines rule-based pattern matching with semantic analysis to catch manipulation, coercion, romance scams, and unhealthy dependency; past a threshold it can warn, restrict personalization, or demand verification.
The paper’s worked example is worth reading as a design brief. Given a user under financial stress and a twin proposing high-return investments, the monitors should jointly register an anxious high-intensity affective state, an asymmetric persuasion relationship, and elevated risk. A companion system that can be talked into a pig-butchering script is the failure this layer exists to prevent.
Memory, and the anchor
Three operational layers hold the history. Working memory is a FIFO buffer of up to 20 recent turns; every 10 turns the oldest block compresses into episodic memory, which stores summaries of events and emotional trends linked back to the original turns for traceability, with periodic consistency checks against the archive. Semantic memory reflects over recent episodes every 50 turns to extract higher-level patterns, and those enter the profile only when the evidence supports them.
Persona preservation rides on top. Instead of a static persona prompt, the system maintains a memory anchor holding identity, speaking style, inferred values, behavioral boundaries, and recent commitments. Every ten turns a dedicated writer rewrites the anchor from its current state, recent dialogue, and trait constraints. Generation then conditions on the refreshed anchor plus local context rather than the entire history, which is what keeps consistency from depending on an ever-growing context window.
Experiments
Six benchmarks cover the modules: PANDORA (2,415) and Essays (2,466) for Big Five, DailyDialog (564) for emotion, PsyScam (730) for manipulation detection, LoCoMo (1,542) for long-session memory QA, and PersonaConflicts (6,012) for relationship-conflict classification. Backbones span API models — GPT-5.4, Qwen-3.6-Max-Preview, DeepSeek-V4-Pro — and local OLMo-3-7B and OLMo-3.1-32B served through Ollama on H100s under an identical output schema.
Calibration carries the result. Point estimation improves everywhere and only slightly: PANDORA MAE moves 0.268 → 0.265 for GPT-5.4 and 0.278 → 0.269 for Qwen; Essays moves 0.451 → 0.438 and 0.469 → 0.448. The paper is candid that Direct prompting stays competitive here, since short texts often carry only a few explicit behavioral cues and one structured estimate captures most of them.
| Dataset / model | Method | MAE ↓ | ECE ↓ | Coverage ↑ |
|---|---|---|---|---|
| PANDORA / GPT-5.4 | Direct | 0.268 | 0.109 | — |
| PANDORA / GPT-5.4 | AI YOU | 0.265 | 0.071 | 0.976 |
| Essays / GPT-5.4 | Direct | 0.451 | 0.326 | — |
| Essays / GPT-5.4 | AI YOU | 0.438 | 0.214 | 0.944 |
| Essays / OLMo-3-7B | AI YOU | 0.454 | 0.251 | 0.921 |
| Essays / OLMo-3-7B | w/o Bayesian | 0.462 | 0.304 | 0.887 |
Calibration error falls for every dataset–backbone pair, often by more than a third, and conformal coverage clears the nominal 90% in every full-pipeline row. The ablations do the mechanistic work: removing Bayesian aggregation on OLMo-3-7B raises PANDORA MAE from 0.270 to 0.274 and drops coverage from 0.946 to 0.917, while removing conformal calibration eliminates the sets and raises ECE. CoT is not reliably better than Direct anywhere, which fits the paper’s reading that an extra reasoning step can amplify weak evidence as easily as it can surface strong evidence.
State monitors are backbone-dependent, and the paper says so. On DailyDialog, AI YOU beats Direct for all three API models (GPT-5.4 0.626 → 0.663, Qwen 0.582 → 0.638, DeepSeek 0.516 → 0.629), though CoT edges it out on GPT-5.4 at 0.676 — a gap the authors decline to call significant on a 564-utterance imbalanced pool. Removing dialogue context hurts every model, confirming that a target utterance often cannot be read alone. Removing intensity modeling barely touches accuracy and sometimes improves ECE, so intensity earns its place operationally rather than as a classification aid: it gives the coordinator a graded signal for how much validation or caution a response needs.
The risk monitor shows the same conditionality more sharply. It improves every API backbone on PsyScam — GPT-5.4 reaches 0.830 partial score, and removing the monitor lowers all three. On OLMo-3-7B, plain Direct prompting scores higher (0.472 against 0.451), and on OLMo-3.1-32B both CoT and the no-monitor variant beat the full pipeline. A model juggling a checklist, a multi-label ontology, and a JSON schema at once can lose more task signal to the scaffolding than the scaffolding returns.
Memory is the component that matters most. On a 300-instance LoCoMo diagnostic set, removing memory entirely produces by far the largest drop — 0.628 → 0.417 for GPT-5.4, 0.641 → 0.429 for Qwen. Removing only the semantic or only the episodic layer costs much less and the ordering flips by backbone. Retrieval is the driver; how the three layers should be weighted depends on the model and the question.
Persona preservation. Two settings test whether the same anchor stabilizes generation. Persistent Personas runs eight fictional roles over 100-turn dialogues under persona-directed and goal-oriented probes; a Werewolf game puts seven persona-conditioned agents through 100 turns of accusation and persuasion with distinct Big Five targets.
| Setting | Metric | Refresh | Static | Δ |
|---|---|---|---|---|
| Persistent Personas | Overall fidelity (1–5) ↑ | 4.48 | 3.61 | +0.87 |
| Persistent Personas | Style consistency ↑ | 4.38 | 3.47 | +0.91 |
| Werewolf / GPT-4o-mini | Big Five drift ↓ | 0.008 | 0.036 | −0.028 |
| Werewolf / Gemini-2.5-Flash | Big Five drift ↓ | 0.015 | 0.013 | +0.002 |
Refresh wins every Persistent Personas dimension and reduces adversarial drift for three of four models. Gemini-2.5-Flash is the exception, which the authors read as a backbone already stable enough at the instruction level to gain little from external anchoring.
Limitations
The authors name four, and the first is the one that governs the rest. These are module-level evaluations rather than a longitudinal study with real users, so nothing here speaks to trust, consent renewal, users correcting a wrong inference about themselves, or behavioral adaptation over months. The persona-preservation benchmarks use fictional roles and a simulated game — good instruments for measuring drift, silent on fidelity to an actual person. The profile is inferred from text, which is sparse and biased and ambiguous, so weakly supported attributes need to stay uncertain and user-controllable. And several comparisons rest on backbone JSON reliability: GPT-5.4 parses at 0.920 on LoCoMo and 0.940 on PersonaConflicts, DeepSeek at 0.910, and the paper marks rows below 0.98 as diagnostic.
One gap the paper does not list is the one in its own title. The economic layer — twins hired by real companies, Lily’s twin working as a school counselor at a price a small school can afford — is described in the system section and never evaluated. Trading is named explicitly as “an imaginative hook.” Read the paper as a calibrated user-state modeling framework with a town sketched around it.
The ethics statement is unusually direct about what a personal digital twin puts at risk: unauthorized impersonation, overconfident inference of sensitive traits, emotional dependency, and persona-conditioned agents used for persuasion. The design answers with null outputs under insufficient evidence and filtering of low-confidence state before generation, and the authors state that real deployment would additionally need consent, AI disclosure, user inspection and correction of inferred state, deletion, and a bar on high-stakes decisions.
References
- Original paper: AI YOU Town: Make Friends and Money with Your Digital Twin
- Online demo: quinnnnnne-ai-you.hf.space
- Conformal prediction primer: Angelopoulos and Bates (2021), arXiv:2107.07511
- Three-layer memory ancestry: Park et al. (2023), Generative Agents — covered earlier in this series
- Persona drift and identity instability: Choi et al. (2024), arXiv:2412.00804; Li et al. (2024), arXiv:2402.10962
- Persistent Personas benchmark: Luz de Araujo et al. (2026), EACL 2026, pp. 5329–5359
- Benchmarks: PANDORA (Gjurković et al. 2021); Essays (Pennebaker and King 1999); DailyDialog (Li et al. 2017); PsyScam (Ma et al. 2025); LoCoMo (Maharana et al. 2024); PersonaConflicts (Shen et al. 2025)
- Beyond the Big Five: Liu et al. (2025), arXiv:2511.03235