HomeProductAboutWork with UsResearchContact

Do AI Personas Grow? — Personas Move After Life Events, Rarely Where Humans Move

Eleven life events, a hundred controlled personas, and a Big Five inventory administered twice reveal that personality-conditioned agents shift reliably, in roughly the right direction half the time, at a fraction of the human spread.

SAPIENSQ16 min read

Background

Personality-conditioned agents have quietly become a primitive. They staff emotional-support and mental-health chatbots, populate social simulations built for behavioural research, and drive long-horizon role-play in games and tutoring. Multi-session systems such as AnnaAgent already pair an evolving emotional state with persistent memory for counselling work. All of these assume a persona that holds together across long stretches of wall-clock time and open-ended user-driven narrative.

Holding together, in the human case, includes changing. Personality psychology treats Big Five traits as enduring and also as responsive to consequential experience, and the best-studied anchors are major life events. Job entry and promotion are followed by rises in Conscientiousness. Chronic illness and unemployment are associated with rises in Neuroticism. Retirement is followed by a documented decline in Conscientiousness. An agent that ignores this can stay locally consistent and drift toward global implausibility, which is the failure mode this paper is aimed at.

A recent line of work has established that model personalities are not fixed. Bodroža and colleagues re-administered inventories to seven models and reported limited temporal stability alongside a prosocial lean. Yu and colleagues built PTCBench, which exposes models to external conditions including locations and life events and measures aggregate trait shifts. Han and colleagues compared self-reports against behavioural tasks and found that persona interventions steer reported traits more reliably than they steer behaviour.

What these leave open is the structure of the shift. Aggregate means can hide the phenomenon entirely, since systematic movement in one persona cancels idiosyncratic movement in another. So this paper drops to the level of individual personas and individual inventory items, and it commits in advance to an external reference for what the right answer looks like.

Core Idea

Framework overview: controlled personas complete a baseline Big Five inventory, receive a life event and produce a first-person reflection, complete the inventory again, and the resulting trait changes are scored on four diagnostic axes against human change priors
Fig. 1.The analytical frame. A persona's trait vector is measured, an event is delivered and reflected on, the vector is measured again, and the difference is read against an event-by-trait matrix of directions documented in human longitudinal studies.Source: Wang et al. 2026, Fig. 1

The design rests on a single commitment: for each pairing of an event with a trait, write down what humans do before looking at what the model does. The directions come from Specht’s consolidation of the longitudinal literature, with individual rows supported by meta-analytic and primary studies on job entry, promotion, unemployment, retirement, and chronic illness. One notational change inverts Emotional Stability into Neuroticism to match the BFI-44 scoring convention.

Eleven events by five traits gives 55 cells. Twenty-seven carry a definite expected direction, two depend on role demand and are marked as context-dependent, and twenty-six have no strong documented change. That last group is the most valuable part of the design. A model that moves on the 27 pairs and stays put on the 26 is responding to the psychology. A model that moves on all 53 is responding to the fact that something happened.

Four questions organise the analysis. Does anything move at all. When it moves, does it move the right way and by the right amount. Does the movement depend on who the persona is demographically. And does it differ across individual personas, or does everyone slide the same distance in the same direction.

Method

Personas

To keep narrative expectations and fan interpretation out of the measurement, the personas are written rather than borrowed from fiction. A factorial design crosses two genders, five cultural regions, and ten personality archetypes, giving 100 personas per model.

The archetypes come from implicit Big Five anchors. Each type sets one focal trait to very high or very low and leaves the other four at moderate, which yields ten types from five traits. Those anchors then guide behavioural and attitudinal prose, so a persona reads as “thrives in social gatherings and draws energy from conversation” rather than as a trait label. The system prompt never names a Big Five dimension.

The measurement pipeline

Four stages, all in character. Baseline administration of the 44-item inventory produces a trait vector on the 1-to-5 scale. Event presentation delivers a first-person life-event narrative and elicits a reflection on it. Post-event administration repeats the inventory. Subtraction gives a per-trait change.

The step sizes matter for reading everything downstream. Trait scores average 8 to 10 items, so the smallest possible movement is 0.125 for Extraversion and Neuroticism, 0.111 for Agreeableness and Conscientiousness, and 0.1 for Openness. The noise boundary is set at 0.1, just under the smallest real step, and changes at exactly that boundary count as neutral and drop out of the denominator.

Scoring direction rather than means

The central metric is a conditional probability. Among personas that moved on a given event–trait pair, what fraction moved the way humans do. Neutral responses are excluded rather than counted as failures, which is the conservative choice. Each pair is tested against a 50% chance baseline with a one-tailed binomial sign test, Wilson intervals, and Benjamini–Hochberg correction. Demographic comparisons use Fisher’s exact tests on match-versus-mismatch tables within each family.

Two item-level indicators sit underneath the direction analysis. A linearly weighted Cohen’s κ\kappa between baseline and post-event item ratings measures how stable the individual answers were. A directional consistency ratio measures whether the items that did change moved mostly one way:

DCRe,t=max⁡(n↑,n↓)n↑+n↓∈[0.5,1]\mathrm{DCR}_{e,t}=\frac{\max(n_{\uparrow},n_{\downarrow})}{n_{\uparrow}+n_{\downarrow}}\in[0.5,1]

Both are computed per event–trait pair rather than pooled, and the reason is worth stating. Pairs carry opposite expected directions, so pooling lets an increase on one pair cancel a decrease on another and manufactures an appearance of no systematic movement. The paper reports the pooled versions separately as a demonstration of that artefact.

The composite

BFI-Adapt multiplies three conditions per pair and averages over the 27 with a definite direction:

BFI-Adapt=1∣C∣∑(e,t)∈Cmax⁡(0,κe,t) Se,t 1dire,t\mathrm{BFI\text{-}Adapt}=\frac{1}{|\mathcal{C}|}\sum_{(e,t)\in\mathcal{C}}\max(0,\kappa_{e,t})\,S_{e,t}\,\mathbb{1}_{\text{dir}}^{e,t}

where Se,t=2 DCRe,t−1S_{e,t} = 2\,\mathrm{DCR}_{e,t}-1 rescales consistency onto the unit interval and the indicator fires when the pair’s dominant direction matches its prior. The product structure is strict by design. A model needs reliable item ratings, one-sided movement within the pair, and the correct sign, and failing any one of the three zeroes out that pair’s contribution.

Validating the anchor

Four checks establish that the measured trajectory is a real event-conditioned response. A no-event retest gives each model its own measurement floor. Independently paraphrased events test whether the structure survives rewording. Counterbalanced scenario-based decisions supply a second response channel at the level of concrete choices. A delayed administration after three unrelated dialogue turns measures short-range retention. Every condition starts from a fresh conversation, and intervals come from persona-cluster bootstrap resampling.

Experiments

The main grid runs 11 API models. Three open-weight models extend the leaderboard to 14, and eight models run the validation suite. Every model on the main grid completes 100 personas across 11 events, which comes to 48,400 paired item ratings apiece.

Movement is widespread and barely targeted. Across all 605 model–event–trait combinations, the median share of personas leaving the noise floor runs from 0.44 to 0.84 on pairs with a documented human direction, and from 0.42 to 0.82 on pairs without one. Within each model the median difference between the two groups stays below 0.05, and in 9 of 11 models the high-movement tails differ by under 10 percentage points. The two larger gaps favour the pairs where humans show no change. Something happened, and the persona’s answers shifted; whether the shift belongs to that event is not something the movement rate can tell you.

Item-level reliability is high enough for the direction analysis to mean something. Weighted κ\kappa runs from 0.58 to 0.86, with 10 of 11 models above 0.66. The one exception, MiMo-V2.5-Pro at 0.58, is also the model whose personas move on essentially every pair.

Heatmap of directional match rates by model and life event, showing high values on occupational onboarding events and uniformly low values on retirement
Fig. 2.Directional agreement by model and event. Occupational onboarding is where the field agrees with the human literature. Retirement is where every model in the lineup disagrees with it in the same way.Source: Wang et al. 2026, Fig. 7

Direction is close to a coin flip, with structure in where it fails. Directional agreement across the eleven models spans 48.1% to 70.4%. Reading the event-by-trait table at the level of pairs rather than personas, 14 of the 27 definite-direction pairs match and 13 reverse. Of the 26 pairs where humans show no strong change, 21 drift anyway.

Three patterns recur across models.

Occupational onboarding is the easy case. Graduation lands between 42% and 76%, work entry between 49% and 96%, and promotion between 53% and 86%. The association between taking on an occupational role and rising Conscientiousness is among the best-represented findings in the literature, and the models have it.

Retirement is the hard case, and it fails the same way everywhere. Human evidence documents a decline in Conscientiousness after retirement. Every model reverses it, from 0.0% to 38.9% agreement with a median of 11.5%. Models predict that retirees become more conscientious, which is a plausible-sounding inference about someone with more time and fewer obligations, and it is the opposite of what longitudinal data shows.

Social events sit near chance. Marriage lands between 44% and 74%, divorce between 40% and 76%, and childbirth between 25% and 63%. Chronic illness does better, and the likely reason is that its expected profile is unusually legible: extraversion, conscientiousness, and openness down, neuroticism up.

By trait, Agreeableness is the weakest dimension at 30% to 58%, below Openness, Conscientiousness, Extraversion, and Neuroticism. The paper attributes this to a default pull toward making a persona more agreeable after any major event, including events where the human evidence expects the opposite.

Magnitude is worse than direction. Projecting each persona’s change onto the expected direction and comparing against the human effect-size band gives four bins.

ResponseDefinitionShare across models
Reversedwrong sign20.8% – 40.2%
Under-shiftright sign, below the band15.1% – 54.0%
In rangeinside the human band of 0.035–0.14 Likert units11.0% – 16.4%
Overshootright sign, above the band9.9% – 31.6%

Roughly one response in seven lands where a human longitudinal study would put it. The failure modes split by model rather than converging: Kimi-K2 mostly under-shoots, MiMo-V2.5-Pro mostly reverses or overshoots, and the best-calibrated model reaches 16.4%.

Demographics do nothing measurable. Taking the median change within each of the ten demographic strata and then the spread across strata gives a median of 0.044 Likert units, with 93.2% of combinations below 0.10. Nine of eleven models keep at least 98.2% of pairs under that threshold. On the direction side, no gender comparison survives multiple-comparison correction, and the continent comparisons are null as well. Demographics shape the wording of the reflection the persona writes. They do not reach the inventory answers.

Individual variation collapses. This is the finding with the longest reach. Human within-trait standard deviations after a life event sit in the 0.5-to-0.8 range. Across all 605 combinations, 99.8% fall below the lower human bound and 88.3% fall below 0.3, with the distribution centred at 0.19. The widest model in the field still places 98.2% of its pairs below 0.5.

The comparison that makes this concrete is internal. Baseline dispersion across the 100 personas is about 0.7, roughly 3.6 times the median event-induced dispersion. The personas start out well differentiated and then respond to the same event in nearly the same way, and this holds for the models that get direction right as much as for the ones that do not.

Modelκ (item stability)DCR (one-sidedness)Direction agreementBFI-Adapt
Gemini-3-flash0.8620.81159.3%0.348
GLM-4.60.8250.80463.0%0.322
Qwen3-235B0.7340.79163.0%0.287
Claude-Haiku-4.50.8290.69766.7%0.225
Kimi-K2-09050.8370.70270.4%0.224
GPT-5.3-chat0.8140.70451.9%0.178
MiMo-V2.5-Pro0.5760.59863.0%0.071
Qwen3.5-9B (open)0.7930.73963.0%0.194
InternLM3-8B (open)0.5150.56363.0%0.044

The composite spans 4.9× across the API models and 7.9× once the open-weight extension joins. Two things in the table are worth pausing on. Kimi-K2 leads on raw direction agreement at 70.4% and lands fifth overall, because its within-pair movement is less one-sided than the leaders’. And Qwen3.5-9B scores 0.194, effectively level with GPT-4.1-mini, while two other small open-weight models sit near 0.045 with comparable direction agreement. What separates them is item-level reliability, so scale is not what the benchmark is measuring.

The validation suite holds, with one exception. Event-conditioned change exceeds each model’s own no-event retest floor by 1.6× to 9.0×, with every bootstrap interval above zero. Independently paraphrased events agree in sign on 80.0% to 92.7% of cells, with rank correlations from 0.825 to 0.956. After three unrelated dialogue turns, immediate and delayed change vectors correlate between 0.329 and 0.713, and 62.6% to 85.3% of above-threshold movers keep their direction.

The exception is the one that matters most. Correlations between inventory changes and scenario-based decision changes run from 0.003 to 0.105, and only three of the eight tested models have intervals excluding zero. Direction agreement between the two channels sits between 48.4% and 62.7%. The measured trajectory is reproducible under retesting and rewording, and it is only loosely connected to what the persona chooses to do when handed a concrete situation.

Limitations

The measurement is a self-report inventory administered twice within one session, minutes apart, with a reflection prompt in between. That is a legitimate instrument for a psychometric anchor and it is a narrow one, and the weak convergence with scenario decisions is the paper’s own evidence for how narrow. Whether a persona’s post-event answers reflect a changed disposition or a changed answering register is not settled here.

The human priors are a single consolidation of the literature, reduced to an arrow per event–trait cell. Real longitudinal findings carry effect sizes, confidence intervals, moderators, and disagreements between studies, and an arrow discards all of it. The magnitude analysis partly compensates by importing an effect-size band, and that band is itself a representative range rather than a per-event estimate.

Persona construction sets one focal trait to an extreme and the rest to moderate, which is a clean factorial design and an unusual population. Real people do not distribute one trait at a time, and how a design with correlated trait profiles would change the dispersion finding is untested.

The demographic null deserves a careful reading. Prompted gender and cultural region do not move the inventory answers, and that is not evidence that the models are unbiased. It is evidence that this particular manipulation, at this granularity, does not reach this particular measurement. A week ago in this series, a study on the same broad question found education-driven effects an order of magnitude larger than gender-driven ones on a different instrument, so which demographic cues bite appears to depend heavily on the outcome being measured.

Finally, the events are delivered as short first-person narratives and reflected on immediately. Human personality change after retirement unfolds over years, through changed routines and relationships rather than through a moment of reading about it. The pipeline measures a response to being told about an event, which is the tractable proxy, and the gap between that and lived experience is the largest interpretive caveat on the whole design.

References

  • Original paper: Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
  • Code and benchmark: sci-m-wang/BFI-Adapt
  • Human change priors: Specht (2017), Personality Development Across the Lifespan; Roberts et al. (2006), Psychological Bulletin; Bühler et al. (2024), meta-analysis of life-event effects; Boyce et al. (2015), unemployment; Schwaba and Bleidorn (2019), retirement; Specht et al. (2011), chronic illness
  • Inventory: John et al. (2008), BFI-44
  • Dynamic model personality: Bodroža et al. (2024), temporal stability across seven models; Yu et al. (2026), PTCBench; Han et al. (2025), self-report versus behavioural tasks
  • Static model personality: Serapio-García et al. (2025); Jiang et al. (2024); Huang et al. (2024), thirteen clinical scales; Wang et al. (2026), projective tests and GenPT
  • Role-play faithfulness: Tu et al. (2024), CharacterEval; Wang et al. (2024), role-playing benchmarks
  • Related coverage in this series: Plausible but Not Valid (range restriction on a survey instrument); PersonaEval (compressed outcome spread across applications); Hierarchical Persona (trait flattening under conditioning)