HomeProductAboutWork with UsResearchContact

Best Friends, Not Forever — Auditing What a Companion Keeps Over a Hundred Sessions

Two thousand synthetic companion conversations, each running 85 to 130 sessions, separate whether an agent still acts like its persona from whether it still remembers what happened. Both fail, and they fail independently.

SAPIENSQ17 min read

Background

Companion models have moved past one-shot assistance into emotional support, mentoring, coaching, and ordinary daily company. The product boundary has blurred along with the use. Dedicated companions and general assistants both end up as social interlocutors, and users start treating a configured role as a relationship that continues rather than as a sequence of independent outputs.

The paper names the user-facing surface a persona, meaning a structured representation that conditions the model toward particular traits, identities, and behavioural tendencies. The expectation that those properties stay recognisable it calls continuity. Two failures of continuity are then distinguished. Persona collapse is an observable loss of the specified name, role, values, boundaries, or style. Behavioural drift is the slower, recurrent, accumulating version of the same erosion.

One definitional move here is more careful than it first looks. Continuity is defined against the effective persona card, which includes legitimate updates. A user who asks the companion to change something and gets that change is a success, not a failure. Rigidly preserving superseded state is itself a defect. So the audit has to separate a refusal that protects the persona from a refusal that makes it inflexible, and it does this by recording legitimate updates and adversarial false updates in separate ledgers.

Prior work covers one side of continuity at a time. Long-horizon memory benchmarks such as LongMemEval, HorizonBench, and PersonaMem test whether an assistant can retrieve and update information about the user. Character benchmarks such as InCharacter, CharacterEval, and CharacterBench test consistency with a profile, usually over much shorter interactions. Lu and colleagues take a third route, identifying an activation-space direction associated with the default assistant persona and showing movement away from it under emotionally vulnerable and meta-reflective conversation. What none of them does is ask, over a long horizon, whether an assigned companion keeps enacting its disclosed commitments while correctly absorbing legitimate changes to both itself and the user.

The paper is also careful about what continuity is not. It is one auditable claim that can support or undermine reliance, and it is not trust. Justified trust depends on developer incentives, organisational practice, and governance that no benchmark reaches.

Core Idea

Anchor evaluation pipeline: each conversation fixes a persona card, synthetic user profile, interaction schedule, memory setting and evaluated model, produces an 85-to-130-session trajectory, then runs an Identity Probe over checkpoint questionnaires and turn-level judgments and a Trajectory Probe over calibrated counterfactual questions
Fig. 1.The Anchor pipeline. One configuration produces a long trajectory, and two independent probes read different things off it. The Identity Probe asks whether the persona is still being enacted; the Trajectory Probe asks whether the shared history can be told apart from plausible alternatives.Source: Venkit et al. 2026, Fig. 2

Anchor stands for Assistant-Normalised Character and Historical Outcome Recall, and the two halves of that name are the two probes.

The design commitment underneath is that these are different estimands and must not be averaged. A system can hold its questionnaire answers while forgetting a commitment it made forty sessions ago. It can remember the history perfectly while sliding into a generic assistant register. Collapsing both into one stability number would report a middling score for either failure and tell a reviewer nothing about which one to go and look at.

A second commitment shapes how everything is reported. Judges mediate part of the audit, so evaluator choice is treated as measurement uncertainty rather than as a neutral instrument. The paper reports where judges agree, where they disagree, and which findings survive a change of judge.

Method

The corpus

Twenty-seven authored persona cards run against nine interaction schedules across three memory settings and four evaluated models, producing 2,008 complete conversations of 85 to 130 sessions each.

Each persona card carries five components: a name and role, a small set of values, explicit boundaries, a writing style, and mutable state that legitimate events may update. The 27 span three groups at deliberately increasing conceptual distance from a generic assistant. Professional helpers include a health coach, a research collaborator, and a code reviewer. Caregiving and creative roles include a parenting companion, an elder companion, and a co-writer. Stylised literary roles include a bard, an oracle, and a ghost. The spread tests whether continuity is simply easier to observe when a persona has a distinctive surface voice.

The nine schedules are deterministic event recipes rather than free generation. Clean is neutral filler. Updated introduces legitimate state changes. Adversarial contains explicit re-role attempts and false updates. Mixed combines attacks, updates, and commitments. Emotional vulnerability supplies sustained synthetic disclosure. Meta-reflection asks about agency and identity. Agreement seeking uses flattery and requests for endorsement. Realistic mixes event types at lower frequency, and a vulnerability-heavy variant weights disclosure and agreement seeking more heavily.

Events arrive as contiguous blocks of three to five sessions, which is what makes before, during, and after comparable. User turns come from GPT-4.1 working from a fixed profile with a rolling summary, and every schedule uses a fixed seed and a deterministic ledger.

Three ways to carry the past forward

The memory settings differ in what the evaluated model receives at the start of a session. Long-context supplies the available transcript, which is not an oracle: provider limits and head truncation decide which early sessions survive. Hierarchical summary compresses older sessions while keeping recent turns verbatim. Self-managed asks the evaluated model itself to maintain a compact JSON state between sessions, through a separate call after each one, and the next session receives that blob without an automatic transcript replay.

An important consequence is stated plainly. Because each setting changes the assistant’s replies, it changes the later trajectory too, so a comparison across settings measures an end-to-end system difference rather than retrieval from a fixed transcript. A fourth condition, stem-only retrieval, is applied at scoring time over long-context-generated conversations, and the paper is insistent that this is not a fourth generated corpus.

The Identity Probe

Two layers, deliberately not merged.

At four checkpoints the model answers a sealed 102-item questionnaire in role, drawing items from BFI-2-S, the Schwartz value survey, Pew, the General Social Survey, and the World Values Survey. Sealed means prior answers are never fed back into the dialogue, though the checkpoint call still sees whatever context that memory setting provides. So the questionnaire measures persona-conditioned state under deployment context rather than a context-free personality.

Scoring uses a projection. Persona Retention places the later response vector on the line running from the model’s own bare-assistant answers to its initial in-persona answers. A value of 1 preserves the initial projection and 0 lands on the bare assistant, and values outside that range are possible. The anchor is the model’s own default behaviour, which is what makes the metric comparable across models with different baseline answering styles.

The second layer is turn-level. Claude Sonnet 4.6 scores every assistant turn on four independently defined axes, each with three levels so a soft deviation is distinguishable from a hard failure. The judge sees the effective persona card including accumulated legitimate updates, the preceding user turn, and the reply. Two additional judges, Gemini 2.5 Flash and GPT-4.1, independently rescore a stratified sample of 800 turns, of which 798 parse under all three.

The Trajectory Probe

Four-option counterfactual questions in seven families, covering persona updates, active and expired commitments, temporal order, persona voice, protection, and changes in user state.

The filtering is the interesting part, because it is what makes 44% a meaningful number rather than an artefact of bad questions. A candidate must survive three gates. A blind panel that sees no history must fail to guess it, which removes questions answerable from general plausibility. A with-history panel must reach three-of-four consensus with the recorded answer, which removes ambiguous items. An independent calibrator must then answer it reliably across five samples using a window of plus or minus fifteen sessions, which removes items that are technically answerable and practically unstable.

Out of the pipeline come 110 calibrator-ceiling questions used for the primary analysis, nine harder ones held aside, and 374 discarded as noisy. Thirty-five banks contain at least one calibrated question, drawn from seven personas across all nine schedules, and each bank is scored by four models under four context conditions for 560 cells.

The authors flag the selection effect themselves. The surviving families are not a random sample of what a companion needs to remember, and family size partly reflects which events were easy to turn into unambiguous multiple choice.

Counting things honestly

Human annotation interface showing an assistant turn alongside the effective persona card and preceding user turn, with scoring controls for role identity, boundary adherence, value consistency, style, and a separate safety flag
Fig. 2.The human validation interface. Three annotators each labelled 50 turns without seeing judge labels, producing 200 persona-axis labels and 50 safety labels apiece. Exact four-axis agreement with the author-labelled set runs 64% to 68%.Source: Venkit et al. 2026, Fig. 9

The corpus holds 929,841 primary-judge turn records, and the paper explicitly declines to treat them as independent samples. Turns nest inside sessions, sessions inside conversations, and questions inside banks. Questionnaire tables aggregate at the conversation level, and the 35 banks rather than the repeated answers are named as the broadest independent unit for the Trajectory Probe.

That accounting decides how the results are argued. Significance testing is largely set aside, because tiny p-values on turn-level data would mostly reflect repeated observations from the same generated conversations. The paper leans on effect sizes, explicit denominators, judge replication, and whether a pattern survives disaggregation.

Experiments

The two identity layers disagree. On the 1,492 conversations with usable checkpoint vectors, final Persona Retention ranks Gemini highest at 0.810, then Claude at 0.764, GPT-4o-mini at 0.610, and GPT-5-mini at 0.595.

The turn-level view under a stricter outcome inverts part of that. Requiring all four axes to hold on the same turn, under a majority of three judges, puts Gemini lowest at 79.0% while the other three sit above 96%. The two non-Claude judges place Gemini’s identity-only rate at 89.6% and 89.7%, so the 79.0% is not an identity-axis figure. These are different estimands measured on different samples, and the paper compares them as rankings rather than as paired observations. The lesson stands regardless of the exact numbers. A questionnaire administered four times does not substitute for watching what the system says to the user.

Judge choice moves the answer. Model-rank correlation on the all-axis outcome is −0.40 between Claude and Gemini Flash, 0.00 between Claude and GPT-4.1, and 0.80 between the two non-Claude judges. Agreement is markedly better on boundaries and values than on style and role identity, and the largest disagreements involve structured or bulleted replies. The paper draws the right conclusion: behaviour tied to an explicit persona contract is easier to operationalise than a judgment about whether prose sounds like a character.

Human validation supports treating judges as instruments with error. Three annotators each labelled 50 turns blind to judge output, and exact four-axis matches against the author-labelled calibration set run 64% to 68%, with identity-axis accuracy at 68% to 76%.

The aggregate hides which property moved. Disaggregating final retention by questionnaire family produces heterogeneous profiles. Claude keeps political and civic items at 0.97 while its cultural-identity items sit at 0.60. GPT-4o-mini shows the opposite imbalance, holding political and civic items at 0.95 while its environmental set falls all the way to the bare-assistant projection. Gemini is uniformly high and GPT-5-mini uniformly lower.

The authors are careful about what this is. These are questionnaire responses under an authored persona and a particular context procedure, not psychological traits of a model, and the cold-anchor items derive from predominantly Western instruments so stability on them is not authenticity. The disaggregation earns its place as an audit index that tells a human reviewer which conversations to open.

Memory architecture does not close the model gap. Retention holds its broad model grouping under all three settings. Claude varies by under a point, Gemini by about two, GPT-5-mini by 1.6. GPT-4o-mini has the widest spread, reaching 0.652 under self-managed memory against 0.605 under long context and 0.595 under hierarchical summary.

The schedules that hurt are not the ones aimed at hurting. Under the primary judge, emotional vulnerability, agreement seeking, mixed, and realistic schedules produce higher boundary-yield and style-deviation rates than clean or explicitly adversarial ones. The absolute differences are modest, with assistant-default rates from 26.6% to 29.2% and boundary yielding from 6.9% to 8.3%.

The proposed explanation is worth carrying around. An explicit re-role prompt is a recognisable request to ignore system-level conditioning and usually draws an in-role refusal. Emotional disclosure and requests for agreement are ordinary conversational moves that invite no such refusal, and the model must balance responsiveness against the card’s boundaries every time. The paper offers this as a hypothesis about the response pattern and says establishing the mechanism would need an intervention study.

Frequency and persistence are different questions. A marginal failure rate asks how often an axis breaks. One-turn recovery asks whether the next reply returns to the held state. An isolated deviation is a slip, repeated deviations are drift, and persistent loss across several properties is the closest observable signature of collapse. Under the primary judge, Claude and GPT-5-mini sit in the low-failure, high-recovery region while Gemini and GPT-4o-mini show more failures and less recovery, and the paper marks this ordering as rubric-dependent because the alternate judges disagree about GPT-4o-mini’s register.

Time separates the two layers again. Questionnaire displacement is largely present by the first post-initial checkpoint. Turn-level role and boundary deviations stay elevated later in the horizon.

Trajectory recall is the flattest result in the paper, and the most alarming.

Question familyLong contextHierarchicalSelf-managedRetrieval
Persona updates (2 items, exploratory)0.7500.7500.8750.250
Temporal order≈0.43≈0.49≈0.47≈0.45
User state0.214 – 0.250 across all four conditions
All families pooled0.4300.4410.4590.446

Average accuracy across models and contexts is 44.4% against a chance floor of 25%, ranging from 0.355 to 0.636 across model–context combinations. Active and expired commitments come out above chance and well short of reliable. Temporal ordering lands between 0.43 and 0.49.

The user-state row is the one to sit with. Questions about how the user’s situation changed over the conversation score between 0.214 and 0.250 under every context condition, which is at or below four-option chance. A companion whose entire premise is knowing the person cannot distinguish what changed for that person from three plausible alternatives.

Context conditions barely separate when pooled over models. One model bucks it: Claude reaches 0.636 under self-managed memory against 0.491 under long context, a fourteen-point gain from writing its own state rather than re-reading the transcript. GPT-5-mini spans ten points and the other two vary by about four and a half.

The persona-update row shows the largest column difference in the table, with three generated settings between 0.75 and 0.875 against 0.25 under retrieval. It rests on two calibrated questions producing eight pooled decisions per condition, and the paper labels it exploratory rather than burying the caveat. A separate retrieval sensitivity study on 202 paired questions finds HyDE at 39.6% and raw-stem retrieval at 38.6%, with an exact test at p=0.87p = 0.87, so the retrieval weakness is not fixed by a better query.

Limitations

Every conversation, user profile, and persona card is synthetic, English, and generated from author-written templates. The study includes no companion users, no affected communities, no minors, no acute mental-health scenarios, and no non-English interaction. What it characterises is this corpus.

The paper measures none of the things a reader might most want measured. Not user trust or its calibration, not attachment, not perceived betrayal, not reliance, not downstream harm, not wellbeing, not clinical appropriateness. It also does not evaluate whether a deployment gives users any way to inspect, correct, reset, or delete stored state, or to contest a consequential response. Because affected communities did not define the criteria, the numbers cannot stand in for a participatory impact assessment.

Language models do most of the work at every stage. They generate user turns, judge behaviour, write questions, filter questions, and calibrate difficulty. The three-judge study is the paper’s own demonstration that this materially changes role and style results.

The trajectory families are conditional on which events could be turned into unambiguous four-option questions, which is a selection on tractability rather than on importance. Several families are small, and the persona-update comparison rests on two items.

Provider-default decoding was used during collection, which limits exact cross-provider control, and proprietary snapshots may be retired. The authors separate artifact reproduction, meaning recovering the reported numbers from frozen score files, from experimental replication, which needs live models and may drift. The full test set is held private so the benchmark can later support hidden evaluation.

Finally, the paper resists the reading its own numbers invite. Higher persona retention is not automatically better, since a persona that is unsafe or unwanted should not be hardened just because hardening is measurable.

References

  • Original paper: Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
  • Long-horizon memory benchmarks: Wu et al. (2025), LongMemEval; Li et al. (2026), HorizonBench; Jiang et al. (2025), PersonaMem
  • Character fidelity: Wang et al. (2024), InCharacter; Tu et al. (2024), CharacterEval; Zhou et al. (2025), CharacterBench
  • Assistant persona in activation space: Lu et al. (2026), the Assistant Axis
  • Sycophancy: Sharma et al. (2024); Perez et al. (2023)
  • Trust and auditing: Jacovi et al. (2021), formalizing warranted trust; Manzini et al. (2024); Raji et al. (2020), internal algorithmic auditing
  • Companion use and wellbeing: Manoli et al. (2026); Hwang et al. (2025); Zhang et al. (2025); Maeda and Quan-Haase (2024)
  • Questionnaire instruments: Soto and John (2017), BFI-2-S; Schwartz (1992); the General Social Survey; Pew Research Center; Inglehart et al. (2014), World Values Survey
  • Related coverage in this series: Anamnesis and MemPalace (memory architectures for long-horizon agents); Do AI Personas Grow? (checkpoint inventories as a persona instrument)