#evaluation
11 posts
- Classic AI Scaffolding — The Agents Negotiated Brilliantly and Never Ordered Lunch
- Best Friends, Not Forever — Auditing What a Companion Keeps Over a Hundred Sessions
- Do AI Personas Grow? — Personas Move After Life Events, Rarely Where Humans Move
- PersonaEval — One Persona Population, Three Kinds of Application
- Plausible but Not Valid — A Gaussian Copula Is a Better Survey Respondent Than Any LLM
- Digital Pantheon — Auditing a Coalition Negotiation Clause by Clause
- From Agent-Based Modeling to Digital Twins — What a Simulation Is Allowed to Claim
- Anamnesis — Survey Simulation With Backstories, Behind a GUI
- ARCANE — Do Role-Playing Agents Stay in Character at the Right Time?
- LifeSide — Benchmarking Agents as Lifelong Digital Companions
- Should LLM Agents Decide in Social Simulations? — Finite-State vs. LLM Policies