Diversity
The range of users, languages, cultures, goals, tasks, and scenarios represented, so simulation spans the target population rather than a narrow slice.
NeurIPS 2026 · Paris · 12 or 13 December
Grounded User Simulation for Model Evaluation and Training: Diversity, Fidelity, and Validity
Simulated users now drive AI evaluation and training, but when can we trust them?
Making user simulation diverse, faithful, and valid.
Organized by Yev Meyer NVIDIA Krisztian Balog University of Stavanger & Google DeepMind Flora Salim UNSW Sydney Victor Barres Mercor Preethi Seshadri UC Irvine & Cohere
The mission
The opportunity is to make user-simulation methods open, validated, and shared across communities. We are bringing together the people who build simulators, the people who depend on them, and the people who test where they fail, helping establish the evidence, standards, and open tools the field can build on.
Why now
User simulation models how people, and increasingly agents acting on their behalf, behave and respond as they interact with AI systems, products, services, and environments. Simulated users now appear in agentic training pipelines, conversational-search evaluation, dialogue systems, recommender systems, and tool-use benchmarks.
Done well, simulation provides scalable and reproducible interaction data, broader coverage of rare and long-tail behavior, finer control over demographic and behavioral variation, and a path to evaluating multilingual and sovereign-AI systems with little in-language production traffic.
But adoption is outpacing validation. In Lost in Simulation, measured agent success shifted by as much as nine points depending on which model played the user. The study also found that overly cooperative simulators can inflate performance and that simulators can be miscalibrated for speakers of dialects including African American Vernacular English and Indian English. Simulator choice is therefore part of the measurement, not a neutral implementation detail.
The challenge is broader than selecting a better model. UserBench models users who reveal preferences incrementally, while τ2-bench represents settings in which users share control of the environment. These designs highlight what realistic simulation must account for: goals that evolve over time, interaction histories, other actors, and the social, digital, and physical environments that constrain behavior. If the simulator is a measurement instrument, both the user and the conditions shaping their behavior need validation.
At the same time, grounded population-scale simulation is newly feasible. The field has an opportunity to establish shared evidence and reporting norms now, before weak assumptions harden into standard practice.
Who it’s for
Relevant work is scattered across labs and communities that rarely meet, often using different terminology, assumptions, and evaluation practices. Without shared, validated simulators, datasets, and reporting norms, results remain difficult to compare and the field cannot build cumulatively. We are convening these communities around a shared challenge: representing users, populations, and the contexts that shape their behavior.
If you build simulators, validate or audit them, or use them to train and evaluate models, this workshop is for you. Work that stress-tests simulation is as welcome as work that advances it. Negative results, replications, and evidence that a simulator fails to transfer are first-class contributions.
Three pillars
We organize the agenda around three pillars, covering the two strategically important uses of simulation: evaluation of interactive and agentic systems, and training via SFT data, preference and process-reward signals, and simulators as RL environments.
The range of users, languages, cultures, goals, tasks, and scenarios represented, so simulation spans the target population rather than a narrow slice.
Whether simulated users match real individual and population-level behavior, including style, strategy, natural variation, and the tendency of real users to be incomplete, inconsistent, impatient, or strategic rather than idealized and cooperative.
Whether conclusions drawn from simulation transfer to reality: do simulated evaluations predict real outcomes, and is the training signal trustworthy?
The centerpiece
People do not act as isolated personas, and no single user stands for a population. A simulation must decide what it represents: individual goals and histories, behavioral variation and population distributions, relationships and other actors, and the social, digital, and physical environments that shape interaction.
This cross-community panel will examine which components matter for different uses, how they should be grounded and evaluated, and how requirements change across evaluation, synthetic data, and training.
On the day
A single full day, provisionally 08:00–17:00. The program combines a keynote and invited talks, contributed spotlights, two poster sessions, the cross-community panel, and a closing breakout to advance draft reporting guidance and the benchmark catalog.
The detailed schedule will be published once all speakers are confirmed. Every session is designed to move from evidence to action: what user simulation needs to capture, where methods fail, and what the community should investigate and report next.
Program
Additional invited speakers and cross-community panelists are being confirmed across validity and fairness, grounded social simulation, reinforcement learning, and simulator-based training.
What we aim to build
We aim to build on open community efforts rather than starting from scratch. The planned outputs are reusable artifacts and draft community guidance on what evidence a simulator should report for a stated use.
A planned open, census-conditioned baseline to build on, extend, or argue with, together with persona and configuration scripts.
A validation checklist and failure-mode taxonomy covering which simulator was used, how it was validated, which populations it represents, and what class of claim the evidence can support.
A proposed index of open benchmarks, datasets, and evaluation environments for user simulation across communities.
A planned non-competitive first-year exercise over selected TREC-style and τ-family scenarios, with a fairness and representativeness lens.
Written breakout notes consolidated into a short public report recording agreements, disagreements, open questions, and research priorities.
Detailed participation mechanics for the calibration exercise will be announced once scenarios and reporting requirements are finalized. The planned catalog will cover work from across the field, including SOTOPIA, UserBench, and ConvApparel, not only benchmarks built by the organizers. A future edition may expand this foundation into a competitive leaderboard.
Call for papers
Simulated users now drive AI evaluation and training, but when can we trust them? User simulators increasingly serve as evaluation subjects, sources of synthetic interaction data and preference signals, and interactive environments for reinforcement learning. Yet overly cooperative, unrepresentative, or poorly calibrated simulators can inflate performance and produce conclusions that do not transfer to real users.
UserSim @ NeurIPS 2026 brings together researchers building, using, validating, and stress-testing user simulators. Our goal is to establish the evidence, standards, and open tools needed to make user simulation diverse, faithful, and valid.
We invite non-archival submissions of 4–8 pages as either extended abstracts or short papers. We welcome empirical studies, benchmark proposals, negative results, replications, tools and demos, position papers, and methodological reports.
UserSim is designed to produce more than a day of presentations. It is an opportunity to help establish the scientific and open-source foundations of trustworthy user simulation.
UserSim is a one-day, in-person workshop in Paris. NeurIPS will provide live streaming and recording for an online audience; details about access and recording availability will be shared when confirmed.
At least one author is expected to present each accepted paper in person. If no author can attend, the paper will remain accepted and listed on the workshop site and OpenReview, but it will not receive a poster slot. Selected papers with an in-person presenter will receive spotlight presentations.
All poster and spotlight presentations must be delivered in person in Paris.
Responsible use
User simulation can reduce reliance on sensitive real-user data, support scalable safety testing, and enable disaggregated fairness analysis without repeatedly exposing real people to risk. But simulators can also inherit bias, reward systems for exploiting their quirks, or create an “easy mode” that masks failures real users would encounter.
That tension is central to the workshop. We are interested not only in stronger simulators, but in validation and reporting practices that make the limits of simulation visible.
Who is behind it
Yev Meyer leads user-simulation and synthetic-data research at NVIDIA and contributes to the Nemotron team’s open models and data, including Nemotron-Personas, a census-grounded persona collection. Previously Chief Scientist at Gretel, he initiated and led Data Designer, an open-source compound-AI framework for synthetic data generation. He holds a PhD in computational neuroscience from Columbia University.
Krisztian Balog is a Professor at the University of Stavanger and staff research scientist at Google DeepMind, with user simulation central to his work on conversational AI, recommenders, and evaluation. He co-authored User Simulation for Evaluating Information Access Systems and leads work on measuring the realism gap in user simulators. He founded usersim.ai and co-organizes the Sim4IA workshop series and TREC User Simulation track.
Flora Salim is a Professor of Computer Science at UNSW Sydney working on robust and trustworthy machine learning, multimodal foundation models, and grounded simulation of human and dynamic systems. Her work includes GenUP, a population-grounded mobility simulator, SOCIA, an agentic simulation-generation platform, and the multi-institutional GenAISim project.
Victor Barres is the first author of τ2-bench and leads the τ-family, which couples user simulators with agent environments to improve evaluation fidelity. At Mercor’s APEX research team, he works on realistic simulations of complex work grounded in labor-market data.
PhD candidate at UC Irvine (advised by Sameer Singh) and member of the Cohere Safety Team; first author of Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. Her research investigates the validity of agentic evaluations and allocational fairness in enterprise settings such as hiring.
Confirmed reviewers
We are continuing to recruit across agentic evaluation, information retrieval, dialogue, HCI, reinforcement learning, and fairness. If you would like to review, please contact the organizers.
Contact
Questions about submitting, joining the calibration exercise, contributing to the shared artifacts, or volunteering to review: organizers@usersim-workshop.org
Before you ask
No. Submissions are open to everyone, wherever you are and whether or not you attend. This workshop is held in Paris. Location matters only for presenting, not for submitting or being accepted.
At least one author is expected to present each accepted paper in person. If no author can attend, the paper remains accepted and listed on this site and OpenReview, but it does not receive a poster slot. Selected papers with an in-person presenter receive spotlight presentations. All poster and spotlight presentations must be delivered in person in Paris.
NeurIPS has announced live streaming and recording for an online audience. Details about access and recording availability will be added when confirmed by NeurIPS.
No. The workshop is non-archival and authors retain all rights. You are free to submit the work to an archival venue afterwards.
Yes. Concurrent submissions are permitted because the workshop is non-archival, provided the other venue’s policies also permit concurrent submission.
Previously published work and work appearing in the NeurIPS 2026 main track are not eligible. Preprints and other non-archival versions are permitted.
No. An existing arXiv or other non-archival preprint is not grounds for rejection. Submissions must still be anonymized for double-blind review.
Organizers may not submit. Sharing a large employer or institution with an organizer does not by itself make someone ineligible. Personal conflicts under NeurIPS policy, such as current advisees or close collaborators, remain ineligible. Organizers and reviewers will recuse themselves from submissions involving their institution or collaborators.