NeurIPS 2026 · Paris · 12 or 13 December

UserSim @ NeurIPS

Grounded User Simulation for Model Evaluation and Training: Diversity, Fidelity, and Validity

Simulated users now drive AI evaluation and training, but when can we trust them?

Making user simulation diverse, faithful, and valid.

Organized by Yev Meyer NVIDIA Krisztian Balog University of Stavanger & Google DeepMind Flora Salim UNSW Sydney Victor Barres Mercor Preethi Seshadri UC Irvine & Cohere

Submissions due
29 August 2026Anywhere on Earth
Notifications
29 September 2026Anywhere on Earth
Workshop
12 or 13 December 2026Paris · exact day to be confirmed

The mission

Accelerate user simulation as an open, rigorous, shared capability

The opportunity is to make user-simulation methods open, validated, and shared across communities. We are bringing together the people who build simulators, the people who depend on them, and the people who test where they fail, helping establish the evidence, standards, and open tools the field can build on.

Why now

Simulation is becoming infrastructure faster than we are learning to trust it

User simulation models how people, and increasingly agents acting on their behalf, behave and respond as they interact with AI systems, products, services, and environments. Simulated users now appear in agentic training pipelines, conversational-search evaluation, dialogue systems, recommender systems, and tool-use benchmarks.

Done well, simulation provides scalable and reproducible interaction data, broader coverage of rare and long-tail behavior, finer control over demographic and behavioral variation, and a path to evaluating multilingual and sovereign-AI systems with little in-language production traffic.

But adoption is outpacing validation. In Lost in Simulation, measured agent success shifted by as much as nine points depending on which model played the user. The study also found that overly cooperative simulators can inflate performance and that simulators can be miscalibrated for speakers of dialects including African American Vernacular English and Indian English. Simulator choice is therefore part of the measurement, not a neutral implementation detail.

The challenge is broader than selecting a better model. UserBench models users who reveal preferences incrementally, while τ2-bench represents settings in which users share control of the environment. These designs highlight what realistic simulation must account for: goals that evolve over time, interaction histories, other actors, and the social, digital, and physical environments that constrain behavior. If the simulator is a measurement instrument, both the user and the conditions shaping their behavior need validation.

At the same time, grounded population-scale simulation is newly feasible. The field has an opportunity to establish shared evidence and reporting norms now, before weak assumptions harden into standard practice.

A real-user trajectory and simulated-user trajectory diverge over six turns.
The simulation-to-reality (Sim2Real) gap: a real user and a simulated one, given the same opening turn, drifting apart as the conversation goes on.

Who it’s for

Across communities, one shared challenge

Relevant work is scattered across labs and communities that rarely meet, often using different terminology, assumptions, and evaluation practices. Without shared, validated simulators, datasets, and reporting norms, results remain difficult to compare and the field cannot build cumulatively. We are convening these communities around a shared challenge: representing users, populations, and the contexts that shape their behavior.

  • Agentic LLM evaluation
  • Information retrieval and conversational search
  • Dialogue and task-oriented systems
  • HCI and social simulation
  • Reinforcement learning and recommender systems
  • Synthetic data and model training
  • Multilingual and lower-resource NLP
  • Fairness, responsible AI and evaluation methodology

If you build simulators, validate or audit them, or use them to train and evaluate models, this workshop is for you. Work that stress-tests simulation is as welcome as work that advances it. Negative results, replications, and evidence that a simulator fails to transfer are first-class contributions.

  • The first NeurIPS workshop centered squarely on user simulation. Prior workshops have explored adjacent topics such as persona modeling, agent environments, and multi-turn interaction. UserSim makes the simulated user the unifying object across evaluation and training.
  • Building across communities. UserSim builds on Sim4IA and TREC while connecting them with agentic LLMs, dialogue, HCI and social simulation, reinforcement learning, and recommender systems.
  • Evaluation and training together. It connects simulation as an evaluation substrate with simulation as training data, preference and process-reward signal, and RL environment, while pairing simulator builders with validity-and-fairness skeptics.

Three pillars

What must a trustworthy simulator demonstrate?

We organize the agenda around three pillars, covering the two strategically important uses of simulation: evaluation of interactive and agentic systems, and training via SFT data, preference and process-reward signals, and simulators as RL environments.

Diversity

The range of users, languages, cultures, goals, tasks, and scenarios represented, so simulation spans the target population rather than a narrow slice.

Fidelity

Whether simulated users match real individual and population-level behavior, including style, strategy, natural variation, and the tendency of real users to be incomplete, inconsistent, impatient, or strategic rather than idealized and cooperative.

Validity

Whether conclusions drawn from simulation transfer to reality: do simulated evaluations predict real outcomes, and is the training signal trustworthy?

A target population with a highlighted subset showing what the simulator covers.
Diversity, stated literally: how much of the target population does a simulator actually stand for?

The centerpiece

What Must User Simulation Capture?

People do not act as isolated personas, and no single user stands for a population. A simulation must decide what it represents: individual goals and histories, behavioral variation and population distributions, relationships and other actors, and the social, digital, and physical environments that shape interaction.

This cross-community panel will examine which components matter for different uses, how they should be grounded and evaluated, and how requirements change across evaluation, synthetic data, and training.

On the day

Schedule

A single full day, provisionally 08:00–17:00. The program combines a keynote and invited talks, contributed spotlights, two poster sessions, the cross-community panel, and a closing breakout to advance draft reporting guidance and the benchmark catalog.

The detailed schedule will be published once all speakers are confirmed. Every session is designed to move from evidence to action: what user simulation needs to capture, where methods fail, and what the community should investigate and report next.

Program

Confirmed invited speakers

  • Craig Boutilier Google Research Decision theory, user modeling, recommender systems, and user simulation
  • Hyunwoo Kim KAIST Kim Jaechul Graduate School of AI NVIDIA-KAIST Joint AI Research Lab; synthetic data, privacy, social reasoning, and agentic AI

Additional invited speakers and cross-community panelists are being confirmed across validity and fairness, grounded social simulation, reinforcement learning, and simulator-based training.

What we aim to build

Shared infrastructure, not just a program

We aim to build on open community efforts rather than starting from scratch. The planned outputs are reusable artifacts and draft community guidance on what evidence a simulator should report for a stated use.

  • A pilot baseline user simulator

    A planned open, census-conditioned baseline to build on, extend, or argue with, together with persona and configuration scripts.

  • Draft reporting guidance

    A validation checklist and failure-mode taxonomy covering which simulator was used, how it was validated, which populations it represents, and what class of claim the evidence can support.

  • A living benchmark catalog

    A proposed index of open benchmarks, datasets, and evaluation environments for user simulation across communities.

  • A community calibration exercise

    A planned non-competitive first-year exercise over selected TREC-style and τ-family scenarios, with a fairness and representativeness lens.

  • A community report

    Written breakout notes consolidated into a short public report recording agreements, disagreements, open questions, and research priorities.

Detailed participation mechanics for the calibration exercise will be announced once scenarios and reporting requirements are finalized. The planned catalog will cover work from across the field, including SOTOPIA, UserBench, and ConvApparel, not only benchmarks built by the organizers. A future edition may expand this foundation into a competitive leaderboard.

Call for papers

Bring your evidence, methods, and hard questions

Simulated users now drive AI evaluation and training, but when can we trust them? User simulators increasingly serve as evaluation subjects, sources of synthetic interaction data and preference signals, and interactive environments for reinforcement learning. Yet overly cooperative, unrepresentative, or poorly calibrated simulators can inflate performance and produce conclusions that do not transfer to real users.

UserSim @ NeurIPS 2026 brings together researchers building, using, validating, and stress-testing user simulators. Our goal is to establish the evidence, standards, and open tools needed to make user simulation diverse, faithful, and valid.

We invite non-archival submissions of 4–8 pages as either extended abstracts or short papers. We welcome empirical studies, benchmark proposals, negative results, replications, tools and demos, position papers, and methodological reports.

Why submit?

UserSim is designed to produce more than a day of presentations. It is an opportunity to help establish the scientific and open-source foundations of trustworthy user simulation.

  • Set the scientific standardHelp define the evidence needed before results from simulated users can be trusted.
  • Build an open foundationContribute to shared simulators, scenarios, validation practices, reporting guidance, and benchmark resources that the community can build on.
  • Reach across communitiesPut your work in conversation with researchers spanning agentic evaluation, information retrieval, dialogue, HCI, reinforcement learning, recommender systems, synthetic data, and responsible AI.
  • Develop and showcase early workReceive focused feedback, present a poster, and be considered for a spotlight while retaining the option to publish an archival version later.

Topics

Diversity

  • Population grounding and representativeness
  • Fairness and demographic, dialectal, and multilingual calibration of simulators
  • Scenario and goal generation, including coverage of rare and long-tail interactions

Fidelity

  • Behavioral realism and cognitive plausibility of simulated users; the Sim2Real gap and validation against real humans
  • Individual- and population-level fidelity, including behavioral variation, style, strategy, and consistency over time
  • Controllability and reliability: steering persona and instruction-following, consistency across runs, and simulator failures that invalidate a measurement or a reward signal

Validity

  • Simulator diversity
  • Train/eval separation and contamination
  • Human calibration
  • LLM-judge auditing

Settings and modalities

  • Agentic and tool-use user simulation; interactive and trajectory-level evaluation
  • Voice and multimodal user simulation: turn-taking, interruption, and full-duplex interaction
  • Conversational search, dialogue systems, recommender systems, HCI, and social simulation

Training and infrastructure

  • Simulation for training: SFT, preference optimization, and process-reward data generation; the user simulator as an interactive RL environment
  • Small, open, and distilled user simulators, including specialized models rather than prompted ones, that the community can run locally
  • Synthetic-data generation and compound-AI pipelines for building, grounding, and scaling simulators (frameworks, persona datasets)

We are especially interested in work that advances

  1. Open, multi-domain benchmarks that move beyond narrow settings and cover diverse use cases
  2. Validation methods that measure how well a simulator matches real human behavior, including proposals for human-likeness scores and calibration against held-out human data
  3. Fairness and calibration benchmarks across dialects, languages, and populations
  4. Guidance on contamination and reward hacking when simulators are used as a training signal
  5. Reporting norms for simulation-based evaluation, including disclosure standards for which simulator was used, how it was validated, which populations it represents, and its computational and API cost
  6. A shared problem taxonomy and terminology that lets methods and results transfer across communities
  7. Small, open, and distilled simulators that make evaluation reproducible without frontier-API dependence, including evidence on whether specialization can substitute for scale
  8. Cross-disciplinary work that connects methods, evidence, or validation practices across machine learning, HCI, information retrieval, dialogue, recommender systems, behavioral and social science, fairness, and related fields

Submission details

Submission type
Non-archival extended abstract or short paper
Paper length
4–8 pages, excluding Limitations, Appendix, and References
Format
PDF using the NeurIPS 2026 LaTeX template
Submission portal
OpenReview
Review
Submissions must be anonymized for double-blind review
Archival status
Non-archival; authors retain all rights
Prior publication
Previously published work and work appearing in the NeurIPS 2026 main track are not eligible
Preprints
Permitted. An existing arXiv or other non-archival preprint is not grounds for rejection. Submissions must still be anonymized for review.
Concurrent submissions
Permitted because the workshop is non-archival, provided the other venue’s policies also permit concurrent submission

Attendance and presentation

UserSim is a one-day, in-person workshop in Paris. NeurIPS will provide live streaming and recording for an online audience; details about access and recording availability will be shared when confirmed.

At least one author is expected to present each accepted paper in person. If no author can attend, the paper will remain accepted and listed on the workshop site and OpenReview, but it will not receive a poster slot. Selected papers with an in-person presenter will receive spotlight presentations.

All poster and spotlight presentations must be delivered in person in Paris.

Important dates

  • Submission deadline: 29 August 2026, Anywhere on Earth
  • Decision notification: 29 September 2026, Anywhere on Earth
  • Workshop date: 12 or 13 December 2026, exact day to be confirmed

Responsible use

Simulation can reduce harm or hide it

User simulation can reduce reliance on sensitive real-user data, support scalable safety testing, and enable disaggregated fairness analysis without repeatedly exposing real people to risk. But simulators can also inherit bias, reward systems for exploiting their quirks, or create an “easy mode” that masks failures real users would encounter.

That tension is central to the workshop. We are interested not only in stronger simulators, but in validation and reporting practices that make the limits of simulation visible.

Who is behind it

Organizers

  • Yev Meyer NVIDIA

    Yev Meyer leads user-simulation and synthetic-data research at NVIDIA and contributes to the Nemotron team’s open models and data, including Nemotron-Personas, a census-grounded persona collection. Previously Chief Scientist at Gretel, he initiated and led Data Designer, an open-source compound-AI framework for synthetic data generation. He holds a PhD in computational neuroscience from Columbia University.

  • Krisztian Balog University of Stavanger · Google DeepMind

    Krisztian Balog is a Professor at the University of Stavanger and staff research scientist at Google DeepMind, with user simulation central to his work on conversational AI, recommenders, and evaluation. He co-authored User Simulation for Evaluating Information Access Systems and leads work on measuring the realism gap in user simulators. He founded usersim.ai and co-organizes the Sim4IA workshop series and TREC User Simulation track.

  • Flora Salim UNSW Sydney

    Flora Salim is a Professor of Computer Science at UNSW Sydney working on robust and trustworthy machine learning, multimodal foundation models, and grounded simulation of human and dynamic systems. Her work includes GenUP, a population-grounded mobility simulator, SOCIA, an agentic simulation-generation platform, and the multi-institutional GenAISim project.

  • Victor Barres Mercor

    Victor Barres is the first author of τ2-bench and leads the τ-family, which couples user simulators with agent environments to improve evaluation fidelity. At Mercor’s APEX research team, he works on realistic simulations of complex work grounded in labor-market data.

  • Preethi Seshadri UC Irvine · Cohere

    PhD candidate at UC Irvine (advised by Sameer Singh) and member of the Cohere Safety Team; first author of Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. Her research investigates the validity of agentic evaluations and allocational fairness in enterprise settings such as hiring.

Confirmed reviewers

Program committee

  • Dane Corneil
  • Ritu Gala
  • Yuncheng “Devin” Hua
  • Jiantao Jiao
  • Aditya Joshi
  • Dhruv Nathawani
  • Venkat Srinivasan
  • Eric Tramel
  • Charles Wang
  • Hao Xue

We are continuing to recruit across agentic evaluation, information retrieval, dialogue, HCI, reinforcement learning, and fairness. If you would like to review, please contact the organizers.

Contact

Get in touch

Questions about submitting, joining the calibration exercise, contributing to the shared artifacts, or volunteering to review: organizers@usersim-workshop.org

Before you ask

Questions

Do I need to be registered for NeurIPS Paris to submit?

No. Submissions are open to everyone, wherever you are and whether or not you attend. This workshop is held in Paris. Location matters only for presenting, not for submitting or being accepted.

What if no author can attend in Paris?

At least one author is expected to present each accepted paper in person. If no author can attend, the paper remains accepted and listed on this site and OpenReview, but it does not receive a poster slot. Selected papers with an in-person presenter receive spotlight presentations. All poster and spotlight presentations must be delivered in person in Paris.

Can I follow the workshop online?

NeurIPS has announced live streaming and recording for an online audience. Details about access and recording availability will be added when confirmed by NeurIPS.

Is the workshop archival?

No. The workshop is non-archival and authors retain all rights. You are free to submit the work to an archival venue afterwards.

Can I submit work that is currently under review elsewhere?

Yes. Concurrent submissions are permitted because the workshop is non-archival, provided the other venue’s policies also permit concurrent submission.

Is previously published work eligible? What about main-track NeurIPS 2026 papers?

Previously published work and work appearing in the NeurIPS 2026 main track are not eligible. Preprints and other non-archival versions are permitted.

I already have a preprint on arXiv. Does that disqualify me?

No. An existing arXiv or other non-archival preprint is not grounds for rejection. Submissions must still be anonymized for double-blind review.

Can the organizers submit? How are conflicts of interest handled?

Organizers may not submit. Sharing a large employer or institution with an organizer does not by itself make someone ineligible. Personal conflicts under NeurIPS policy, such as current advisees or close collaborators, remain ineligible. Organizers and reviewers will recuse themselves from submissions involving their institution or collaborators.