Between Intention and Experience
The Promise and Limits of Simulated Users
31.03.2026
Simulated users in research: promising, but only with guardrails
There’s a lot of hype around AI‑driven “users”. Most of it skips one crucial fact: Simulated users (LLM‑based user agents) are models of users, not “virtual people”.

They generate plausible narratives, not ground truth.
Simulated users in research may be promising, but only with guardrails in place.

User‑simulation research is very detailed on this:
“Simulators do not need to be perfect mirrors of human behaviour, but instead simply need to be good enough… output from simulations should correlate well with human assessments on a given task… The main requirement is reproducibility.”

Simulated users can support research, but they cannot replace listening to real people.

In practice, studies like UXAgent and recent work on customer “digital twins” show that simulated users can pilot studies and stress‑test designs… but fail spectacularly on real decisions:
“Only ~2% of human-agent pairs chose the same product, despite similar task completion rates”

How Simulated Users Can Support Research

Used well, they are great at making our human research more effective:

  • Before the fieldwork: pilot interview guides and usability tasks; expose leading questions and broken flows.
  • In discovery: explore “what if” scenarios and edge cases; generate hypotheses worth taking to real people.
  • In analysis: cluster quotes, draft themes or personas from real data so that researchers can focus on interpretation.

Guardrails for Responsible Use

Used badly, user simulation becomes a high‑confidence “innovation theatre” – something Ricardo Martins dissects sharply when he shows how tool marketing can over‑promise and under‑deliver [written in 🇵🇹 ].

That’s why I recommend a simple internal policy:

Allowed: study piloting, early‑stage exploration, journey stress‑testing, analytical support on existing data.

Not allowed as sole evidence: launch‑critical decisions, user behaviour, pricing, policy and inclusion, accessibility, work with marginalised or statistically less represented.

📌 Always document: purpose (“exploratory” vs “confirmatory”); model and prompts used (ideally compare different models; be wary of polished platforms whose claims you cannot independently validate); key assumptions; and whether results have been checked against real users (including simple control groups where possible).

Important reminder: clearly label all outputs as simulated vs validated with humans.

The point is not to be anti‑AI. It’s to be honest about what user simulation CAN and CANNOT do (don’t expect suppliers to be very transparent about the second) – and to make sure our enthusiasm for new tools does not quietly replace the hard, slow work of actually listening to people.

AI disclosure: The accompanying illustration was generated using Gemini AI. AI was also used to support the outlining and editing of this post. The final content and revisions are my own.
Also published on LinkedIn
Daniel T Santos
See also
    Made on
    Tilda