Principle

Simulate users who behave like users

A simulator that helps the agent too much is grading a different job.

When it fits

  • A multi-turn agent looks excellent in synthetic conversations but real users still produce surprising failures.

When to avoid it

  • Synthetic users remain approximations. Do not replace production monitoring or real-user research with simulation alone.

Why it matters

Model the target user's information, patience, ambiguity, mistakes and willingness to cooperate. Do not let the simulator volunteer missing details merely because an assistant model knows they would help. Compare simulated failures with production traces and refine the user model when the two diverge.

An example

A support simulator should sometimes say 'it still doesn't work' instead of naming the exact error code the agent needs to solve the case.

Check your result

The simulator can reproduce at least some real user friction without turning into a hidden co-pilot for the agent.

Keep this limit in mind

  • Synthetic users remain approximations. Do not replace production monitoring or real-user research with simulation alone.

Evidence and sources

Supports

Multi-turn agent evaluation can become unrealistically easy when the simulated user cooperates like a helpful assistant instead of behaving like the target user population.

No simulator perfectly reproduces real users; production evidence remains necessary to check whether simulated behavior covers the failure modes that matter.

Build Evals That Actually Matter · 7:20-17:37, simulate the whole interaction and avoid an unrealistically helpful user simulator

All sources (1)