All docs
EvaluationCompare versionsCompare two prompt versionsdashboard

Compare two prompt versions

Run both against the same traffic and get a statistical verdict rather than an impression.

What you'll have

An experiment with two variants, a scoring profile, and a success metric — ready for your app to start reporting against.

You'll need
  • A project
  • Two prompt versions worth comparing

A prompt change that looks better on three hand-picked examples is not evidence. Experiments run two versions against the same live traffic and tell you whether the difference is real.

What you get

  • Both variants scored on the same 49 dimensions
  • Statistical significance, not a raw average difference
  • Per-dimension breakdown, so you can see what got better

A prompt that raises overall score while lowering grounding is a trade, not a win. The per-dimension view is where you notice.

Assignment is sticky

A given user stays on the same variant for the life of the experiment, so you are measuring the variant rather than the switching.

When to reach for this instead of Evaluations

Evaluations tell you how your system is doing right now. Experiments tell you whether a change you are considering is an improvement. Use Evaluations to find the problem and Experiments to confirm the fix.