Evaluation set
A versioned collection of representative, authorized cases used with a stable rubric to test a workload.
Why it matters
It turns model and architecture changes into repeatable comparisons tied to the work.
Measure it with
- Case coverage
- Acceptance rate
- Failure categories
- Reviewer agreement
Do not assume
A small set proves performance for every user, language, edge case, or future model version.