Measure AI beyond a fluent answer
Correctness, evidence and practical usefulness need to be checked separately from how an answer reads.
Start with real questions
Collect frequent questions, incomplete examples and difficult cases from the people who will use the system. A small evaluation set helps preserve expectations across releases.
Make expectations explicit
Check sources for a document question, labels for classification and required fields for a proposed action. Success differs by task. Acknowledging missing information also matters.
Compare changes
When the model, instructions or sources change, rerun the same examples. Look for both improvements and regressions, combining human review with system records.
