Recorded with the result
Dataset identity and version, run time, assistant revision, criteria, item results, aggregate values where available, and review state.
Assistant evaluation
Evaluation is workspace-specific. Results apply only to the recorded test set, assistant configuration, connected knowledge, model context, and evaluation date.
Use representative questions, expected evidence, prohibited behavior, and known edge cases. Keep the set versioned and separate from day-to-day prompting.
Execute each item through the same retrieval, model, instructions, and policy path used by the evaluated assistant.
Measure answer correctness, citation support, refusal behavior, latency, and failure state where those criteria are configured. A single aggregate score is not treated as universal quality.
Inspect failed and borderline items, identify retrieval or instruction causes, and require a human decision before promotion.
Rerun the same versioned set after material changes. Record configuration, dataset version, model context, date, and reviewer with the result.
Interpretation
A completed run is evidence about a defined configuration and dataset. It is not a guarantee for every future question, model provider, connector state, or data change.
Dataset identity and version, run time, assistant revision, criteria, item results, aggregate values where available, and review state.
Teams decide acceptable thresholds and inspect failures before publishing or expanding access. High-impact workflows should retain approval and fallback paths.
Model behavior, source freshness, permissions, and provider availability can change. Tests should be rerun after material configuration or knowledge changes.
Ask for the test definition and item-level outcome, not only a headline score. Keep targets, internal measurements, and independently verified results clearly labeled.