Evaluation Suite
Kind: technique
Layer: Correctness Core
Record: lexicon:evaluation-suite
Canonical: Lexicon
A technique for scoring model output against a fixed set of cases on every change to the model, prompt or data.
Listed in Lexicon terms, after Prompt Versioning and before Centralized Model Configuration.
Category
Refactors
- Prompt Sprawl
- Artificial Intelligence Architecture
- Model Governance
- Model Evaluation
- Vector Search
- Model Safety