Evaluation — Advanced
How do we detect model drift in production?
Run a live evaluation harness on a sampled slice of production traffic (or on a synthetic re-run of yesterday's prompts against today's model). Alert when the eval score drops below a threshold or delta. Drift sources include: provider model updates, prompt template edits, upstream data changes, and evaluation-metric decay.
More on Evaluation — Advanced
Related on this site
Framework dimensions
Free tools
Whitepapers
Browse the full FAQ for 164 answers, or start a free GenAI maturity assessment to see where your organisation stands.