Evaluation & Observability
What is LLM-as-judge and when should we use it?
LLM-as-judge uses a strong LLM to evaluate outputs from another (or the same) LLM against a rubric. Cheaper and faster than human eval; suffers from positional bias, sycophancy toward the model being evaluated, and mode collapse. Mitigate with rubric prompts, reference answers, and swapping order in pairwise comparisons. Use for scale evaluation; keep humans in the loop for high-stakes decisions.
More on Evaluation & Observability
Related on this site
Framework dimensions
Free tools
Whitepapers
Browse the full FAQ for 164 answers, or start a free GenAI maturity assessment to see where your organisation stands.