Evaluation & Observability

What is LLM-as-judge and when should we use it?

LLM-as-judge uses a strong LLM to evaluate outputs from another (or the same) LLM against a rubric. Cheaper and faster than human eval; suffers from positional bias, sycophancy toward the model being evaluated, and mode collapse. Mitigate with rubric prompts, reference answers, and swapping order in pairwise comparisons. Use for scale evaluation; keep humans in the loop for high-stakes decisions.

More on Evaluation & Observability

Browse the full FAQ for 164 answers, or start a free GenAI maturity assessment to see where your organisation stands.