Evaluation — Advanced

What is trajectory quality for agent evaluation?

A measure of whether the intermediate steps an agent takes are sensible — did it call reasonable tools in a reasonable order, did it correctly interpret tool outputs, did it revise its plan when it hit an obstacle. Cannot be captured by task-success-rate alone (an agent can luck into a right answer via wrong steps). Typically graded with LLM-as-judge against a rubric per trajectory.

More on Evaluation — Advanced

Browse the full FAQ for 164 answers, or start a free GenAI maturity assessment to see where your organisation stands.