How to build great evals – part 10 Measure the steps, not just the result. Much like high school math, it isn’t sufficient just to get the right answer, the steps to get there are critical. Two agent trajectories migh…

Madhu Guru 在“如何构建优秀 evals”系列第 10 篇中提出,评估 AI Agent 不能只看最终答案对不对,还要看它走完任务的每一步,因为两条轨迹可能给出同样结果,质量却相差悬殊。








