
Risk Analysis11 min read
Your LLM judge agrees with humans 71% of the time. That is not enough.
We graded 2,400 model outputs with three human raters and four LLM judges. The best judge hit 71% agreement — and disagreed with humans most on exactly the failures that mattered.
Read the article

