Amazon Science’s Post

Eight out of ten LLM judges agree. But how independently did they arrive at that answer? When judges share a prompt template, training lineage, or model family, a majority vote can make evidence look far more convincing than it really is. The vote count inflates confidence without adding independent signal. Amazon researchers introduce dependence-aware label aggregation using Ising models to address this. The method models both individual judge reliability and pairwise dependencies, allowing redundant agreement to be discounted without discarding votes entirely. Across three tasks with 10-judge panels, the approach improved accuracy by 9–14% over weighted-majority-vote baselines.

  • No alternative text description for this image
  • No alternative text description for this image
  • No alternative text description for this image
  • No alternative text description for this image
  • No alternative text description for this image

When LLM judges agree, should we believe them? https://amzn.to/3UnXvfS

The dependence map is the useful part. Most eval setups treat "add another judge" as a free accuracy win, when a lot of the time you're just buying a second copy of the same failure mode: same pretraining data, same RLHF preferences, same blind spots. Curious how stable the learned structure is across tasks, and what happens when a provider ships a new checkpoint. Also wondering how much labeled data the coupling estimation needs before it actually beats weighted majority. The slide mentions "largest training split," so I'd guess that crossover point matters if you only have a few hundred labels. The ladder framing feels right though. Uniform majority is fine until your judge pool gets homogeneous, which is usually when people stop questioning it.

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories