Est.

Calibrating Judge Confidence Scores Against Human Labels

LLM judges need calibration against human labels to produce trustworthy scores, not assumptions.

Staff Writer & Data Correspondent · · 10 min read
Cover illustration for “Calibrating Judge Confidence Scores Against Human Labels”
LLM-as-Judge Design · October 10, 2026 · 10 min read · 2,232 words

Calibrating judge confidence scores against human labels is the practice of measuring, and then correcting, the gap between what an LLM judge reports and what a human reviewer would actually find. A judge's score is only meaningful once that gap has been checked and closed, not assumed away.

Why LLM judge scores are not trustworthy measures of agent quality without calibration

A team ships an LLM judge to score agent traces, the dashboard turns green, and it stays green for months. When that expert's labels are compared against the judge's scores, the resulting Cohen's kappa comes out to 0.31, well under the 0.6 threshold generally treated as the floor for production use.

An LLM judge score is only meaningful when the judge's behavior lines up with what humans would decide, systematically and verifiably. Absent that check, a judge score is a number that looks like signal. It might be signal. It might also be noise, or worse, a consistent distortion that happens to look stable because it is wrong in the same direction every time.

Three structural mechanisms push a judge toward miscalibration before anyone runs a single calibration test. The first is family and verbosity bias: a judge built on the same model family as the agent it is scoring tends to rate that family's outputs more generously, and longer answers tend to score better regardless of whether they are more accurate. The second is overconfidence baked into the reward model itself. Work out of Meta AI, presented at the ACL 2026 Industry Track, tackles this directly by calibrating LLM judges with linear probes built for fast, reliable uncertainty estimation, precisely because a judge's stated confidence and its actual hit rate often don't match. The third is drift: a model update, a prompt tweak, or a shift in the kind of work flowing through the system can change how a judge behaves without producing any alert, dashboard change, or log entry that would tell anyone to look.

For AI agents, the cost of missing this is higher than it is for scoring a single static model response. None of those failures throw an error code. They pass through the judge clean, and the dashboard stays green right up until someone outside the system goes looking. That gap between the judge's report and what a human would find on inspection is the problem the rest of this piece works through.

What a calibrated judge score means, and the agreement metrics used to measure it

A judge is calibrated when its scores, run over a labeled set of examples, line up with human judgment at a level that clears a documented threshold. The main tool for measuring that correlation is Cohen's kappa, or one of its variants for cases with more than two labelers.

Three properties separate a judge ready for production from one that only looks ready: reproducibility, calibration, and bias control. Bias control means position, verbosity, and family effects have been measured and addressed. Calibration is the middle of the three, and it's the one teams skip most often, usually because it requires building something (a human-labeled gold set) rather than just running a prompt through a few test cases and eyeballing the output.

The thresholds themselves are concrete. A Cohen's kappa above 0.6 between judge and human labels is the generally accepted line for production use. Above 0.8 counts as strong agreement. Below 0.5, the judge should be treated as advisory only, feeding into dashboards and reviews but not triggering automated decisions on its own. Where a gold set has three or more human annotators rather than two, Fleiss' kappa or Krippendorff's alpha takes the place of Cohen's kappa. Where the judge outputs a continuous score rather than a categorical label, weighted kappa or mean absolute error does the same job.

Calibration and accuracy are not the same thing, and conflating them is a common mistake. Research on calibrating autoraters to preference distributions makes this distinction concrete: a judge that's correct a substantial share of the time but expresses much higher confidence on every call is miscalibrated even though its accuracy looks fine on a scoreboard. That distinction matters because confidence is what determines whether a system can safely route a case to automation or needs to send it to a human.

The raw score a judge reports is, in general, a biased estimate of true human agreement. Because judges have imperfect sensitivity and imperfect specificity, the uncorrected fraction of traces a judge marks correct is not the same number a human panel would produce on the same traces. That bias is the specific quantity the rest of this workflow is built to measure and then correct.

Diagram: The Three Kappa Thresholds That Gate Judge Decisions. Visualizes: Visualize the three Cohen's kappa threshold bands that determine how a judge's output may be used in production.

Building the human gold set that anchors the entire calibration workflow

Every later step in this process depends on one artifact: a representative, well-labeled set of traces that stands in for human ground truth. Without it, kappa has nothing to be computed against, bias correction has no reference point, and "calibration" is just a word on a slide.

The scale that works in practice is 200 to 500 hand-labeled traces per workload per rubric, each one labeled by two or three humans working against the rubric. That range isn't arbitrary. The "per workload per rubric" qualifier carries real weight too: a gold set built to check a customer-support agent against a faithfulness rubric doesn't transfer to checking that same agent against an instruction-following rubric. Different rubrics test different things, and a gold set is only as good as its match to the specific question being asked.

Humans need to label against the exact same rubric prompt the judge will see, not against their own sense of what "good" looks like. If the humans and the judge are effectively answering two different questions, the resulting kappa measures the distance between two unrelated judgments, not the thing anyone actually wants to know.

Inter-annotator agreement, measured before the judge even enters the picture, works as a diagnostic on the rubric itself. If kappa between human labelers comes in below 0.4, the rubric is ambiguous, and the fix is to rewrite the rubric before touching the judge. A kappa between 0.4 and 0.6 means the rubric is workable, usually improved by sharpening it with added examples. No amount of judge calibration fixes a badly specified question.

A calibration framework that favors calibrating over curating adds a useful wrinkle to how the gold set gets used once it exists. Rather than hunting for the single most accurate judge and discarding the rest, the stronger approach calibrates a full panel of judges, including the weaker ones, against the labeled set. The practical implication for gold set construction is that it doesn't need to be exhaustive. It needs to be representative of the actual workload and correctly labeled.

A gold set built only from simple single-turn exchanges will miscalibrate a judge that spends most of its time scoring long, branching agentic traces, because the failure modes in a five-step tool-calling sequence (a dropped parameter, a stale retrieval, a plan that silently changes mid-execution) are invisible in a single back-and-forth exchange.

Running the calibration measurement and diagnosing what the numbers reveal

With a gold set in hand, the next step is to run the production judge over it and compute kappa between the judge's scores and the majority human label on each trace, then compare that number to the 0.6 threshold. That four-step loop (run the judge, compute kappa, compare to threshold, decide on next steps) is the core of the measurement phase, and it produces more than a pass or fail.

The kappa value by itself is a gate. The pattern of where the judge disagrees with humans is the diagnostic. If the judge consistently scores longer outputs higher than humans do, that points to verbosity bias, and the fix is a length penalty built into the rubric or some form of token-count normalization. If a model scores its own outputs higher when it's serving as both generator and judge, that's self-enhancement bias, and the only real fix is to stop using the same model for both roles.

Kappa measures agreement in aggregate, while some decisions need a guarantee on the individual case. Conformal prediction applied to individual judge scores, as shown in work by Sheng and colleagues at EMNLP 2025, produces interval guarantees on single evaluations. That distinction matters when one judge call is being used to gate a single high-stakes decision, like approving a refund or releasing a flagged response, rather than to estimate how often the judge is right across a large batch.

A further layer of diagnosis separates two kinds of uncertainty that look identical on a dashboard but call for completely different responses. Some uncertainty comes from the judge being unclear on what the rubric is actually asking. Other uncertainty reflects a case where humans themselves genuinely disagree, because the trace sits in a gray area the rubric was never built to resolve cleanly. Decomposing LLM-judge uncertainty into these two sources matters because the first case calls for clarifying the rubric, while the second calls for routing the case to human review rather than trying to force the judge toward an artificial consensus that doesn't exist among people either.

One more caution belongs in this phase before moving to correction: small gold sets can produce a kappa that looks healthy purely by chance, especially with the low end of the 200-trace range. A finite-calibration regime map for LLM judge panels lays out the specific conditions under which a given human-labeling budget is better spent on a simpler reliability model versus a richer joint output table for aggregating a full judge panel. The practical use of that work is knowing when a kappa reading is trustworthy at the sample size a team actually has, rather than assuming any number above 0.6 settles the question.

Correcting systematic bias and adjusting the score stream to match human ground truth

Once the gold set has revealed that a judge's raw scores are biased, the fix is not to throw out the judge and start over. The fix is to statistically adjust what the judge outputs, using the calibration set as the correction key that translates judge behavior into true human-agreement rates.

The core of this correction is a plug-in framework that estimates two numbers from the human-labeled calibration set: the judge's sensitivity, the probability it marks a truly correct response as correct, and its specificity, the probability it marks a truly incorrect response as incorrect. Call those q₁ and q₀. Once both are estimated, the relationship between the judge's raw accepted fraction and the true accuracy can be inverted using the corrected estimator θ̂ = (p̂ + q̂₀ − 1) / (q̂₀ + q̂₁ − 1). That formula recovers the true human-agreement rate even when the judge's raw score runs too high or too low.

This isn't a theoretical exercise. In empirical testing on public chatbot comparison data using one model as the judge, the raw score carried measurable bias across all six target models evaluated. After applying the correction, bias moved toward zero across all six, with results averaged over many random splits of the data into a majority test set and a smaller calibration set.

The approach also holds under distribution shift, where simpler fixes tend to fall apart. If the mix of correct-to-incorrect responses in the calibration set differs from the mix in the live traffic being scored, the misclassification-adjusted estimator stays unbiased, as long as the judge's behavior conditional on the true label stays stable. That's a meaningful advantage over other calibration-only methods, which lose their accuracy guarantees once the underlying data shifts between the calibration set and the live traffic it's meant to represent.

Where a team controls the judge model directly rather than calling an opaque API, a second correction path becomes available. Other published research on linear probes trains lightweight probes on the judge model's hidden states using proper-scoring-rule supervision, producing calibrated uncertainty estimates without needing multiple sampling passes. This method needs access to the judge's internal hidden states and a labeled dataset of ground-truth verdicts, so it works for a self-hosted or fine-tuned judge, not for a judge accessed only through a third-party API.

A further gain comes from changing what gets passed downstream once scores are corrected. Rather than collapsing a judge's output into a hard pass or fail, propagating the full calibrated win probability into a pairwise ranking procedure brings LLM-derived ratings within 17.9 Elo mean absolute error of ratings built from human judgments, averaged across 55 held-out models on a public leaderboard. Soft, calibrated probabilities carry more information than a binary verdict, and that extra information compounds any time judge output feeds into a ranking system or a reward signal used for further training.

When kappa remains low even after bias correction, the remediation path runs back through the rubric before it runs through the model. Tightening the rubric prompt closes gaps that let the judge interpolate its own criteria where the instructions were vague. Swapping to a different judge checkpoint or model family addresses cases where the bias traces back to the judge itself. And refreshing the gold set addresses the case where the workload has simply changed since the set was built, since a gold set that no longer reflects current production traffic will miscalibrate even a well-built judge running against a well-written rubric.

Sources

  1. Calibrate, Don’t Curate: Label-Efficient Estimation from Noisy LLM Judges
  2. Judging with Confidence: Calibrating Autoraters to Preference Distributions
  3. Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
  4. Decomposing LLM-Judge Uncertainty to Target Expert Labels
  5. Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
  6. A Finite-Calibration Regime Map for LLM Judge Panels
  7. Robust LLM Performance Certification via Constrained Maximum Likelihood Estimation
  8. Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

More in LLM-as-Judge Design