LLM as a Judge: Sampling, Calibration, and Evaluation Drift

Ask how an LLM feature will be evaluated in production, in a design review or a system design interview, and one answer is to have another model score the output. The judge is a component of your system with its own cost, latency, failure modes, and version history. Proposing an LLM judge without saying how you validate it adds an unmeasured dependency to the layer that is supposed to do the measuring.
An LLM judge is a measuring instrument, valid only while it is versioned and calibrated. It works as one of three layers used together: deterministic checks for anything objectively verifiable, a judge for the subjective residue, and periodic human labels to check the judge. Applying LLM as a judge best practices starts with stating where those layers meet: which criteria stay deterministic, which eligible traces reach the judge, and what decision the judged sample supports.
Why a judge is the right starting instinct
Human review covers only what reviewers have time to label. Rule-based checks can confirm that output is well-formed, not that an answer is good. User feedback mixes reactions to latency, tone, and whether the problem was solved, and users may have no way to check faithfulness to the documentation. When reviewing every production trace would exceed reviewer capacity, a model scoring sampled answers against a rubric is a practical approach to evaluating AI output at scale.
A running example: a retrieval-backed support assistant
Consider a support assistant that retrieves help-center articles and returns JSON with an answer and cited document IDs. Evaluation runs as an asynchronous worker over logged traces, outside the request path, so replies do not wait for judge results.
Some criteria are objective: the output parses, the citations refer to documents that were actually retrieved. Others are subjective: whether each claim is supported by the cited article (faithfulness), and whether the answer resolves the user's question (helpfulness).
Deterministic checks run on all traffic
Everything objectively verifiable belongs in code that runs on every trace:
- The response parses as JSON and matches the schema.
- The citation list is non-empty, and every cited document ID exists in the set retrieved for that trace. A model can emit a plausible ID that was never retrieved.
- Forbidden content is absent: internal hostnames, ticket-system links, strings matching account-number patterns.
- Quoted spans and specific values, such as a refund window in days, appear in the cited document after normalization.
These checks are cheap next to a judge call and repeatable. Repeatability alone does not tell you why the failure rate changed: a generator update, a retrieval or index change, an edited check rule, or a traffic shift can each move it. Logging generator, index, and check versions and the segment on each trace helps separate these causes.
A citation check confirms that the cited ID was retrieved, not that the document supports the sentence next to it. Exact matching works for verbatim quotes and extractable values and fails on paraphrase, which is much of what a helpful answer contains. Deciding whether a paraphrase is supported is semantic work, the residue the judge exists for.
Judge the subjective residue on a designed sample
For judge inference alone:
judge spend per day = judged items per day
× criteria per item
× tokens per judgment
× price per token
That boundary excludes human labeling, trace storage, worker compute, and retries. In this assistant, retrieved context can make up most of the tokens per judgment, because a faithfulness judge reads the source articles along with the question, answer, and rubric. A single judgment can cost as much as the generation it evaluates, or more.
The sample depends on the decision the scores support. For tracking the overall faithfulness rate, a uniform random sample works, because its precision depends on how many items are judged rather than on total traffic. Rare failures in a specific segment need different arithmetic. If a failure occurs in a fraction p of a segment's traffic, and the judge flags it with recall r, then n randomly sampled items from that segment surface at least one flagged instance with probability:
P(detect) = 1 − (1 − p·r)^n
This holds only if sampling is random within the segment and r is measured against human labels rather than assumed. When p·r is small, detection depends roughly on n·p·r, so halving prevalence or recall requires roughly doubling the sample for the same detection probability. A uniform sample under-represents small segments, such as a new locale or a newly indexed document set. Stratify by segment, oversample where a miss is expensive, and weight strata by traffic share when reporting the overall rate.
The configured rate is not the achieved rate. Keep four numbers separate: the configured sample rate, the eligible traffic it applies to (for example, traces that passed deterministic checks), the judge provider's rate limit, and the throughput your worker pool sustains in testing. Achieved judged volume is the configured rate applied to eligible traffic, capped by whichever of the last two binds first. If the rate limit binds at peak, the configured rate is only a ceiling and the sample may skew toward off-peak traffic. Monitor judged-item counts per segment.
When judging every response is warranted
If the judge makes a per-response decision, such as blocking an answer in a refund or account-security workflow before it is sent, or routing a response to a human reviewer, it has to see every response in that workflow. Sampling cannot gate a response it never evaluated.
In the request path, the judge's latency adds to the user's wait. Its outages need an explicit policy: fail open and send unchecked answers, or fail closed and block them. Its false-positive rate becomes user-visible friction. A per-response gate on a high-risk workflow and a sampled monitor over all traffic can run side by side; both need the calibration below.
Known judge biases, and how to measure yours
Zheng et al. (2023), the paper that introduced MT-Bench, examined position, verbosity, and self-enhancement bias. The authors found some of these minor or mitigable, reported that GPT-4 as a judge agreed with human preferences about as often as humans agreed with each other, and saw signs of self-preference their data could not firmly establish. Wang et al. (2023) showed how strong position bias can be in pairwise comparison: with ChatGPT as the evaluator, reordering the responses let Vicuna-13B beat ChatGPT on 66 of 80 queries.
Those results come from open-ended chat benchmarks, largely pairwise comparisons, with 2023-era models. They show the biases exist; they do not tell you their size for your judge, rubric, and traffic. Each maps to a measurement:
- Position. For pairwise comparisons, evaluate both orderings and either average the scores, as Wang et al. propose, or treat inconsistent verdicts as ties, as Zheng et al. do. The example's pointwise faithfulness judge is less exposed.
- Verbosity. On the human-labeled set, check whether judge scores correlate with answer length more strongly than human scores do. Zheng et al. found verbosity sensitivity varied considerably across judge models.
- Self-preference. Where practical, use a judge from a different model family than the generator, or measure whether it scores same-family outputs higher than humans do.
- Inconsistency. Rerun the judge on the same items and measure how often it agrees with itself. Items whose verdicts flip are low-confidence and make good candidates for human labeling.
Humans judge the judge, on a sample and a cadence
For the support assistant, that means a small set of recent production traces, refreshed on a schedule and labeled against the same rubric by the people whose judgment matters: support leads for helpfulness, documentation owners for faithfulness. From those labels, compute the judge's precision and recall on its "unsupported claim" flag, per criterion and per segment, plus a chance-corrected agreement statistic such as Cohen's kappa. This judge model calibration supplies r in the detection formula, estimated per failure criterion and segment from the labeled items.
Have two people label an overlapping subset. Where they disagree on a criterion, adjudicate those items, clarify the rubric, and revise the labels before scoring the judge against them.
Calibration has to be periodic even when the judge is pinned, because traffic keeps moving. New product areas, document sets, and question types can erode agreement without any change to the judge. Tie the cadence to how quickly those inputs change, and relabel after events such as a generator model change or a large documentation update.
Changing the judge is a measurement migration
The judge is a model version, a prompt, a rubric, decoding settings, and the code that parses its output. Changing any of them, or a provider updating the model behind an alias, can change the instrument. A score shift caused that way is LLM as judge evaluation drift, and comparability with historical scores has to be checked across each change.
Treat judge changes like a schema migration:
- Pin an exact, dated model version instead of an alias the provider can repoint, and plan for provider deprecations.
- Store the judge version identifier with every score.
- Before cutting over, run old and new judges in parallel for an overlap period on a frozen human-labeled set and a slice of recent traffic. Because both score the same items, their differences reflect the instrument. Compare each judge's precision and recall against the human labels by criterion and segment, and compare paired verdicts on recent traffic to see where scores shift.
- Use the overlap results to decide whether thresholds and alerts still hold, need adjusting, or need a new baseline, and mark the cutover on dashboards.
A faithfulness rate that falls the week the judge was upgraded may reflect a stricter instrument rather than a worse assistant. Before rolling back a generator or retrieval change, check the judge's version history and whether recent human labels show the same drop.
Disclosure: Formation publishes this blog. Formation's mock interviews let you practice defending a design like this with experienced engineers asking the follow-up questions.