Skip to content

BlogAI

LLM Monitoring: Why Fast, Valid Answers Can Still Be Wrong

LLM Monitoring: Why Fast, Valid Answers Can Still Be Wrong

A customer asks a refund assistant whether they can return an opened blender bought 40 days ago. The service answers in 400 ms with HTTP 200 and JSON that validates against the response schema: {"eligible": true, "reason": "Opened items can be returned within 60 days", "source": "refund-policy#4"}. The latency monitor passes, the error-rate monitor passes, and the parse-failure counter stays at zero. The answer is wrong. In this hypothetical, the policy moved to a 30-day window for opened items last quarter, and the chunk the model cited came from the superseded version, which was never removed from the index.

The defining failure of an LLM system is a response that is fast, successful, well-formed, and wrong. Monitoring that measures only delivery reports it as healthy. Observing these systems takes two things beyond infrastructure metrics: traces that record what each stage did, and evaluations that judge what the output says. Groundedness, the natural first quality metric for a retrieval system, depends on both, because it compares the answer's claims against the context the trace captured.

Why infrastructure signals miss it

Latency, error rate, throughput, and saturation answer whether the system delivered a response. They miss semantic errors in conventional software too, but a wrong database row usually traces to a defect you can reproduce, bisect, and fix.

An LLM that produces a plausible wrong answer is doing what the component does at some nonzero rate. No exception is thrown and no timeout fires. The output parses because structured output constraints guarantee shape, and shape says nothing about truth. The same request can succeed today and fail next week if the context, prompt, or model changes.

None of this is entirely new. Search and recommendation teams have long paired offline relevance evaluation with online quality metrics, because a ranked list can also be fast and wrong. What changes with generated text is that the output makes free-form claims, and checking them means reading them against whatever was in the context window for that request.

Traces and evaluations answer different questions

A trace tells you the request took 1.4 seconds, made three tool calls, and returned 800 tokens. An evaluation tells you the answer was ungrounded. Neither answers the other's question. Traces alone explain why a request was slow or expensive but not whether it was correct. Evaluations alone can report that quality dropped on Tuesday with no record of what changed.

The two meet at groundedness. To score whether an answer is supported by its context, the evaluator needs the exact context the model saw, and only the trace has it. That dependency determines what the trace has to store.

What the trace has to capture

For a retrieval-augmented system, a useful trace records, per request:

  • Request input and generated answer, including the resolved prompt variables and conversation context sent to the model, retained under applicable privacy and retention controls or recoverable through permitted durable references to versioned content.
  • Prompt template version and model identifier. Record both the model you requested and the model the provider reports serving the response, since an alias can resolve to a different snapshot over time.
  • Sampling parameters such as temperature and max tokens, because a config change can shift output behavior with no code change.
  • Retrieved chunk IDs, their text or content hashes, source document versions, and the index version. Include reranker scores and the final set that entered the context, in addition to the candidates.
  • Tool calls with their arguments and results.
  • Per-span token counts and latency, so cost and time are attributed to the stage that incurred them.

The retrieval fields are the easiest to omit, and they are what separate a retrieval failure from a generation failure after the fact. Chunk IDs alone are insufficient, because indexes get rebuilt and documents get edited after the request. Re-running the query next week returns whatever the index holds next week. In the refund case, the trace needs to show that refund-policy#4 came from the pre-change document version and that the index build predated the policy update.

These fields usually span several services: a retriever, a reranker, and possibly a model gateway. Trace context has to propagate across each hop, or the retrieval span ends up in another service's logs with no shared trace ID.

OpenTelemetry's GenAI semantic conventions define standard attributes for part of this, including gen_ai.request.model, gen_ai.response.model, and gen_ai.usage.input_tokens. They are still in Development status, so names can change. The attribute registry also says message content capture should require explicit opt-in because it may contain personal data. If policy forbids retaining full context, store references to immutable document versions the policy allows keeping; a hash alone cannot reconstruct text, and discarded input cannot be replayed.

Offline evaluation gates deploys

Offline evaluation is the test suite. It runs a fixed, versioned dataset of inputs through the system and checks expected properties: the answer states the 30-day window, cites a current policy chunk, or declines when the question falls outside policy. The suite runs in CI and blocks the deploy when a metric falls below its threshold.

Thresholds should be set per metric and per slice. An aggregate score can rise while refund-policy questions regress, if a prompt change helps a larger category. Versioning the dataset keeps scores comparable across runs.

A prompt change is a deploy. So is a change to sampling parameters, retrieval configuration, or the corpus behind the index. The refund failure came from index contents, which no code review would have seen. A model version change is a deploy you did not initiate. Pinning a dated model snapshot, where the provider offers one, turns that change into something you schedule and gate. It does not cover everything: aliases move, snapshots are retired on the provider's timeline, and some serving-side changes may not surface in any identifier you receive. Recording the response-reported model catches the visible changes, and online scoring is the backstop for the rest.

Online scoring monitors live traffic

Online evaluation is the monitor. It scores a sample of production traces asynchronously, outside the request path, so judge latency and judge outages never reach users. Three numbers define it: the sample rate, the scored count per stratum per time window, and the judge cost, which is scored count times the input- and output-token cost of every judge call per answer.

Precision comes from the scored count. To estimate a failure rate near 2% within ±1 percentage point at 95% confidence, a nominal binomial approximation calls for about 750 scored answers in each stratum and window (1.96² × 0.02 × 0.98 / 0.01² ≈ 753), whether the service handles ten thousand requests a day or ten million. That assumes independent, randomly sampled, correctly labeled outcomes; judge error and calibration uncertainty widen the interval around the true failure rate. If refund questions are a small share of traffic, a uniform sample may score too few of them to detect a regression, so set a per-stratum target. Judge spend then scales with the number of strata and windows rather than with traffic. Assuming one 3,000-token judge call per answer (input plus output), 750 scores come to about 2.25 million judge tokens per stratum per window.

Write scores back onto the trace, as span attributes or evaluation records keyed by trace ID, so quality is queryable next to latency and version fields: groundedness by index version, failure rate by prompt template, scores before and after a model change. That join is what makes a bad week diagnosable.

Four quantities need to stay separate: the alert threshold you configure, the observed score the judge reports, judge-human agreement measured by having people label a sample of the same traces, and the true failure rate, estimated after correcting for judge error. The MT-Bench study found GPT-4 judges reached over 80% agreement with human preferences, similar to agreement between humans, and also documented position, verbosity, and self-enhancement biases. That setting was open-ended chat preference, so the agreement rate for your judge on your task still has to be measured.

Online scoring complements in-path guardrails and offline gates. A guardrail, such as a rule that rejects answers citing no chunk, runs on every request and adds latency. The offline gate catches known regressions before deploy. Online scoring catches drift and unanticipated failures after the fact. The three are combined rather than chosen between, and the design question is which failures each layer owns.

Groundedness and what it does not catch

For a retrieval system, groundedness asks whether each claim in the answer is supported by the retrieved context. Ragas defines its faithfulness metric as the fraction of response claims the context supports. It suits online scoring because it needs no reference answer, only the output and the context the trace captured.

Whether it is cheap depends on the implementation. Claim extraction plus per-claim verification with an LLM judge can take several model calls per scored answer; a smaller NLI-style classifier costs less and may agree with human reviewers less often. Price it with the formula above, counting every call's tokens.

The refund answer shows its main limit: every claim is supported by the retrieved chunk, so it scores as fully grounded and is still wrong, because the context itself was stale. Groundedness measures generation against context. It says nothing about whether the context was correct.

That limit is also what makes the pairing with traces diagnostic. Low groundedness points at generation: the model added or distorted claims. High groundedness on a wrong answer points at retrieval or the corpus, and only a trace with document versions can confirm which. For the refund failure, the cheapest check needs no model: compare the document versions recorded in the trace against the current policy version and flag answers built on superseded sources.

Whether groundedness is the metric that matters most depends on the system. For an assistant answering from an authoritative policy corpus, it is central. For open-ended drafting or code generation, where the model is expected to go beyond the context, it captures less of what users care about.

Why user feedback is not enough

The strongest objection is that users already report bad answers. Thumbs-down ratings and escalations are real signal, cheap to collect, and tied to actual harm.

They can also be sparse and skewed: rating is optional, raters self-select, and a rating may reflect tone or helpfulness rather than correctness. In the refund example, a customer told they are eligible for a refund they should not get has little reason to complain. Feedback is useful as a stratum to oversample for scoring and as a source of regression cases, but not as the primary quality monitor.

Turning failing traces into regression cases

The loop is short. A trace is flagged by an online score, a feedback signal, or a freshness check. A person confirms the failure. The input, the context or conditions that produced it, and the expected property become a new entry in the versioned dataset, and that entry becomes a CI case.

For the refund failure, that means two cases: an end-to-end case asserting that an opened item at 40 days is not eligible under the 30-day window, and a retrieval case asserting that superseded policy documents are not returned. The fix, removing or filtering superseded documents at index time, ships behind both. From then on, every prompt edit, model change, and index rebuild runs against a failure that has already happened once.

An LLM feature is observable when you can say what the trace stores, what the offline gate checks and at what thresholds, how many live answers are scored per stratum and how well the judge agrees with people, and how a flagged trace becomes a test. If you are asked to design an AI feature in a system design interview, those answers are the monitoring design. Latency and error-rate alerts still belong alongside them, but they cover delivery, and they would have reported the refund answer as healthy.

Share this post