Why a Better LLM Can Make Your AI Product Worse

Here is a failure mode worth testing for before your next model migration.

An agent searches whenever it is unsure. Nobody wrote that rule down. The prompt and a handful of examples shape that behavior, and over time the workflow quietly comes to depend on it.

Then the model is replaced with a stronger one. The new model handles more of those questions directly, so it chooses the search path less often. An offline evaluation focused on answer quality reads that as progress. Eventually a user asks about something whose answer changed after the training cutoff, and the model answers directly, fluently, and out of date.

No individual component malfunctioned. The prompt was followed, the search service was healthy, and the new model really was more capable than the one it replaced. The product still gave the worse answer.

One unwritten assumption disappeared: uncertainty used to trigger retrieval. In this workflow, the search step was doing more than compensating for weak reasoning. It was the mechanism that kept current information inside the context window, and it was gated on a behavior nobody had specified.

The same pattern appears well beyond retrieval. Prompts, parsers, validators, thresholds, timeouts, and tool descriptions form an implicit compatibility layer around the model. It was tuned against the behavioral distribution of the model you had, often without anyone writing that down. Replace the model and you have changed one side of a contract that was never written.

This is why the obvious objection, that all of this is just integration testing, only half lands. The API contract holds across the swap: same endpoint, same schema, valid responses. What moved is the behavioral contract, the distribution of what the model tends to do across many inputs. Integration tests that assert on outputs for fixed inputs will pass while that distribution shifts underneath them.

For a migration, I find it useful to check four places: what a request is allowed to spend, what information reaches the model, how the model is allowed to act, and how you check any of it.

A model upgrade is a system change

Three levels often get collapsed into “model quality,” and they sit at different boundaries. Capability is what a model does on a defined task under a defined evaluation. Application behavior is what it does inside your prompts, context, tool schemas, and retry policy. Product outcome is whether the user finished what they came to do, how long it took, and what it cost. Improvement does not pass between those boundaries for free.

A pipeline from user through context and retrieval, the model, validation, and tools to the outcome. A model benchmark measures only the model box, while swapping M1 for M2 propagates behavior changes into validation, tools, and the outcome.

Swapping M1 for M2 changes one box, and that box produces the input to everything downstream of it. Backend engineers already know this shape. You replace a query with a faster one and the service gets slower, because the new plan changed the cache hit pattern and moved load somewhere else.

The mistake is not trusting the benchmark. It is asking the benchmark a question it was never built to answer.

What a request is allowed to spend

A new model can change how much of a request’s latency and cost budget gets used. More reasoning tokens or fewer, longer outputs or shorter, an extra tool step or one less. The direction is not the point. The point is that your latency and cost budgets were set against the old distribution.

Say a single model call has a 1 percent chance of crossing a latency threshold. If those events were independent, a request that makes four calls would have roughly a 4 percent chance that at least one crosses it (1 − 0.99⁴ ≈ 3.9 percent). In practice they are often correlated, since a slow provider or a long prompt slows every call at once, so treat that as an illustration of counting. A sequential workflow also has an exposure the formula ignores, because the deadline applies to the sum, and four individually unremarkable calls can still exceed it. Both are reasons to measure the end-to-end distribution instead of inferring it from per-call numbers.

Dean and Barroso described the fan-out form of this problem in The Tail at Scale, where a request that depends on many parallel operations becomes sensitive to the slowest of them. The architecture is different, but the measurement lesson transfers. Once a request depends on many operations, mean component latency becomes a poor proxy for the latency the user experiences.

Cost moves more indirectly. If the new model costs more per resolved task, the compensation may come out of something else. Shorter context. Fewer retrieved documents. One less retry. A verification step quietly dropped. Any of those can give back more accuracy than the model gained, and none of them appear in the model comparison.

What to measure. End-to-end p95 and p99, not mean model latency. Model calls per completed task. Cost per successful task, not cost per token. Timeout and fallback rates, before and after.

What information reaches the model

A model can only reason over what it receives, and an upgrade may arrive with a larger context window that makes it possible to send more.

Liu et al. tested how context length and the position of relevant information affect model performance. In multi-document question answering, and for several models in key-value retrieval, performance varied substantially with where the relevant information sat in the context, tending to be higher near the beginning or the end. They also observed reader performance saturating before retriever recall did, so adding documents stopped helping before retrieval stopped improving.

Long-context models have advanced since those experiments, whose main runs used GPT-3.5-Turbo, Claude-1.3, MPT-30B-Instruct, and LongChat-13B. Read the result as evidence about a failure mode and not as a standing rule for current models. The narrower conclusion is the useful one. Supporting a long context window and using the information inside it effectively are different capabilities, which makes context construction part of your application’s quality and not something you delegate to the model.

  • If retrieval returns the wrong documents, a better generator can give you a more fluent wrong answer.
  • If you raise the context limit because the new model permits it, you have changed how evidence position and irrelevant context affect the answer, and you have not measured how.
  • If your prompt accumulated conflicting instructions across a year of patches, a stronger reasoner may resolve those conflicts differently than the old one did. Same prompt, different resolution, no diff to review.

What to measure. Grounding or attribution, not just answer quality. Retrieval quality at the k actually passed to the model. Accuracy as a function of evidence position. Whether adding more documents improves the end-to-end result or only retriever recall.

How the model is allowed to act

The opening example belongs here, and this is the part of the contract I would worry about most. With tool use the consequences become direct, because generated output proposes an execution path. Which tool to call, with what arguments, and whether to stop are all model outputs. Whether those proposals become actions is a runtime decision, and that runtime was tuned against the proposals the old model used to make.

Arguments carry the same risk in smaller packaging. A model may populate optional fields it used to omit, choose differently between two tools with overlapping descriptions, or pass a date as a natural-language string where the old one passed ISO 8601. Schema validation catches the malformed cases. It will not catch well-formed calls that mean something different from what they used to mean. Catching those takes validating meaning as well as shape.

None of this is a model defect. It is what happens when a probabilistic component sits behind a fixed interface. No one replaces a database driver because its benchmark improved and then skips testing the behavior it depends on, and a model that emits tool calls is a dependency of that kind.

What to measure. Tool-selection distribution per intent class. Argument validity and argument semantics against golden traces, since the second does not follow from the first. Steps per completed task. Escalation rate. Recovery rate after a failed call. Traffic share on paths that used to be rare.

How you check any of it

Offline scores can improve while end-to-end task success falls, often because the eval set sits at a narrower boundary than the product does. Answers judged in isolation, never after retrieval. Tool-call syntax validity standing in for workflow completion. Mean latency where the deadline is set by the tail.

Automated judges have a boundary of their own. Zheng et al. found several biases in LLM judges, including a preference for verbose answers, alongside high agreement with human preference in the setting they tested. If a judge rewards verbosity, your evaluation can favor behavior that your latency budget has reason to penalize. The two measurements then pull model selection in opposite directions.

Migrations also expose weaknesses that were already there. A tool path the eval set never covered. A fallback branch the old model almost never took. A retrieval problem hidden because the old model asked a clarifying question first. The new model changed the conditions under which those became visible, which is a different claim from saying it created them. Without that distinction, every post-migration regression gets logged as “the new model is worse,” the team rolls back, and nobody learns which constraint was binding.

Treat the upgrade as a hypothesis

“Model B will improve our product” is a hypothesis about a system. A benchmark result is evidence for one term in it. Before migrating, answer four questions in writing, and answer them before you see any results.

QuestionWhat to write down
Expected gainWhich failure class should shrink?
Regression budgetWhat are we willing to trade for it?
Decision metricWhich end-to-end result determines ship or rollback?
Hidden slicesWhere could the average lie?

The first one carries most of the weight. If you cannot name the failure class that should shrink, you are upgrading on vibes.

Then evaluate in layers instead of one pass. The layers expose different failure modes.

  1. Component eval isolates capability.
  2. Integrated replay runs representative traces through your actual retrieval, prompts, tools, validators, and policies. Include failure paths, not only happy paths. For an agent, model behavior shapes which paths get proposed, and so which ones the runtime ever sees.
  3. Production experiment measures the product outcome and the system budgets under load.

Skipping integrated replay leaves the gap between component capability and production outcome largely untested.

Sometimes the answer is still to migrate and then retune everything around the model. That is a legitimate outcome, and it makes the point: the unit of cost and evaluation is the integrated system, not the API name in your config.

Sometimes the model is the bottleneck

Sometimes the model is the binding constraint. If it misreads instructions, cannot perform the reasoning the task requires, or fails in a language you serve, then piling orchestration, retries, and more retrieval around it can make the system more expensive and more fragile without fixing the thing that is broken. The response might be a different model, a specialized tool for the step it cannot do, or routing that sends those cases somewhere else.

The discipline is telling that case apart from the others. If you can isolate a capability, show that it limits end-to-end performance, and show that improving it improves the product under realistic conditions, the evidence is strong. At that point, a model change is no longer a reflex. It is a justified intervention. If the argument is that the candidate leads a leaderboard, you know considerably less than it feels like you know.

So the question at migration time is not which model is best. It is which constraint currently limits this product, and whether the candidate relieves it.

When the model changes, test the system that changed with it.


I develop this system-level view of AI quality at more length in Beyond the Model.

Beyond the Model book cover

References

Dean, J., and Barroso, L. A. “The Tail at Scale.” Communications of the ACM, 56(2), 74-80, 2013.

Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics, 12, 157-173, 2024.

Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.” NeurIPS 2023 Datasets and Benchmarks Track.

Comments

One response to “Why a Better LLM Can Make Your AI Product Worse”

  1. […] moves into a region nobody measured. None of them changes the model. It is the pattern from Why a Better LLM Can Make Your AI Product Worse: the arrangement around a model is tuned against behavior nobody wrote down, and the calibration […]

Leave a Reply

Discover more from Dongsun Moon

Subscribe now to keep reading and get access to the full archive.

Continue reading