Quality Locus is a diagnostic coordinate system for production AI quality. It locates which part of a deployed AI system currently limits the quality users get: the model, or one of the six surfaces around it. The term was introduced by Dongsun Moon in Beyond the Model: A Systems Theory of Modern AI Quality (2026).
Production quality is not contained in the deployed model. It is assembled across the system that prepares, surrounds, measures, and controls it.
Beyond the Model, Section 0.3

Definition
A Quality Locus is a place where the quality of a production AI system can be observed, verified, and changed. The term names both a single such surface and the six-surface diagnostic map as a whole. When a system disappoints, the map turns the question “which model should we buy?” into a different one: where did the quality go, which layer is failing, and which surface do we not yet measure, own, or control?
The deployed model sits at the center of the map but is not one of the six surfaces. It is the baseline the surfaces are tested against, and it carries a standing hypothesis that capability itself is the limit. A diagnosis sets that hypothesis against the alternatives instead of assuming it.
The six surfaces
| Surface | What it covers | In the book |
|---|---|---|
| Data | The corpus and examples used to train, tune, or otherwise shape the deployed model. Evidence assembled at request time belongs to memory, runtime, and inference instead. | Chapter 6 |
| Training | The process that produced the deployed model, and the one surface a running system cannot reach into. Changing it means training or procuring a different model. | Chapters 6, 7 |
| Inference | Where the frozen model runs: test-time compute, decoding policy, and the thinking budget. Each moves quality with the weights untouched. | Chapter 8 |
| Memory | What the system remembers, forgets, and retrieves, including retrieval ranking and external stores. | Chapter 9 |
| Runtime | Where model output meets the outside world. It opens into five sub-surfaces: orchestration, tool and action execution, validation, retrieval ranking, and context assembly. The prompt is a control artifact at the inference-runtime boundary, not a seventh surface. | Chapters 10, 13 |
| Evaluation | What measures the system, and decides what it can detect, compare, reject, ship, and improve. Permanently staffed work, not a gate cleared once at launch. | Chapter 11 |
The surfaces overlap on purpose. Retrieval ranking belongs to both memory and runtime. A diagnostic map has no reason to partition the system it describes.
How to use it
Before treating a model swap as the fix for a quality problem, set two hypotheses against each other. One says the model’s capability is the limit. The other says the cause lives in a surrounding surface. Then run the smallest, safest experiment that tells them apart. Sometimes a head-to-head model test is the sharpest experiment available, and sometimes the model really is the limit. The claim is about ordering: frame the question before reaching for a new model.
In practice the map becomes a short triage. Is this the model? The context being assembled? The retrieval? The tool or action being executed? The validation? An evaluation case nobody wrote? Repeated, the triage becomes a loop: define what passing means, raise competing hypotheses, find the measurement that separates them, intervene where doing so narrows the field most, and keep what you learn.
Five propositions
The book closes with five propositions that follow from the map. They form a scoped explanatory frame, not mathematical theorems.
- Composition. Production quality is jointly produced by the model and the system around it, and neither alone determines the outcome.
- Displacement. The symptom you observe and the surface that causes it frequently sit in different layers. A prompt symptom can reflect a retrieval cause.
- Constraint. A surface that sits below what the contract requires sets a ceiling on the whole system, and improving the model does not raise that ceiling.
- Reallocation. Improving one surface usually reveals the next bottleneck. The need for system design relocates rather than retires.
- Boundary. Some failures are genuine model bottlenecks. Telling them apart from system bottlenecks is the diagnostic task itself. The verdict is earned by a counterfactual that changes the model and nothing else.
What it is not
It is not an argument against better models. Capability is necessary, and the model is often the most visible source of it and one of the most expensive components a system runs. The Quality Locus disputes one belief only: that a good model finishes the job. The label is new. Many of the underlying ideas, that systems have properties no component has and that a symptom and its cause can sit in different places, are older, and the book traces them.
Posts that apply it
- Why a Better LLM Can Make Your AI Product Worse. A stronger model can break behavior the surrounding system was quietly tuned for.
- Calibration Does Not Ship with the Model. Confidence has to be measured on your own traffic.
- Why Your Eval Picks the Model That Guesses. The evaluation surface decides what ships.
How to cite
First published on 25 September 2026 in Beyond the Model.
Moon, D. (2026). Beyond the Model: A Systems Theory of Modern AI Quality. Beyond Moon Press. ISBN 9781067722326. The Quality Locus is defined in Section 0.3 and Appendix E.
@book{moon2026beyond, author = {Moon, Dongsun}, title = {Beyond the Model: A Systems Theory of Modern AI Quality}, publisher = {Beyond Moon Press}, year = {2026}, isbn = {9781067722326}, url = {https://dongsunmoon.com/quality-locus/}, note = {Introduces the Quality Locus (Section 0.3, Appendix E)}}