Why Your Eval Picks the Model That Guesses.

Where confident guessing comes from, and the part of it you choose.

Suppose you are choosing a model for an assistant that answers questions from your team’s engineering documentation and runbooks. You have two candidates and an evaluation set of 2,000 questions with known answers. Here are the results, scaled to 100 questions.

Per 100 questionsCorrectWrong“I don’t know”Accuracy
Model A78101278%
Model B8218082%

On 2,000 questions, that four-point gap is 80 more correct answers. An accuracy ranking picks Model B, because accuracy counts an “I don’t know” exactly the way it counts a wrong answer.

Whether that is the right pick depends on what each outcome is worth to you. Score a correct answer at 1, a wrong answer at minus k, and an abstention at minus c. Here k is the cost of a confident wrong answer, measured in correct answers, and c is the cost of saying “I don’t know” and sending the user somewhere else. Per 100 questions, Model A scores 78 – 10k – 12c and Model B scores 82 – 18k, so Model A comes out ahead whenever k > 0.5 + 1.5c.

Cost of an abstention, cModel A wins when a wrong answer costs more than
00.5
0.20.8
0.51.25

The formula does not name a winner. It names the question the ranking skipped: what does a confident wrong answer cost you, compared with an honest “I don’t know”? For an assistant whose answers someone may follow in the middle of an incident, it is hard to argue that a wrong answer costs less than half of what a right one is worth. Even if an abstention costs half a correct answer, the bar is only 1.25.

But what if Model B knows which 12 answers it should abstain on?

Let Model B abstain on the 12 questions per 100 where it is least confident, and the comparison turns on a single number: how many of those 12 it would have gotten wrong.

Wrong answers among B’s 12 least confidentModel B with a threshold (correct / wrong / abstain)Compared with Model A
676 / 12 / 12Model A is better, for any nonnegative k and c
878 / 10 / 12Identical to Model A
1080 / 8 / 12Model B is better, for any nonnegative k and c

If Model B’s confidence carried no information, about 2 of those 12 would be wrong, in line with its 18 percent error rate. Tying Model A takes 8. Model B’s rank depends entirely on how well its confidence picks out its own mistakes.

Accuracy alone cannot tell you whether Model B can do that. If you record only accuracy, the information is gone, and an accuracy score gives no reason to record more. When “I don’t know” earns no more than a wrong answer, an abstention is just another error, and guessing gains an advantage. Whether a model’s confidence can make that call on your traffic is what the previous post examined.

The condition shows up upstream in training rewards, in the benchmarks that labs and buyers use to choose between models, and in the eval you own. Each acts on a model in a different way.

Rewards shape the weights

After pretraining, two common kinds of reward change a model’s weights: preferences collected from people, and automatic verifiers that check an answer. Both can lean toward answering.

Zhou and colleagues found that the preference data used to align models is biased against text that expresses uncertainty. The models they studied rarely expressed uncertainty unless asked, even when they were wrong. When asked, they favored expressions of certainty, and their confident answers were wrong 47 percent of the time on average. People relied on the answers whether or not the model had marked them as certain. The common factor between RLHF (reinforcement learning from human feedback) and DPO (direct preference optimization) is this data, not any particular reward model. One fits a reward model to the preferences and the other optimizes on the preference pairs directly, and both inherit the pressure against an honest “I don’t know” when the data ranks it below a confident answer.

Leng and colleagues looked at the reward itself. They attached a randomly chosen high or low confidence score to otherwise identical responses, and the two reward models in their main test, one built for PPO (proximal policy optimization) and a DPO model used as an implicit one, preferred the high score whether the response had originally been chosen or rejected. They also found models trained with PPO or DPO more overconfident than their counterparts after supervised fine-tuning (SFT) alone. That evidence covers numeric confidence the model writes out, mainly on open models of 7 and 8 billion parameters. It does not settle other ways of reading confidence, or frontier-scale models.

Correctness-only verifiers lean the same way by construction. A reward that checks only whether the final answer is right gives a lucky guess full credit and charges a wrong guess nothing extra, so a policy that answers everything is never worse off than one that holds back. Damani and colleagues show that this binary reward degrades calibration, and note that almost all successful applications of reinforcement learning for reasoning use exactly that kind of reward.

Not every reward has to work this way. The same paper introduces reinforcement learning with calibration rewards (RLCR), which adds a Brier score on the model’s stated confidence to the correctness reward, and proves that with a bounded proper scoring rule the optimum is both accurate and calibrated.

None of this means models were honest before post-training. Kalai and colleagues argue that pretraining on its own produces confident errors through statistical pressure, with no reward for guessing involved. What post-training can change is how well a model’s confidence tracks its errors. In the GPT-4 technical report, the pre-trained model was highly calibrated on a subset of MMLU, with an expected calibration error of 0.007, and after post-training the error was 0.074. The measurement at the end of this post found the same kind of shift inside one open recipe, although most of it came at the supervised fine-tuning step, before any preference or reward training.

Benchmarks select

A benchmark computes no gradients. It acts on a model through the choices people make with it.

Kalai and colleagues point out that most widely used benchmarks grade each answer as right or wrong. Of the ten popular benchmarks they reviewed, only one gave any credit for “I don’t know,” and only partial credit. An abstention scores zero, the same as a wrong answer, so a model that guesses whenever it is unsure can only raise its expected score. A lab that compares checkpoints, data mixes, and recipes on those numbers will tend to keep whichever variant answers more, and the variant it keeps becomes the starting point for the next recipe. That is how a preference expressed on a scoreboard can end up in the weights without the scoreboard ever touching them.

Some benchmarks keep the cases apart. SimpleQA grades each answer as correct, incorrect, or not attempted, and reports accuracy on attempted questions next to overall accuracy. Even so, its authors note that under the F-score combining the two, a model that scores below 50 percent should always answer when it is at least half sure, whatever a wrong answer costs where the model is used. Every headline number implies a price on errors, and when nobody chose that price, it tends to be low.

You make the same kind of choice when you adopt a vendor model because it leads a leaderboard, which the first post in this series argued tells you less than it feels like it does.

Your eval decides what ships

The eval you own is the last selection before users see a model, and the only one in this chain you can change this week.

The same condition turns up there in ordinary forms. A golden set scored by exact match gives “I don’t know” the same zero as a wrong answer. An LLM judge with a preference for longer answers, a bias the first post discussed, has no reason to prefer an answer that stops at “I don’t know.” A prompt A/B test picks the variant that answers more questions, because nobody wrote down what an abstention was worth. None of these teaches the model anything. Each one decides which model, prompt, or threshold goes to production.

Changing that starts with the record. For every item, keep whether the system answered, abstained, or broke the output format, and the confidence it acted on. With that record you can later draw how the error rate among answered items falls as the system answers less, the risk-coverage calculation that the previous post walks through. None of it can be recovered from an accuracy column.

Next, put prices on the outcomes. Set k and c from what actually happens after an answer: a wrong step executed during an incident, or a handoff to a person who then looks it up. Score with those numbers. The arithmetic from the opening also gives a rule for each question. If a model’s chance of being right is p, answering beats abstaining when p > (k – c) / (1 + k). With k = 1 and c = 0 the line sits at 0.5, and with k = 3 and c = 0 it sits at 0.75.

Then tell the model the rule. Kalai and colleagues propose that evaluations state an explicit confidence target in the instructions, for example: “Answer only if you are >t confident, since mistakes are penalized t/(1-t) points, while correct answers receive 1 point, and an answer of ‘I don’t know’ receives 0 points.” With c = 0 that is the same line as above, so a wrong answer that costs 3 points means answering only above 0.75. The peer-reviewed version of their paper in Nature calls these open-rubric evaluations, and pairs a threshold of 0.75 with a penalty of 3. In their tests, models abstained more as the stated threshold rose. In the measurement below, a prompt that gave the points but not the threshold raised abstention by only about 4 points, and the 7B models kept answering questions they got wrong about half the time while rating almost every answer 95 or higher. State the rule, then measure whether your model follows it, and whether its confidence deserves the rule.

Finally, compare candidates at operating points, or across the whole risk-coverage curve, instead of at one accuracy number. Model A against Model B at 78 and 82 percent compares two different operating points as if they were one: Model A answers 88 percent of the questions and Model B answers all of them. With confidence recorded, the useful comparison is Model A against Model B at every threshold B could have used.

Confidence is not one number

Recording confidence raises a question the accuracy column never did: which confidence? A model offers several, and they move for different reasons.

SignalWhat it isWhat moves itBefore you threshold on it
Choice-token probabilityProbability over the listed options, read from the model and normalizedPost-training, prompt and template, option order and labelsRecalibrate on held-out traffic
Verbalized numeric confidenceA number the model writes about its own answerElicitation prompt, preference and reward trainingMeasure calibration per task and model version
Linguistic confidenceWords like “likely” or “I’m certain”Style, preference data, SFTDo not threshold on it directly
Sample agreementHow often repeated samples land on the same answerTemperature, output diversity, decodingValidate it separately, and do not assume it is confidence

In a free-form answer, the probability of a single token is not the confidence of the whole answer, which is why the first row applies only when the options are listed.

Two findings that look contradictory show why the channel matters. Tian and colleagues found that for models fine-tuned with human feedback, a confidence the model writes out as a number was typically better calibrated than the model’s conditional probability of its answer, often cutting calibration error by half. Leng and colleagues, above, found that stated confidence in RLHF models ran too high. The first result compares two channels on the same models. The second compares one channel with accuracy, on different models, tasks, and prompts. A channel can beat another and still be overconfident. Calibration belongs to a channel measured on a task, and a result for one channel or task need not hold for another.

The last row deserves suspicion. Kirk and colleagues found that RLHF sharply reduces the diversity of outputs for the same input compared with supervised fine-tuning. If post-training narrows the output distribution, agreement across samples may rise even when correctness does not. In the measurement below it rose about as much as the share of correct answers did, which says nothing about your recipe. Check agreement against correctness on your own items before you let it stand in for confidence.

What changed across one open post-training recipe

Published comparisons usually cover one step, such as SFT against RLHF, or set models from different recipes side by side. To see every stage of a single recipe, I ran the same questions through four public checkpoints of Olmo 3 7B from Ai2: the base model, then the Instruct model after SFT, after DPO, and after reinforcement learning, which produced the released Instruct model. Ai2 publishes every stage, its training data, and a technical report, which makes this one of the few recipes where such a comparison is possible. Two details of the recipe matter for reading the results. According to the technical report, the Instruct SFT stage started from Ai2’s reasoning-focused Think SFT model, not directly from the base model, so the first step is more than SFT alone. And the reinforcement learning stage mixes verifier rewards for math, code, and instruction following with a language-model judge for general chat, so it is not a clean test of either kind of reward.

The first test used 1,000 four-option questions from MMLU-Redux, a hand-checked subset of MMLU, using only questions its annotators found error-free. Each was asked in four option orders. Instead of parsing the model’s text, I read the probability it put on each option letter.

Accuracy did not improve across the recipe. It fell 3 points, from 63.7 to 60.5 percent, all of it at the first step. Mean confidence rose from 63 to 85 percent. The base model’s expected calibration error was 0.045 and the final model’s was 0.252, with most of the rise also at the first step. Rescaling the final model’s probabilities with one temperature, fitted on half the questions and checked on the other half, brought its error down to about the level the base model reaches under the same procedure, so much of the calibration error could be repaired by rescaling the probabilities. The ranking was another matter: how well confidence separated right answers from wrong ones also fell, from an AUROC of 0.81 to 0.77, and no single temperature can repair that.

The second test used 493 short factual questions from TriviaQA. Each post-trained model gave an answer with a confidence from 0 to 100, or said “I don’t know,” and answers were graded automatically against TriviaQA’s list of accepted answers.

With no scoring rule stated, the SFT model said “I don’t know” to 61 percent of the questions. After preference optimization that fell to 42 percent, and wrong answers rose from 19 to 34 percent. The reinforcement learning stage won some of it back, to 48 and 29 percent, but not to where SFT had been. When the models did answer, they almost always rated themselves 95 to 100, while only 41 to 51 percent of those answers matched an accepted answer.

I also sampled five more answers per question at temperature 0.7. The share that matched the main answer went from 75 to 78 percent between the SFT and final models, about as much as the share of correct answers rose, and the difference was within noise.

Then the same questions came with the rule stated: a correct answer earns 1 point, a wrong one loses 3, and “I don’t know” scores 0. The prompt gave the points, not the threshold. By the arithmetic above, a model should then answer only when it is more than 75 percent sure. Abstention rose by about 4 points. The questions the models still answered matched an accepted answer only 43 to 53 percent of the time, so under the rule they had just been given, every stage scored worse than it would have by saying “I don’t know” to everything. Yet they rated those answers about 96 on average, so by their own numbers they were well above the line. A stated rule only helps if the confidence behind it is calibrated.

Public benchmark questions may have been in the training data. I checked all but the 54 questions shorter than six words against the four public datasets used to post-train these checkpoints. A question counted as overlapping if any 13 consecutive words from it, or all of it when it was shorter, appeared word for word in a training example, ignoring case and punctuation. Of the 1,493, 37 overlapped, and removing them changed none of the conclusions. I did not check the pretraining and midtraining data.

This is an observational comparison across the stages of one public recipe, not a causal experiment. Each stage changes the objective and the data together, so no difference here can be assigned to DPO or to reinforcement learning as such, and the changes did not all run in one direction. The models have 7 billion parameters, the questions are in English, and nothing guarantees that frontier models behave the same way. The questions also test what a model knows, while the assistant in the opening mostly needs to notice when its documents do not hold the answer, and automatic grading misses some correct answers phrased at length and counts hedged answers as abstentions.

What the comparison can show is whether the pattern the literature describes appears in one fully documented recipe. For the probabilities on the option letters, it does, but most of the drift came at the first step, before any preference or reward training, and in this recipe that step includes the reasoning warm start. Stated confidence, which I measured only from that step on, was already near the top of the scale and stayed there. Preference optimization added to the drift in the option probabilities and cut abstention. Reinforcement learning barely moved the option probabilities and gave some abstention back. In this recipe, then, the reward stages were not where most of the calibration was lost, which is one more reason to measure the model you ship instead of inferring its calibration from its recipe.

Setup: the four checkpoints at their released weights, run locally in bf16 with MLX on one laptop. Option probabilities were read from the next token after the prompt. The chat models got their own chat template with the same system message (“You are a helpful assistant.”) because the template otherwise inserts its own, and the base model got a plain completion prompt that ends where the answer letter goes. Giving all four models the plain prompt moved calibration error the same way, though in that format the chat models put much of their probability outside the option letters. Calibration error used 15 equal-width bins, computed for each option order and averaged. Short answers were generated greedily with at most 48 new tokens. Questions were drawn with a fixed seed, round-robin over the 57 MMLU-Redux subjects and from the TriviaQA validation split without context.

A model does not have to learn to guess for guessing to reach production. Training can shape it, a benchmark can select it, and your eval can ship it.


Why evaluation works as the training surface of a production system even though it computes no gradients is the subject of Chapter 11, “Evaluation Is the New Training,” and Section 12.4 follows the same incentive into the weights, both in Beyond the Model.

Beyond the Model book cover

References

Ai2. Model cards for Olmo-3-1025-7B, Olmo-3-7B-Instruct-SFT, Olmo-3-7B-Instruct-DPO, and Olmo-3-7B-Instruct. 2025.

Damani, M., Puri, I., Slocum, S., Shenfeld, I., Choshen, L., Kim, Y., and Andreas, J. “Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty.” ICLR 2026. https://arxiv.org/abs/2507.16806

Gema, A. P., Leang, J. O. J., Hong, G., Devoto, A., Mancino, A. C. M., Saxena, R., He, X., Zhao, Y., Du, X., Ghasemi Madani, M. R., Barale, C., McHardy, R., Harris, J., Kaddour, J., Van Krieken, E., and Minervini, P. “Are We Done with MMLU?” NAACL 2025, pages 5069-5096. https://aclanthology.org/2025.naacl-long.262/

Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. “TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.” ACL 2017, pages 1601-1611. https://aclanthology.org/P17-1147/

Kalai, A. T., Nachum, O., Vempala, S. S., and Zhang, E. “Why Language Models Hallucinate.” 2025. https://arxiv.org/abs/2509.04664

Kalai, A. T., Nachum, O., Vempala, S. S., and Zhang, E. “Evaluating Large Language Models for Accuracy Incentivizes Hallucinations.” Nature 653, 1047-1051, 2026. https://www.nature.com/articles/s41586-026-10549-w

Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., and Raileanu, R. “Understanding the Effects of RLHF on LLM Generalisation and Diversity.” ICLR 2024. https://arxiv.org/abs/2310.06452

Leng, J., Huang, C., Zhu, B., and Huang, J. “Taming Overconfidence in LLMs: Reward Calibration in RLHF.” ICLR 2025. https://arxiv.org/abs/2410.09724

OpenAI. “GPT-4 Technical Report.” 2023. https://arxiv.org/abs/2303.08774

Team Olmo. “Olmo 3.” Technical report, 2025. https://arxiv.org/abs/2512.13961

Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D. “Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.” EMNLP 2023, pages 5433-5442. https://aclanthology.org/2023.emnlp-main.330/

Wei, J., Karina, N., Chung, H. W., Jiao, Y. J., Papay, S., Glaese, A., Schulman, J., and Fedus, W. “Measuring Short-Form Factuality in Large Language Models.” 2024. https://arxiv.org/abs/2411.04368

Zhou, K., Hwang, J. D., Ren, X., and Sap, M. “Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty.” ACL 2024, pages 3623-3643. https://aclanthology.org/2024.acl-long.198/

Comments

Leave a Reply

Discover more from Dongsun Moon

Subscribe now to keep reading and get access to the full archive.

Continue reading