The First Token Is Not the Verdict: Hidden Costs of Reading LLM

Overview


A growing class of evaluation harnesses—especially those built around constrained decoding and likelihood scoring—reads an LLM judge's verdict directly from the logits of its first generated token. This readout is cheap, requires no autoregressive generation, and is trivial to instrument. In a paper submitted on 4 Sep 2026 (arXiv:2610.00054, cs.CL), Gnaneswar Villuri, Hashmath Shaik, and Alex Doboli show that this shortcut systematically distorts position bias, and that the distortion runs in a single direction: the method overstates position bias in every condition tested, meaning any figure obtained this way behaves as an upper bound rather than an estimate.


The Mechanism


The core problem is an unexamined assumption: that a judge always leads with a verdict token. In practice, it often does not.


  • Across three Qwen3 judges, judges failed to lead with a verdict token on 12% to 49% of evaluated pairs.
  • For Llama-3.1-8B and Phi-3.5-mini, the same rate stayed under 3%.

When a verdict token is absent from the first position, forcing a read from the initial logits returns whichever response was shown first rather than an actual judgment. Pooled across the 924 pairs where a judge did not commit to a verdict, this forced read flips on 89.7% of pairs when the two responses are swapped—compared with only 47.5% when the read is taken after full generation (paired difference +0.422, 95% CI [+0.365, +0.467]).


What the Distortion Actually Affects


The distortion is highly specific to what is being measured. It shifts position-bias estimates by as much as 42 points, yet moves judge accuracy by under one point in seven of ten conditions. The practical implication is that this measurement artifact misleads whoever audits a judge—researchers benchmarking bias and reliability—rather than whoever simply uses one to score model outputs.


A Second, Smaller Failure Mode


Even when a judge does lead with a verdict token, a secondary failure can occur: the model sometimes opens with one letter and then reasons its way to the other. This happens on 0 to 5.5% of pairs, at a rate uncorrelated with the judge's overall compliance, making it difficult to predict or filter out from compliance metrics alone.


Recommendation


The authors propose a lightweight diagnostic: report the rate at which a judge leads with a verdict token. This metric costs a single forward pass and requires no labels, making it practical to compute alongside any position-bias figure. In the context of 2026 evaluation practice—where LLM-as-a-judge pipelines are increasingly used for benchmarking, ranking, and safety auditing—treating first-token readouts as ground truth without this diagnostic risks publishing inflated bias numbers that never reflect the judge's actual behavior.

via ArXiv CL+LG

Related