What happens before learning becomes visible?
Small validation sets can assign the same score to different checkpoints. I study what those ties hide about generalization and model selection.
A continuous signal can distinguish checkpoints that share an exact-match score.
Conceptual curves with no numerical scale. Token likelihood is not a universal measure of understanding.
Reading wuxia novels, I noticed a distinction between lacking inner power and lacking the insight needed for a breakthrough. It reminded me of grokking: learning can precede visible generalization. This led me to ask what an exact-match score might miss while a language model is still learning.
Outer practice and inner cultivation accumulate. The next realm can still remain out of reach.
Power may already be there. What is missing is the insight that makes a way through the barrier visible.
Reaching a new realm made me think of grokking: a long period of learning before generalization becomes visible.
Study design
The paper investigates a concrete measurement problem: several checkpoints can receive the same exact-match score on a small labeled out-of-distribution validation set. We compare exact match, token likelihood, and the latest checkpoint among ties across semantic parsing and morphological reinflection settings. The central audit uses 600 draws from 30 training runs.
Findings
With eight labeled references, all early checkpoints score zero in 48% of draws, and the maximum exact-match score is tied in 57%. Token likelihood distinguishes some tied candidates. Choosing the latest tied checkpoint performs better on average for mostly improving trajectories. We propose a reporting protocol that records how ties are handled during checkpoint selection.
What a tied score can conceal.

These are the paper’s measured results. Panel A is a representative COGS trajectory; panels B and C summarize the checkpoint-selection audit. The criteria and reference budgets must be read together.
A / Learning signals
Orange token likelihood moves while blue exact match is still near zero. The signals reveal different aspects of the same trajectory.
B / Selection regret
Lower is better. The plot compares development EM, hard OOD EM, and soft OOD likelihood across labeled reference budgets.
C / Unresolved ties
Brown marks all-zero early windows; purple marks tied maxima. Pale bars count checkpoints distinguished only by token likelihood.
Ryu & Lee · Findings of EMNLP 2026 · Figure 1 from the accepted manuscript.
The wuxia analogy describes the origin of the question. The experiments do not establish separate mechanisms equivalent to inner power and insight. Token likelihood needs labeled target outputs, and it is not a universally superior selection rule. The paper concerns the resolution of checkpoint selection under a small labeled validation budget.
Out-of-Distribution Checkpoint Selection Has a Resolution Problem: Auditing Sparse Exact Match with Token Likelihood
Kunhee Ryu and Keeheon Lee
Findings of EMNLP 2026 · Accepted
BibTeX citation
@inproceedings{ryu2026resolution,
title={Out-of-Distribution Checkpoint Selection Has a Resolution Problem: Auditing Sparse Exact Match with Token Likelihood},
author={Ryu, Kunhee and Lee, Keeheon},
booktitle={Findings of the Association for Computational Linguistics: EMNLP 2026},
year={2026},
note={Accepted}
}