← All research
Findings of EMNLP 2026Accepted

What happens before learning becomes visible?

Small validation sets can assign the same score to different checkpoints. I study what those ties hide about generalization and model selection.

Learning signalTraining →Same exact-match scoreExact matchToken likelihood

Conceptual curves with no numerical scale. Token likelihood is not a universal measure of understanding.

Research motivation

Reading wuxia novels, I noticed a distinction between lacking inner power and lacking the insight needed for a breakthrough. It reminded me of grokking: learning can precede visible generalization. This led me to ask what an exact-match score might miss while a language model is still learning.

Cultivation & breakthrough
Outer practice + inner cultivationThe next realmThe barrier
Training the body. Cultivating inner power.

Outer practice and inner cultivation accumulate. The next realm can still remain out of reach.

A personal analogy behind the research.
01 / The approach

Study design

The paper investigates a concrete measurement problem: several checkpoints can receive the same exact-match score on a small labeled out-of-distribution validation set. We compare exact match, token likelihood, and the latest checkpoint among ties across semantic parsing and morphological reinflection settings. The central audit uses 600 draws from 30 training runs.

02 / What we found

Findings

With eight labeled references, all early checkpoints score zero in 48% of draws, and the maximum exact-match score is tied in 57%. Token likelihood distinguishes some tied candidates. Choosing the latest tied checkpoint performs better on average for mostly improving trajectories. We propose a reporting protocol that records how ties are handled during checkpoint selection.

Inside the paper

What a tied score can conceal.

Enlarge
Original Figure 1 with three measured panels: normalized token log probability rises before substantial exact-match improvement in a representative COGS run; checkpoint-selection regret varies with labeled OOD budget; early-window zero scores and tied maxima become less frequent as the budget increases.

These are the paper’s measured results. Panel A is a representative COGS trajectory; panels B and C summarize the checkpoint-selection audit. The criteria and reference budgets must be read together.

A / Learning signals

Orange token likelihood moves while blue exact match is still near zero. The signals reveal different aspects of the same trajectory.

B / Selection regret

Lower is better. The plot compares development EM, hard OOD EM, and soft OOD likelihood across labeled reference budgets.

C / Unresolved ties

Brown marks all-zero early windows; purple marks tied maxima. Pale bars count checkpoints distinguished only by token likelihood.

Ryu & Lee · Findings of EMNLP 2026 · Figure 1 from the accepted manuscript.

What a tied score can conceal.

Original Figure 1 with three measured panels: normalized token log probability rises before substantial exact-match improvement in a representative COGS run; checkpoint-selection regret varies with labeled OOD budget; early-window zero scores and tied maxima become less frequent as the budget increases.

Ryu & Lee · Findings of EMNLP 2026 · Figure 1 from the accepted manuscript. Open image ↗

Scope & limitations

The wuxia analogy describes the origin of the question. The experiments do not establish separate mechanisms equivalent to inner power and insight. Token likelihood needs labeled target outputs, and it is not a universally superior selection rule. The paper concerns the resolution of checkpoint selection under a small labeled validation budget.

Further questions

Which representations make new relationships learnable, and which measurements can tell us whether those representations will support transfer?

Research directions
Publication details

Out-of-Distribution Checkpoint Selection Has a Resolution Problem: Auditing Sparse Exact Match with Token Likelihood

Kunhee Ryu and Keeheon Lee

Findings of EMNLP 2026 · Accepted

BibTeX citation
@inproceedings{ryu2026resolution,
  title={Out-of-Distribution Checkpoint Selection Has a Resolution Problem: Auditing Sparse Exact Match with Token Likelihood},
  author={Ryu, Kunhee and Lee, Keeheon},
  booktitle={Findings of the Association for Computational Linguistics: EMNLP 2026},
  year={2026},
  note={Accepted}
}
Next paperMulti-agent meta-analysis