← All research
Findings of EMNLP 2026Accepted

Can an AI explore scientific possibilities through patterns and edits?

I study whether an allowed edit to a word, composition, or protein sequence can make the final goal unreachable.

StartGoalA route remainsNo feasible completionThe next edit changes what remains possible.

Conceptual illustration of recoverability. This graph is not experimental data.

Research motivation

I was interested in whether LLMs could contribute to scientific research by recognizing patterns and editing candidate solutions. To study this, I looked at sequences of small changes and asked whether each change still left a way to reach the goal.

01 / The approach

Study design

We use a shared editing protocol across Word Ladder, an alloy composition proxy, and a restricted GB1 protein fitness landscape. Each task separates the admissibility of the next edit from the delayed objective. We then audit recoverability: whether at least one successful completion remains within the available edit budget. GB1 permits exact enumeration; the alloy setting uses bounded probes.

Inside the paper

One editing protocol, three kinds of search.

Enlarge
Original Figure 1: Word Ladder edits one letter, the alloy proxy transfers composition mass, and GB1 edits a protein sequence. Each follows partial artifact, local edit, and delayed objective. Evidence is exact on the finite Word Ladder and GB1 graphs; the alloy audit uses bounded probes and rounded states.

Figure 1 brings the three environments onto the same editing protocol. A locally admissible change and a successful final artifact are separate checks.

The same decision

At each step, the agent changes a partial artifact. The objective is evaluated after a sequence of edits.

Different audit limits

Read the evidence bands: Word Ladder and GB1 permit exact graph checks; the alloy proxy uses bounded probes and a rounded-state audit.

Ryu, Lee & Lee · Findings of EMNLP 2026 · Figure 1 from the accepted manuscript.

One editing protocol, three kinds of search.

Original Figure 1: Word Ladder edits one letter, the alloy proxy transfers composition mass, and GB1 edits a protein sequence. Each follows partial artifact, local edit, and delayed objective. Evidence is exact on the finite Word Ladder and GB1 graphs; the alloy audit uses bounded probes and rounded states.

Ryu, Lee & Lee · Findings of EMNLP 2026 · Figure 1 from the accepted manuscript. Open image ↗

02 / What we found

Findings

Across six models, many trajectories contain only well-formed, locally admissible actions but still fail the final objective. Among these surface-clean episodes, success is 16.0% in the alloy proxy and 6.8% in GB1. In the exact GB1 audit, 51.4% of surface-clean episodes cross a boundary beyond which the goal is no longer reachable within the remaining budget.

Scope & limitations

These are controlled editing environments. The alloy task is a proxy, and the GB1 analysis covers a restricted finite landscape. The findings diagnose obstacles to scientific search; they do not demonstrate autonomous scientific discovery or generalize directly to unrestricted materials and protein design.

Further questions

Can we learn representations that predict how an edit changes future possibilities, so an agent can preserve useful paths before they disappear?

Research directions
Publication details

Evaluating LLM Agents Beyond Local Edit Validity

Kunhee Ryu, Chi-Guhn Lee, and Keeheon Lee

Findings of EMNLP 2026 · Accepted

BibTeX citation
@inproceedings{ryu2026localvalidity,
  title={Evaluating LLM Agents Beyond Local Edit Validity},
  author={Ryu, Kunhee and Lee, Chi-Guhn and Lee, Keeheon},
  booktitle={Findings of the Association for Computational Linguistics: EMNLP 2026},
  year={2026},
  note={Accepted}
}
Next paperCheckpoint selection with limited validation data