← All research
AAMAS 2026 · Main trackPublished

How can agents combine evidence without losing its provenance?

Study-centered agents extract, critique, and revise evidence before a statistical module synthesizes the results.

StudiesStudy agentsVerifyPoolCritique

Conceptual workflow. Agents verify evidence; a separate module performs statistical pooling.

01 / The approach

Study design

One agent is assigned to each primary study. Agents extract page-anchored data, exchange critiques, and revise disputed entries under a shared protocol. A separate statistical module performs pooling. We compare this workflow with a single-LLM pipeline using eight cardiology meta-analyses comprising 114 primary studies.

Inside the paper

From study-level evidence to a pooled estimate.

Enlarge
Original AutoMETA framework: document preprocessing, study-agent initialization, iterative extraction and critique, and statistical synthesis.

Figure 2 shows how study-centered agents extract and check evidence before a separate statistical module combines the results. Figure content is unchanged; surrounding page material was cropped.

Keep the study attached

Each agent works with one primary study and preserves page-anchored evidence for its extracted entries.

Check before pooling

Agents critique and revise disputed entries. Statistical synthesis is a separate step after that verification.

Lee & Ryu · AAMAS 2026 · Figure 2. Paper ↗ CC BY 4.0 ↗

From study-level evidence to a pooled estimate.

Original AutoMETA framework: document preprocessing, study-agent initialization, iterative extraction and critique, and statistical synthesis.

Lee & Ryu · AAMAS 2026 · Figure 2. Open image ↗

02 / What we found

Findings

Median relative effect-size error against human-authored references is 6.4% for AutoMETA, compared with 34.0% for the single-LLM baseline under the same protocol. Ablations show that critique and protocol enforcement both contribute to the result.

Scope & limitations

The evaluation uses a specific cardiology corpus. Stricter verification can exclude more studies and destabilize estimates of between-study variation. Agreement on a pooled estimate does not by itself establish complete or unbiased evidence coverage.

Further questions

When does another agent contribute useful evidence, and how should we evaluate the effects of verification on the final decision?

Research directions
Publication details

AutoMETA: A Multi-Agent LLM System for Autonomous Meta-Analysis

Keeheon Lee and Kunhee Ryu

AAMAS 2026 · Main track · Published

BibTeX citation
@inproceedings{lee2026autometa,
  title={AutoMETA: A Multi-Agent LLM System for Autonomous Meta-Analysis},
  author={Lee, Keeheon and Ryu, Kunhee},
  booktitle={Proceedings of AAMAS 2026},
  pages={2968--2977},
  year={2026},
  doi={10.65109/HXKA2256}
}
Next paperHow clinicians verify AI answers