Pareto logo

LEAPBench:
How efficiently do LLMs learn?

LEAPBench evaluates twelve frontier LLMs across 55 optimization tasks to identify the best learners.

Marilyn Zhang Tianfeng Chen Fabián Barzuna Ankita Rathod Mark E. Whiting
Explore results Read the paper

Two models can reach the same answer,
but one takes far more tries to get there

Winner changes (25) Winner holds (30)

What models are the most efficient?

The leaderboard below ranks every model — including Claude Opus 4.8 — by how efficiently it searches: how often it beats the classical GP-UCB baseline across the full 30-iteration trajectory (best-so-far AUC), its overall efficiency relative to that baseline, and whether scientific context helps. Click any model to view its learning curve over the 30-iteration budget.

By-model comparison

ModelClick a row to learn more
Tasks scored
Beats baseline
Efficiency vs GP-UCB
Scientific context

Explore the task set

Scientific optimization tasks differ widely in how models approach them. Some tasks yield the same winner regardless of the metric used, while others reveal sharp disagreements between final scores and path efficiency.

TaskClick a row to learn more
Winner changes?
Beats baseline?
Does context help?
Endpoint winner
Trajectory winner

Citation

When referencing LEAPBench in academic or technical work, use the BibTeX entry below.

BibTeX
@misc{zhang2026leap,
  title         = {LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design},
  author        = {Marilyn Zhang and Tianfeng Chen and Fabián Barzuna and Ankita Rathod and Mark E. Whiting},
  year          = {2026},
  eprint        = {2605.15341},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2605.15341}
}