Two models can reach the same answer,
but one takes far more tries to get there
What models are the most efficient?
The leaderboard below ranks every model — including Claude Opus 4.8 — by how efficiently it searches: how often it beats the classical GP-UCB baseline across the full 30-iteration trajectory (best-so-far AUC), its overall efficiency relative to that baseline, and whether scientific context helps. Click any model to view its learning curve over the 30-iteration budget.
By-model comparison
| ModelClick a row to learn more | Tasks scored |
Beats baseline |
Efficiency vs GP-UCB |
Scientific context |
|---|
Explore the task set
Scientific optimization tasks differ widely in how models approach them. Some tasks yield the same winner regardless of the metric used, while others reveal sharp disagreements between final scores and path efficiency.
| TaskClick a row to learn more | Winner changes? |
Beats baseline? |
Does context help? |
Endpoint winner |
Trajectory winner |
|---|
Citation
When referencing LEAPBench in academic or technical work, use the BibTeX entry below.
@misc{zhang2026leap,
title = {LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design},
author = {Marilyn Zhang and Tianfeng Chen and Fabián Barzuna and Ankita Rathod and Mark E. Whiting},
year = {2026},
eprint = {2605.15341},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2605.15341}
}