PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems arXiv is now an independent nonprofit! Learn more× # PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Jiayu Liu*, Qihan Lin*, Cheng Qian, Rui Wang, Emre Can Acikgoz, Xiaocheng Yang, Jiateng Liu, Zhenhailong Wang, Xiusi Chen, Heng Ji, Dilek Hakkani-Tür...
Core contribution
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems frames "PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems" as research in the Agents category. The central contribution should be read through the attached source evidence rather than as an announcement: what matters is the paper, benchmark, system design, or evaluation claim that can be inspected and reproduced.
Technical approach
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems arXiv is now an independent nonprofit! Learn more× # PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Jiayu Liu*, Qihan Lin*, Cheng Qian, Rui Wang, Emre Can Acikgoz, Xiaocheng Yang, Jiateng Liu, Zhenhailong Wang, Xiusi Chen, Heng Ji, Dilek Hakkani-Tür University of Illinois Urbana-Champaign {jiayul12,hengji,dilek}@illinois.edu Code Dataset Project Page ###### Abstract LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons. However, existing benchmarks rarely evaluate planning under retrieval-limited tool visibility. To address this gap, we introduce PlanBench-XL, an interactive benchmark o...
For research in the Agents category, the review should specifically look for agent loop or architecture, tool-use environment, planning horizon, task suite, success and failure modes.
Evaluation setup
- Learn more× # PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Jiayu Liu*, Qihan Lin*, Cheng Qian, Rui Wang, Emre Can Acikgoz, Xiaocheng Yang, Jiateng Liu, Zhenhailong Wang, Xiusi Chen, Heng Ji, Dilek Hakkani-Tür University of Illinois Urbana-Champaign {jiayul12,hengji,dilek}@illinois.edu Code Dataset Project Page ###### Abstract LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons.
- However, existing benchmarks rarely evaluate planning under retrieval-limited tool visibility.
- To address this gap, we introduce PlanBench-XL, an interactive benchmark of 327 retail tasks over 1,665 tools that tests whether agents can iteratively retrieve usable tools, invoke them to uncover intermediate evidence for subsequent calls toward the final goal.
- Experiments on ten leading LLMs show that massive-tool planning remains challenging: while GPT-5.4 achieves 51.90% accuracy in block-free settings, it collapses to 11.36% under the most severe blocking condition.
- Learn more× # PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Jiayu Liu*, Qihan Lin*, Cheng Qian, Rui Wang, Emre Can Acikgoz, Xiaocheng Yang, Jiateng Liu, Zhenhailong Wang, Xiusi Chen, Heng Ji, Dilek Hakkani-Tür University of Illinois Urbana-Champaign {jiayul12,hengji,dilek}@illinois.edu Code Dataset Project Page ###### Abstract LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons.
- However, existing benchmarks rarely evaluate planning under retrieval-limited tool visibility.
Results and metrics
The strongest quantitative or technical signals found in the supplied excerpts are listed above. Treat them as source claims until the canonical paper, project page, or code release is checked directly.
Reproducibility notes
At least one attached source appears to be a code, model, or benchmark repository. Confirm license, setup instructions, evaluation scripts, and whether the reported results can be reproduced from the public artifacts.
Limitations and caveats
This dossier should separate what the authors or source documents claim from what can be independently inferred. If the sources omit baseline selection, benchmark construction, failure cases, or deployment constraints, those omissions should remain visible in the public research page.
Why this matters for AI builders
For builders tracking agents work, the useful question is whether this changes what to test, how to evaluate systems, or which assumptions to revisit. The candidate should help readers decide whether to inspect the paper/project more deeply, not just understand that it exists.
Source trail
- PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems: https://arxiv.org/html/2606.22388v1 - PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems arXiv is now an independent nonprofit! Learn more× # PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Ji...
- JiayuJeff/PlanBench-XL: https://github.com/JiayuJeff/PlanBench-Xl - # JiayuJeff/PlanBench-XL Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems - Stars: 37 - Forks: 1 - Watchers: 37 - Open issues: 0 - Homepage: https://planbench-xl.github.i...
- Jiayu Liu, Qihan Lin, Cheng Qian, Rui Wang, Emre Can Acikgoz, Xiaomin Yang, et al.: https://doi.org/10.48550/arxiv.2606.22388 - # PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems arXiv (Cornell University). Published: 2026-06-21. Preprint. 0 citations. ## Authors - Jiayu Liu: h-index 0; 0 citations - Qihan Lin: h-index 0; 0 citatio...
Source Information
Kainotomic Team
Published Jul 29, 2026, 12:00 AM
By Kainotomic Team
Published Jul 29, 2026, 12:00 AM