Research library

Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.

Search and filter

Paper list

T01. LLM agent orchestration and harness design

External research on how multiple language-model agents are planned, routed, and evaluated.

  1. T01-13Similar problemSince 2025

    Why Do Multi-Agent LLM Systems Fail?

    Authors
    Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica
    Year
    2025
    Venue
    NeurIPS 2025 Datasets and Benchmarks Track spotlight

    Korean review of T01-13

  2. T01-19Similar problemSince 2025

    tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment

    Authors
    Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Karthik Narasimhan
    Year
    2026
    Venue
    arXiv 2025-06-09

    Korean review of T01-19

  3. T01-12Similar problemSince 2025

    tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

    Authors
    Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan
    Year
    2025
    Venue
    ICLR 2025 Poster

    Korean review of T01-12

  4. T01-7Similar problem

    WebArena: A Realistic Web Environment for Building Autonomous Agents

    Authors
    Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, Graham Neubig
    Year
    2024
    Venue
    ICLR 2024

    Korean review of T01-7

  5. T01-8Similar problem

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    Authors
    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, Karthik Narasimhan
    Year
    2024
    Venue
    ICLR 2024

    Korean review of T01-8

  6. T01-16Similar problem

    AgentBench: Evaluating LLMs as Agents

    Authors
    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, Jie Tang
    Year
    2024
    Venue
    ICLR 2024

    Korean review of T01-16

  7. T01-17Similar problem

    GAIA: a benchmark for General AI Assistants

    Authors
    Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, Thomas Scialom
    Year
    2024
    Venue
    ICLR 2024

    Korean review of T01-17

  8. T01-18Similar problem

    OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    Authors
    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, Tao Yu
    Year
    2024
    Venue
    NeurIPS 2024

    Korean review of T01-18

From research to product use

Operating capabilities, pilots, and technologies in development are identified separately.

View technology