Research library

Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.

Search and filter

Paper list

T01. LLM agent orchestration and harness design

External research on how multiple language-model agents are planned, routed, and evaluated.

Show this topic only
  1. T01-13Similar problemSince 2025

    Why Do Multi-Agent LLM Systems Fail?

    Authors
    Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica
    Year
    2025
    Venue
    NeurIPS 2025 Datasets and Benchmarks Track spotlight

    Korean review of T01-13

  2. T01-19Similar problemSince 2025

    tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment

    Authors
    Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Karthik Narasimhan
    Year
    2026
    Venue
    arXiv 2025-06-09

    Korean review of T01-19

  3. T01-12Similar problemSince 2025

    tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

    Authors
    Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan
    Year
    2025
    Venue
    ICLR 2025 Poster

    Korean review of T01-12

T03. Tool use, function calling, and the Model Context Protocol

External research and specifications for letting a model call outside tools.

Show this topic only
  1. T03-13Similar problemSince 2025

    Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions

    Authors
    Xinyi Hou, Yanjie Zhao, Shenao Wang, Haoyu Wang
    Year
    2026
    Venue
    ACM Transactions on Software Engineering and Methodology (TOSEM), 3796519

    Korean review of T03-13

  2. T03-12Similar problemSince 2025

    The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models

    Authors
    Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, Joseph E. Gonzalez
    Year
    2025
    Venue
    ICML 2025

    Korean review of T03-12

  3. T03-9Similar problemSince 2025

    Tool Learning with Foundation Models

    Authors
    Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding and 41 others in total (corresponding: Zhiyuan Liu, Maosong Sun)
    Year
    2025
    Venue
    ACM Computing Surveys vol. 57 no. 4, 101, pp. 40

    Korean review of T03-9

T05. Human review and approval workflows

Human-in-the-loop and approval research, including automation-induced complacency.

Show this topic only
  1. T05-21Similar problemSince 2025

    Human-In-the-Loop Software Development Agents (HULA)

    Authors
    Wannita Takerngsaksiri and 9 others (Atlassian, Monash University)
    Year
    2025
    Venue
    ICSE-SEIP 2025

    Korean review of T05-21

T06. RAG, knowledge graphs, ontologies, and provenance

Retrieval-augmented generation and structured knowledge with traceable sources.

Show this topic only
  1. T06-14Similar problemSince 2025

    Enhancing retrieval-augmented generation for interoperable industrial knowledge representation and inference toward cognitive digital twins

    Authors
    Dachuan Shi, Jianzhang Li, Olga Meyer, Thomas Bauernhansl
    Year
    2025
    Venue
    Computers in Industry vol. 171, 2025, 104330

    Korean review of T06-14

T07. Structured output, schema validation, and document understanding

Research on forcing machine-checkable output and reading business documents.

Show this topic only
  1. T07-11Similar problemSince 2025

    JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models

    Authors
    Saibo Geng, Hudson Cooper, Michał Moskal, Samuel Jenkins, Julian Berman, Nathan Ranchin, Robert West, Eric Horvitz, Harsha Nori
    Year
    2025
    Venue
    arXiv

    Korean review of T07-11

T08. OCR, vision-language models, and industrial display reading

Research on reading meters, indicators, and shop-floor displays from images.

Show this topic only
  1. T08-1Similar problemSince 2025

    Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench

    Authors
    Fenfen Lin, Yesheng Liu, Haiyu Xu, Chen Yue, Zheqi He, Mingxuan Zhao, Miguel Hu Chen, Jiakang Liu, JG Yao, Xi Yang
    Year
    2026
    Venue
    arXiv (cs.CV, cs.AI)

    Korean review of T08-1

T12. Video understanding and evidence selection with vision-language models

Research on explaining what a camera saw and pointing back to the evidence frame.

Show this topic only
  1. T12-13Similar problemSince 2025

    Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    Authors
    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen
    Year
    2025
    Venue
    CVPR 2025, pp. 24108-24118

    Korean review of T12-13

T16. Manufacturing knowledge graphs and semantic layers

Research on giving plant data a shared meaning across systems.

Show this topic only
  1. T16-7Similar problemSince 2025

    Fault Cause Identification across Manufacturing Lines through Ontology-Guided and Process-Aware FMEA Graph Learning with LLMs

    Authors
    Sho Okazaki, Kohei Kaminishi, Takuma Fujiu, Yusheng Wang, Manu Sasidharan, Jun Ota
    Year
    2026
    Venue
    arXiv (cs.IR)

    Korean review of T16-7

T17. Natural language to SQL and grounded report generation

Research on turning a question into a checked query and a sourced report.

Show this topic only
  1. T17-5Similar problemSince 2025

    Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows

    Authors
    Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao
    Year
    2025
    Venue
    ICLR 2025 Oral

    Korean review of T17-5

From research to product use

Operating capabilities, pilots, and technologies in development are identified separately.

View technology