Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T01. LLM agent orchestration and harness design
External research on how multiple language-model agents are planned, routed, and evaluated.
- T01-13Similar problemSince 2025
Why Do Multi-Agent LLM Systems Fail?
- T01-19Similar problemSince 2025
tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- T01-12Similar problemSince 2025
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- T01-7Similar problem
WebArena: A Realistic Web Environment for Building Autonomous Agents
- T01-8Similar problem
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- T01-16Similar problem
AgentBench: Evaluating LLMs as Agents
- T01-17Similar problem
GAIA: a benchmark for General AI Assistants
- T01-18Similar problem
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology