Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T01. LLM agent orchestration and harness design
External research on how multiple language-model agents are planned, routed, and evaluated.
- T01-13Similar problemSince 2025
Why Do Multi-Agent LLM Systems Fail?
- T01-14Candidate approachSince 2025
Multi-Agent Collaboration via Evolving Orchestration
- T01-19Similar problemSince 2025
tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- T01-12Similar problemSince 2025
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- T01-1Structural parallel
ReAct: Synergizing Reasoning and Acting in Language Models
- T01-2Structural parallel
Toolformer: Language Models Can Teach Themselves to Use Tools
- T01-3Structural parallel
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
- T01-4Candidate approach
Generative Agents: Interactive Simulacra of Human Behavior
- T01-5Candidate approach
Reflexion: Language Agents with Verbal Reinforcement Learning
- T01-6Candidate approach
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- T01-7Similar problem
WebArena: A Realistic Web Environment for Building Autonomous Agents
- T01-8Similar problem
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- T01-9Structural parallel
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations
- T01-10Structural parallel
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- T01-11Structural parallel
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- T01-15Structural parallel
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- T01-16Similar problem
AgentBench: Evaluating LLMs as Agents
- T01-17Similar problem
GAIA: a benchmark for General AI Assistants
- T01-18Similar problem
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology