Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T01. LLM agent orchestration and harness design
External research on how multiple language-model agents are planned, routed, and evaluated.
Show this topic only- T01-13Similar problemSince 2025
Why Do Multi-Agent LLM Systems Fail?
- T01-19Similar problemSince 2025
tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- T01-12Similar problemSince 2025
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
T03. Tool use, function calling, and the Model Context Protocol
External research and specifications for letting a model call outside tools.
Show this topic only- T03-13Similar problemSince 2025
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
- T03-12Similar problemSince 2025
The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models
- T03-9Similar problemSince 2025
Tool Learning with Foundation Models
T05. Human review and approval workflows
Human-in-the-loop and approval research, including automation-induced complacency.
Show this topic only- T05-21Similar problemSince 2025
Human-In-the-Loop Software Development Agents (HULA)
T06. RAG, knowledge graphs, ontologies, and provenance
Retrieval-augmented generation and structured knowledge with traceable sources.
Show this topic only- T06-14Similar problemSince 2025
Enhancing retrieval-augmented generation for interoperable industrial knowledge representation and inference toward cognitive digital twins
T07. Structured output, schema validation, and document understanding
Research on forcing machine-checkable output and reading business documents.
Show this topic only- T07-11Similar problemSince 2025
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
T08. OCR, vision-language models, and industrial display reading
Research on reading meters, indicators, and shop-floor displays from images.
Show this topic only- T08-1Similar problemSince 2025
Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
T12. Video understanding and evidence selection with vision-language models
Research on explaining what a camera saw and pointing back to the evidence frame.
Show this topic only- T12-13Similar problemSince 2025
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
T16. Manufacturing knowledge graphs and semantic layers
Research on giving plant data a shared meaning across systems.
Show this topic only- T16-7Similar problemSince 2025
Fault Cause Identification across Manufacturing Lines through Ontology-Guided and Process-Aware FMEA Graph Learning with LLMs
T17. Natural language to SQL and grounded report generation
Research on turning a question into a checked query and a sourced report.
Show this topic only- T17-5Similar problemSince 2025
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology