Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T01. LLM agent orchestration and harness design
External research on how multiple language-model agents are planned, routed, and evaluated.
Show this topic only- T01-13Similar problemSince 2025
Why Do Multi-Agent LLM Systems Fail?
- T01-19Similar problemSince 2025
tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- T01-12Similar problemSince 2025
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- T01-7Similar problem
WebArena: A Realistic Web Environment for Building Autonomous Agents
- T01-8Similar problem
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- T01-16Similar problem
AgentBench: Evaluating LLMs as Agents
- T01-17Similar problem
GAIA: a benchmark for General AI Assistants
- T01-18Similar problem
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
T02. Recursive language models and long-context handling
Research on reading documents that do not fit in one context window.
Show this topic only- T02-3Similar problem
Walking Down the Memory Maze: Beyond Context Limit through Interactive Reading (MemWalker)
- T02-6Similar problem
Lost in the Middle: How Language Models Use Long Contexts
- T02-7Similar problem
RULER: What's the Real Context Size of Your Long-Context Language Models?
- T02-11Similar problem
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
T03. Tool use, function calling, and the Model Context Protocol
External research and specifications for letting a model call outside tools.
Show this topic only- T03-13Similar problemSince 2025
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
- T03-12Similar problemSince 2025
The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models
- T03-9Similar problemSince 2025
Tool Learning with Foundation Models
- T03-6Similar problem
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
T04. Planning, reflection, judging, and self-improving agents
Research on step-by-step reasoning, self-critique, and model-as-judge evaluation.
Show this topic only- T04-9Similar problem
Let's Verify Step by Step
- T04-11Similar problem
Large Language Models Cannot Self-Correct Reasoning Yet
- T04-16Similar problem
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
T05. Human review and approval workflows
Human-in-the-loop and approval research, including automation-induced complacency.
Show this topic only- T05-21Similar problemSince 2025
Human-In-the-Loop Software Development Agents (HULA)
- T05-1Similar problem
Ironies of Automation
- T05-2Similar problem
The Out-of-the-Loop Performance Problem and Level of Control in Automation
- T05-3Similar problem
Humans and Automation: Use, Misuse, Disuse, Abuse
- T05-5Similar problem
Complacency and Bias in Human Use of Automation: An Attentional Integration
- T05-13Similar problem
Trust in Automation: Designing for Appropriate Reliance
- T05-14Similar problem
Guidelines for Human-AI Interaction
- T05-15Similar problem
Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-Assisted Decision Making
- T05-16Similar problem
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
- T05-17Similar problem
To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- T05-18Similar problem
The flaws of policies requiring human oversight of government algorithms
T06. RAG, knowledge graphs, ontologies, and provenance
Retrieval-augmented generation and structured knowledge with traceable sources.
Show this topic only- T06-14Similar problemSince 2025
Enhancing retrieval-augmented generation for interoperable industrial knowledge representation and inference toward cognitive digital twins
- T06-2Similar problem
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- T06-4Similar problem
Retrieval-Augmented Generation for Large Language Models: A Survey
- T06-9Similar problem
Measuring Attribution in Natural Language Generation Models
- T06-12Similar problem
Knowledge Graphs
- T06-13Similar problem
A benchmark dataset with Knowledge Graph generation for Industry 4.0 production lines
T07. Structured output, schema validation, and document understanding
Research on forcing machine-checkable output and reading business documents.
Show this topic only- T07-11Similar problemSince 2025
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
- T07-3Similar problem
TableFormer: Table Structure Understanding with Transformers
- T07-5Similar problem
FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
- T07-6Similar problem
DocVQA: A Dataset for VQA on Document Images
- T07-12Similar problem
Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance
- T07-13Similar problem
Grammar-Aligned Decoding
T08. OCR, vision-language models, and industrial display reading
Research on reading meters, indicators, and shop-floor displays from images.
Show this topic only- T08-1Similar problemSince 2025
Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- T08-3Similar problem
Convolutional Neural Networks for Automatic Meter Reading
- T08-4Similar problem
Utilizing Smartphone-Based Machine Learning in Medical Monitor Data Collection: Seven Segment Digit Recognition
T09. Industrial time series, anomaly detection, and predictive maintenance
Sensor-driven fault detection, remaining useful life, and condition monitoring research.
Show this topic only- T09-1Similar problem
Current Time Series Anomaly Detection Benchmarks are Flawed and are Creating the Illusion of Progress
- T09-2Similar problem
Towards a Rigorous Evaluation of Time-series Anomaly Detection
- T09-3Similar problem
The Elephant in the Room: Towards A Reliable Time-Series Anomaly Detection Benchmark
- T09-4Similar problem
Volume Under the Surface: A New Accuracy Evaluation Measure for Time-Series Anomaly Detection
- T09-5Similar problem
A Review on Outlier/Anomaly Detection in Time Series Data
- T09-7Similar problem
Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
- T09-9Similar problem
A Dataset to Support Research in the Design of Secure Water Treatment Systems (SWaT)
- T09-13Similar problem
Damage Propagation Modeling for Aircraft Engine Run-to-Failure Simulation
- T09-14Similar problem
A review on machinery diagnostics and prognostics implementing condition-based maintenance
- T09-15Similar problem
Machinery health prognostics: A systematic review from data acquisition to RUL prediction
T10. AutoML, model selection, evaluation design, and experiment tracking
Research on choosing and validating models instead of shipping a single fitted model.
Show this topic only- T10-1Similar problem
Random Search for Hyper-Parameter Optimization
- T10-9Similar problem
AMLB: an AutoML Benchmark
- T10-10Similar problem
Statistical Comparisons of Classifiers over Multiple Data Sets
- T10-11Similar problem
On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation
- T10-12Similar problem
Evaluating time series forecasting models: An empirical study on performance estimation methods
T11. Object detection, multi-object tracking, and video understanding
Detection and tracking backbones behind camera-based safety and inspection work.
Show this topic only- T11-3.2Similar problem
Simple Online and Realtime Tracking with a Deep Association Metric (DeepSORT)
- T11-5.2Similar problem
SlowFast Networks for Video Recognition
T12. Video understanding and evidence selection with vision-language models
Research on explaining what a camera saw and pointing back to the evidence frame.
Show this topic only- T12-13Similar problemSince 2025
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- T12-6Similar problem
Can I Trust Your Answer? Visually Grounded Video Question Answering (NExT-GQA)
- T12-9Similar problem
QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries (Moment-DETR)
- T12-12Similar problem
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- T12-14Similar problem
Evaluating Object Hallucination in Large Vision-Language Models (POPE)
- T12-15Similar problem
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
T13. Edge AI, streaming inference, and resource scheduling
Running models near the machine under limited compute and latency budgets.
Show this topic only- T13-2Similar problem
Live Video Analytics at Scale with Approximation and Delay-Tolerance
- T13-3Similar problem
AWStream: Adaptive Wide-Area Streaming Analytics
- T13-4Similar problem
Chameleon: Scalable Adaptation of Video Analytics
- T13-5Similar problem
Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video Analytics
- T13-18Similar problem
MillWheel: Fault-Tolerant Stream Processing at Internet Scale
- T13-19Similar problem
The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing
T14. Industrial protocol translation, code generation, and program synthesis
Research on generating and checking the code that talks to plant equipment.
Show this topic only- T14-1Similar problem
Evaluating Large Language Models Trained on Code
- T14-2Similar problem
ChatGPT for PLC/DCS Control Logic Generation
- T14-6Similar problem
Automated Control Logic Test Case Generation using Large Language Models
- T14-7Similar problem
Automated generation of OPC UA information models - A review and outlook
- T14-15Similar problem
Program Synthesis with Large Language Models
T15. OPC UA, Asset Administration Shell, MQTT, and manufacturing interoperability
Specifications and research for describing equipment and moving its data.
Show this topic only- T15-2.1Similar problem
The Future of Industrial Communication: Automation Networks in the Era of the Internet of Things and Industry 4.0
- T15-2.2Similar problem
Insights into Mapping Solutions Based on OPC UA Information Model Applied to the Industry 4.0 Asset Administration Shell
- T15-2.4Similar problem
OPC UA versus ROS, DDS, and MQTT: Performance Evaluation of Industry 4.0 Protocols
- T15-2.9Similar problem
Streaming Machine Generated Data via the MQTT Sparkplug B Protocol for Smart Factory Operations
T16. Manufacturing knowledge graphs and semantic layers
Research on giving plant data a shared meaning across systems.
Show this topic only- T16-7Similar problemSince 2025
Fault Cause Identification across Manufacturing Lines through Ontology-Guided and Process-Aware FMEA Graph Learning with LLMs
- T16-3Similar problem
Knowledge Graphs in Manufacturing and Production: A Systematic Literature Review
- T16-12Similar problem
The Industry 4.0 Standards Landscape from a Semantic Integration Perspective
T17. Natural language to SQL and grounded report generation
Research on turning a question into a checked query and a sourced report.
Show this topic only- T17-5Similar problemSince 2025
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
- T17-1Similar problem
Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
- T17-2Similar problem
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
- T17-4Similar problem
Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs
- T17-17Similar problem
ToTTo: A Controlled Table-To-Text Generation Dataset
- T17-18Similar problem
QTSumm: Query-Focused Summarization over Tabular Data
T18. AI quality management, model monitoring, and audit trails
Research and standards for keeping a deployed model accountable over time.
Show this topic only- T18-1Similar problem
Hidden Technical Debt in Machine Learning Systems
- T18-7Similar problem
A Survey on Concept Drift Adaptation
- T18-8Similar problem
Learning under Concept Drift: A Review
- T18-10Similar problem
Operationalizing Machine Learning: An Interview Study
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology