Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T12. Video understanding and evidence selection with vision-language models
Research on explaining what a camera saw and pointing back to the evidence frame.
- T12-13Similar problemSince 2025
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- T12-6Similar problem
Can I Trust Your Answer? Visually Grounded Video Question Answering (NExT-GQA)
- T12-9Similar problem
QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries (Moment-DETR)
- T12-12Similar problem
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- T12-14Similar problem
Evaluating Object Hallucination in Large Vision-Language Models (POPE)
- T12-15Similar problem
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology