Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T12. Video understanding and evidence selection with vision-language models
Research on explaining what a camera saw and pointing back to the evidence frame.
- T12-2Structural parallelSince 2025
Adaptive Keyframe Sampling for Long Video Understanding
- T12-1Structural parallelSince 2025
Frame-Voyager: Learning to Query Frames for Video Large Language Models
- T12-13Similar problemSince 2025
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- T12-3Structural parallel
Self-Chained Image-Language Model for Video Localization and Question Answering (SeViLA)
- T12-4Structural parallel
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- T12-5Candidate approach
A Simple LLM Framework for Long-Range Video Question-Answering (LLoVi)
- T12-6Similar problem
Can I Trust Your Answer? Visually Grounded Video Question Answering (NExT-GQA)
- T12-7Structural parallel
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
- T12-8Structural parallel
VTimeLLM: Empower LLM to Grasp Video Moments
- T12-9Similar problem
QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries (Moment-DETR)
- T12-11Structural parallel
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- T12-12Similar problem
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- T12-14Similar problem
Evaluating Object Hallucination in Large Vision-Language Models (POPE)
- T12-15Similar problem
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology