Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T12. Video understanding and evidence selection with vision-language models
Research on explaining what a camera saw and pointing back to the evidence frame.
- T12-2Structural parallelSince 2025
Adaptive Keyframe Sampling for Long Video Understanding
- T12-1Structural parallelSince 2025
Frame-Voyager: Learning to Query Frames for Video Large Language Models
- T12-3Structural parallel
Self-Chained Image-Language Model for Video Localization and Question Answering (SeViLA)
- T12-4Structural parallel
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- T12-7Structural parallel
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
- T12-8Structural parallel
VTimeLLM: Empower LLM to Grasp Video Moments
- T12-11Structural parallel
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology