Research library

Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.

Search and filter

Paper list

T12. Video understanding and evidence selection with vision-language models

Research on explaining what a camera saw and pointing back to the evidence frame.

  1. T12-2Structural parallelSince 2025

    Adaptive Keyframe Sampling for Long Video Understanding

    Authors
    Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye
    Year
    2025
    Venue
    CVPR 2025

    Korean review of T12-2

  2. T12-1Structural parallelSince 2025

    Frame-Voyager: Learning to Query Frames for Video Large Language Models

    Authors
    Sicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen, Sheng Jin, Zhongrong Zuo, Xiaolei Xu, Zhenbang Sun, Bingni Zhang, Jiawei Wu, Hao Zhang, Qianru Sun
    Year
    2025
    Venue
    ICLR 2025

    Korean review of T12-1

  3. T12-13Similar problemSince 2025

    Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    Authors
    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen
    Year
    2025
    Venue
    CVPR 2025, pp. 24108-24118

    Korean review of T12-13

From research to product use

Operating capabilities, pilots, and technologies in development are identified separately.

View technology