Research library

Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.

Search and filter

Paper list

T12. Video understanding and evidence selection with vision-language models

Research on explaining what a camera saw and pointing back to the evidence frame.

  1. T12-2Structural parallelSince 2025

    Adaptive Keyframe Sampling for Long Video Understanding

    Authors
    Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye
    Year
    2025
    Venue
    CVPR 2025

    Korean review of T12-2

  2. T12-1Structural parallelSince 2025

    Frame-Voyager: Learning to Query Frames for Video Large Language Models

    Authors
    Sicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen, Sheng Jin, Zhongrong Zuo, Xiaolei Xu, Zhenbang Sun, Bingni Zhang, Jiawei Wu, Hao Zhang, Qianru Sun
    Year
    2025
    Venue
    ICLR 2025

    Korean review of T12-1

  3. T12-3Structural parallel

    Self-Chained Image-Language Model for Video Localization and Question Answering (SeViLA)

    Authors
    Shoubin Yu, Jaemin Cho, Prateek Yadav, Mohit Bansal
    Year
    2023
    Venue
    NeurIPS 2023

    Korean review of T12-3

  4. T12-4Structural parallel

    VideoAgent: Long-form Video Understanding with Large Language Model as Agent

    Authors
    Xiaohan Wang, Yuhui Zhang, Orr Zohar, Serena Yeung-Levy
    Year
    2024
    Venue
    ECCV 2024. Computer Vision - ECCV 2024, LNCS, pp. 58-76

    Korean review of T12-4

  5. T12-7Structural parallel

    TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

    Authors
    Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, Lu Hou
    Year
    2024
    Venue
    CVPR 2024

    Korean review of T12-7

  6. T12-8Structural parallel

    VTimeLLM: Empower LLM to Grasp Video Moments

    Authors
    Bin Huang, Xin Wang, Hong Chen, Zihan Song, Wenwu Zhu
    Year
    2024
    Venue
    CVPR 2024, pp. 14271-14280

    Korean review of T12-8

  7. T12-11Structural parallel

    BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

    Authors
    Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi
    Year
    2023
    Venue
    ICML 2023. PMLR vol. 202 pp. 19730-19742

    Korean review of T12-11

From research to product use

Operating capabilities, pilots, and technologies in development are identified separately.

View technology