Skip to main navigation Skip to search Skip to main content

LIGHTWEIGHT TEMPORAL CONTEXTUAL FINE-TUNING METHOD OF LARGE MULTIMODAL MODEL FOR VIDEO MOMENT RETRIEVAL

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Video moment retrieval (VMR) tasks require a comprehensive understanding of the video-language features of an input video, based on large multimodal models (LMMs). In this paper, we introduce temporal contextual prompts and provide contextual information into the LMM model to improve the generalization ability of the VMR. We integrate temporal contextual prompts into conventional prompts and obtain integrated tokens through an embedding module. The temporal contextual prompts are transformed into temporal conditional tokens, and these are fine-tuned as the embedding representation of the temporal correlation between the video and the given query text. In addition, we quantize the model and apply LoRA to demonstrate efficient learning for limited resources. Experimental results on Charades-STA show improvements of 3.78 percent in mIoU and 0.85 percent in [email protected] over state-of-the-art performance.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Image Processing, ICIP 2025 - Proceedings
PublisherIEEE Computer Society
Pages2880-2885
Number of pages6
ISBN (Electronic)9798331523794
DOIs
StatePublished - 2025
Event32nd IEEE International Conference on Image Processing, ICIP 2025 - Anchorage, United States
Duration: 14 Sep 202517 Sep 2025

Publication series

NameProceedings - International Conference on Image Processing, ICIP
ISSN (Print)1522-4880

Conference

Conference32nd IEEE International Conference on Image Processing, ICIP 2025
Country/TerritoryUnited States
CityAnchorage
Period14/09/2517/09/25

Bibliographical note

Publisher Copyright:
©2025 IEEE.

Keywords

  • Large Language Model
  • Moment Retrieval
  • Multimodal
  • Temporal Conditional Token
  • Temporal Contextual Fine-tuning

Fingerprint

Dive into the research topics of 'LIGHTWEIGHT TEMPORAL CONTEXTUAL FINE-TUNING METHOD OF LARGE MULTIMODAL MODEL FOR VIDEO MOMENT RETRIEVAL'. Together they form a unique fingerprint.

Cite this