AN ALGORITHM FOR VIDEO LECTURE SCENE SEGMENTATION BASED ON VISUAL FRAME EMBEDDING COMPARISON

Authors

DOI:

https://doi.org/10.14529/ctcr260201

Keywords:

video lecture, scene segmentation, video data segmentation, visual embeddings, multi-model data processing, transformer models, visual content analysis, automated processing of educational content

Abstract

With the rapid growth of educational content in the form of video lectures, the task of their automatic transformation into a written format often providing better comprehension has become increa¬singly relevant. Manual annotation of video lectures is highly labor-intensive, which necessitates the deve¬lopment of algorithmic methods for segmenting video lectures into semantically meaningful fragments based on visual analysis. Objective. To develop an algorithm for segmenting video lectures into scenes based on the comparison of visual embeddings of frames. The proposed approach is aimed at identifying temporal boundaries within a video lecture where visual content remains stable, allowing such intervals to be interpreted as scenes corresponding to logically complete fragments of instructional material. Materials and Methods. In modern research, automated processing of video lectures is often addressed using multimodal large language models capable of capturing relationships between audio and visual information. However, the use of such models is associated with limitations related to interpretability, computational complexity, and data requirements. In this study, a multi-model video processing approach is employed, based on the separate analysis of lecture modalities using specialized models. This approach enables more accurate processing by accounting for the specific characteristics of each data type. For visual analysis, transformer-based embedding models, specifically DINOv2 and CLIP, are used to obtain stable and semantically informative representations of frames, which are then compared to detect scene boundaries. Results. As a result of the study, a multi-stage algorithm for segmenting video lectures into scenes based on visual frame embeddings was developed. The best performance was achieved by methods based on the DINOv2 model. The segmentation quality was evaluated by comparing predicted scene boundaries with ground truth annotations using precision, recall, and F1-score metrics. Conclusion. The obtained metric values confirm the effectiveness of the proposed algorithm for automatic segmentation of video lectures into scenes.

Author Biographies

Milan E. Ismagulov, Yugra State University, Khanty-Mansiysk, Russia

3rd year postgraduate student in the field of 2.3.1 “Systems analysis, mana¬gement and information processing, statistics”, Engineering School of Digital Technologies, Yugra State University, Khanty-Mansiysk, Russia

Andrey V. Melnikov, Ugra Research Institute of Information Technologies, Khanty-Mansiysk, Russia

Dr. Sci. (Eng.), Prof., Engineering School of Digital Technologies, Yugra State University, Khanty-Mansiysk, Russia; Director, Ugra Research Institute of Information Technologies, Khanty-Mansiysk, Russia

Published

2026-05-07

Issue

Section

Informatics and Computer Engineering