Novel multimodal fusion approach enhances abnormal behavior detection, suggesting improved accuracy and efficiency.
Aiming at the problems of spatiotemporal alignment, insufficient modal interaction, and weak modeling of long-term dependencies in abnormal behavior detection based on video and location, a multi-modal feature fusion method for abnormal behavior detection based on attention mechanisms is proposed. The proposed method consists of three modules: feature extraction, multi-modal feature fusion, and anomaly detection. Firstly, in the feature extraction module, we utilize the ViViT (Video Vision Transformer) model to extract video features, capturing spatiotemporal action continuity through 3D block embedding; we utilize the ST-GCN (Spatial Temporal Graph Convolutional Networks) to extract localization features, constructing the spatiotemporal graph with the location of target personnel as nodes and spatiotemporal movement correlation as edges. Then, in the multi-feature fusion module, we utilize the cross-attention mechanism to fuse video and localization features to improve the expression ability of cross-modal data features, and reduces the number of network parameters by sharing parameter strategies. Additionally, we apply the spatiotemporal separation attention mechanism to jointly the model the spatiotemporal correlation between video action features and trajectory of target personnel. At the same time, a dynamic gating fusion strategy is introduced to dynamically adjust weights based on the quality of video and localization data during the training. Finally, in the anomaly detection module, a multi-granularity decoder is used to parallelly generate frame-level video anomaly probabilities and trajectory segment anomaly scores, with the spatiotemporal consistency loss function constraining their alignment, which is to ensure the spatiotemporal matching of behavior and trajectory anomalies. This framework offers an efficient and robust multi-modal analysis approach for security monitoring systems.
No takes yet. Share an insight, caveat, or question.
Liu et al. (2025) studied this question.