Font Size: a A A

Research On Video Action Recognition Algorithm Based On Multi-scale Spatiotemporal Feature Extraction

Posted on:2023-10-20Degree:MasterType:Thesis
Country:ChinaCandidate:Y M LuoFull Text:PDF
GTID:2568306620985569Subject:Computer technology
Abstract/Summary:
Video data has a higher feature dimension,which not only provides more spatiotemporal information,but also needs more computation.Video action recognition methods based on deep learning have a large number of parameters,and are difficult to extract discriminant features from different timing speed and timing scale,resulting in inadequate recognition performance.Aiming at the issue of insufficient spatiotemporal feature extraction,this thesis introduces a multi-scale spatiotemporal feature extraction and fusion strategy,to obtain motion descriptions under different spatiotemporal receptive fields and enhance the correlation and complementarity between different feature scales.The main contribution of this thesis are as follows:(1)A Channel-Gating and Long-range Temporal Correlation Method is proposed,which aims to solve the problems of low flexibility of temporal feature movement,insufficient extraction ability of long-term motion feature and weak discrimination of temporal and spatial feature in the process of frame sequence reconstruction.In this method,an Adaptive Temporal Shift Module is constructed,which can not only improve the flexibility of feature movement,but also realize the differential extraction of spatiotemporal features from different layers.At the same time,the Temporal Position Embedding Mechanism is used to ensure the temporal stability of the feature during the moving process.In addition,a Long-term Temporal Correlation Module is constructed to explicitly establish the interdependence between the sequences and improve the long-term motion feature extraction capability of the model.In this method,a Temporal Pyramid Module is constructed on the top layer,which enriches the semantic diversity of the spatiotemporal features by integrating the multi-scale temporal features of the hierarchical pyramid structure extracted from different stages of the network.Experimental results on three real scene data sets show that this method can effectively reduce the number of model parameters and improve the performance of action recognition.(2)A Feature Excitation and Adaptive Feature Fusion Method is proposed to solve the problem of feature redundancy and feature semantic mismatch.In this method,a Short-term Motion Difference Modeling Module is constructed to improve the expression ability of temporal features,and a feature selection mechanism is used to enhance the feature activation value of motion-sensitive time series.At the same time,a Spatial Feature Excitation Module is constructed to suppress redundant spatial features between different time series to improve the robustness of the model.In addition,an Adaptive Spatiotemporal Feature Fusion Module is constructed to solve the problem of semantic mismatch between spatial features and long and short motion features in the fusion process.Experimental results show that the performance of this method is better than that of similar comparison models.
Keywords/Search Tags:Video Action Recognition, Convolutional Neural Network, Multi-scale, Feature Excitation, Temporal Position Embedding
Related items