Skip to content
Open access

TCFNet: an end-to-end framework for multimodal action quality assessment via temporal enhancement and contrastive fusion

Aug 2026 · Multimedia Systems · Vol 32 · 0 citations · 48 references

Abstract

Existing Action Quality Assessment (AQA) methods have limitations, such as over-reliance on unimodal, insufficient long-term temporal modeling, and modality alignment biases in multimodal models. To address these issues, we propose TCFNet, an AQA approach from the perspective of multimodal fusion. Compared to previous methods, TCFNet comprehensively integrates complementary information from multiple modalities, including RGB, optical flow, and audio. It also incorporates specialized modules to capture long-range temporal dependencies and enhance rhythmic consistency. In the unimodal processing stage, we first introduce a Temporal Feature Enhancement Module (TFEM) to capture the sequential dependencies. This is followed by a three-layer pyramid network to extract multi-scale features. Then, the isomorphic multimodal fusion network receives the resulting RGB, optical flow, and audio features as input. We incorporate the cross-trimodal Information Noise Contrastive Estimation loss into the loss function. This promotes feature similarity alignment and alleviates semantic and temporal discrepancies between modalities. Experimental results demonstrate that TCFNet achieves average Spearman’s rank correlation coefficients of 0.862 and 0.840 on the RG and Fis-V datasets, respectively. Compared to the strongest multimodal baseline PAMFN, our method yields significant SRCC improvements of 4.3% and 1.8% on these two datasets.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.