A multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features, demonstrating that reliable risk-quantification is an e...
Alperen Kantarcı, Visvanathan Ramesh, Gemma Roig· Proceedings of the 28th Inte...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.