Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state. Omni-modal large language models excel at multimodal understanding but produce only linguistic outputs, leaving a critical gap in embodied response generation. We iden...
H. Agarwal, Xavier Alameda-Pineda, Olivier Perrotin· 0 citations
Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive pri...
S. Kammoun, Simon Leglaive, Xavier Alameda-Pineda et al.· 1 citation
This work establishes a task-agnostic analysis from three intrinsic per-layer perspectives: compression, geometry, and robustness to perturbations, and introduces the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability.