Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models
With the rapid development of multimodal large language models (MLLMs) and the cost of large amounts of data and resources, researchers are more inclined to use public datasets and fine-tune open-source MLLMs to achieve excellent performance in audiovisual related tasks. While this trend accelerates progress, it also i...