A Review of Reinforcement Learning With Multimodal Large Language Models for Vision-Based Human Activity Recognition
Human activity recognition (HAR) is a well-established research domain in which traditional deep learning approaches often struggle with complex human–object interactions and generalization capabilities. While multimodal large language models (MLLMs) provide strong perceptual encoding for vision-based HAR, they remain...