Skip to content
Review Open access

A Survey on Action Recognition: Multimodal Approaches, Ethical Considerations, and Feedback Mechanisms

Jul 2026 · Electronics · Vol 15, pp. 3139 · 0 citations · 45 references

TL;DR

This paper provides a comprehensive survey of the current state of action recognition, focusing specifically on three open-world challenges: the integration of multimodalities, the ethical and social implications of these technologies, and the utilization of feedback mechanisms to enhance model performance.

Abstract

Action recognition has emerged as a critical area of research within the realm of computer vision, driven by the increasing demand for intelligent human–machine systems capable of understanding and interpreting human behaviors in the real world. The ability to decipher intricate details of human actions holds immense potential to improve system design, predictive modeling, data-informed decision-making, and real-time operational improvements across a wide variety of domains. Some examples of applications range from surveillance and real-time management of public spaces and infrastructure systems, to development of predictive modeling and robotic systems for individualized healthcare interventions, to implementing effective human–computer interaction in both professional and recreational settings. This paper provides a comprehensive survey of the current state of action recognition, focusing specifically on three open-world challenges: the integration of multimodalities, the ethical and social implications of these technologies, and the utilization of feedback mechanisms to enhance model performance. We delve into the evolution of action recognition, from early feature-based approaches to the deep learning revolution, emphasizing how the incorporation of multiple sensory modalities—such as visual, audio, and depth data as well as other cues—has advanced the field. Furthermore, we examine the ethical challenges associated with deploying these technologies in the public domain, particularly regarding privacy, bias, and societal impact, and discuss the need for responsible development and regulation. The third focus of the paper is the use of top-down and bottom-up feedback mechanisms within deep learning architectures, exploring how these strategies can mimic human cognitive processes to improve accuracy and reliability in action recognition systems. By identifying current gaps and proposing future research directions, this paper aims to inspire continued innovation in this dynamic and impactful field for intelligent systems.

Read PDF

Similar papers

Review

Computer Vision and Image Understanding

This work aims to establish a unified perspective on vision-based mistake analysis in procedural activities, highlighting its potential across diverse domains and aspects and categorizing approaches based on their use of procedural structure, supervision levels and learning strategies.

Konstantinos Bacharidis, Antonis A. Argyros, Hazel Doughty · 0 citations
Aug 2026

Preparing for the Augmented Intelligences: A Roadmap for Physicians in the Era of Artificial Intelligence and Robotics.

The integration of artificial intelligence and robotics into clinical medicine is no longer a question of whether but of how, and physicians currently caring for older adults with complex multimorbidity need practical guidance for the transformation ahead. This commentary offers a framework organized around three funda...

R. Stefanacci · 0 citations
Aug 2026

Introduction to the Special Issue on Large Action Models (LAMs): Theory, Implementation, and Applications

Large Action Models (LAMs) extend the capabilities of AI systems beyond text generation toward perception, reasoning, and action, enabling applications across robotics, autonomous systems, smart manufacturing, healthcare, and the Internet of Things. This Special Issue of ACM Transactions on Multimedia Computing, Commun...

M. Gabbouj, Jin Li, Xin Lin et al. · 0 citations
Preprint Open access Aug 2026

Design of a Human-Assistance Robot System with Contextual Action Recognition

This paper presents a conceptual design for a proactive human assisting robot system capable of recognizing human activities and responding proactively. The system leverages contextual human activity recognition to interpret human actions across diverse contexts, while behavior trees are utilized to define dynamic and...

Amanuel Ergogo, Teresa Zielińska · 0 citations
Open access 2021

Human Values in the Age of Autonomous Technologies

The rapid development of autonomous technologies such as artificial intelligence (AI), machine learning, and robotics has transformed the relationship between humans and machines. While these systems enhance efficiency and productivity, they raise critical concerns about preserving human values in automated decision-ma...

Shalini Gupta · 0 citations
Review Open access 2026

A Review of Reinforcement Learning With Multimodal Large Language Models for Vision-Based Human Activity Recognition

Human activity recognition (HAR) is a well-established research domain in which traditional deep learning approaches often struggle with complex human–object interactions and generalization capabilities. While multimodal large language models (MLLMs) provide strong perceptual encoding for vision-based HAR, they remain...

Wenqi Zheng, Yutaka Arakawa · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.