May 2026· arXiv.org· Vol abs/2605.09625· 0 citations· 91 references
Computer Science
TL;DR
A novel multimodal framework that integrates egocentric vision, pupillometry, eye-gaze tracking, posture detection, heart activity, and the inferencing capabilities of large language models (LLMs) to create a proactive and context-aware ecosystem that dynamically adapts to users' psychophysiological states while analyzing temporal patterns and behavioral tendencies to provide personalized and timely interventions.
Abstract
Information workers'productivity is significantly influenced by their cognitive states and physiological responses. AI assistants such as ChatGPT, Copilot, and others have become integral components of knowledge-intensive workplaces. These AI assistants utilize pre-defined user preferences and chat interaction histories, thus confining themselves to reactive exchanges, lacking sufficient adaptability. Consequently, they fail to cater to individual user preferences and are unable to adapt to their psychophysiological states, diminishing potential productivity gains. To bridge this gap, we introduce AwareLLM, a novel multimodal framework that integrates egocentric vision, pupillometry, eye-gaze tracking, posture detection, heart activity, and the inferencing capabilities of large language models (LLMs) to create a proactive and context-aware ecosystem. AwareLLM dynamically adapts to users'psychophysiological states while analyzing temporal patterns and behavioral tendencies to provide personalized and timely interventions. We evaluated AwareLLM through a user study with 20 participants, comparing it to a standard LLM assistant across multiple tasks. Our results show statistically significant improvements in task performance, along with reductions in cognitive fatigue and mental demand. Participants described AwareLLM's personalized interventions as timely and relevant, helping them boost their confidence and deepen engagement with their work. AwareLLM opens new avenues for Human-AI collaboration where technology adapts to our needs rather than us adhering to technological constraints.
ChatGPT for Intelligent Human–AI Interaction: Opportunities and Limitations provides a comprehensive review of the technological foundations, capabilities, applications, and constraints associated with ChatGPT, emphasizing its role in enabling intelligent human–AI collaboration.
Antoine Morel, Camille Laurent· International Bulletin of Ap...· 0 citations
Aura, a framework that enables LLM systems to dynamically modulate output based on a user's evolving emotions, is introduced and indicates that real-time, context-sensitive interventions can improve learning efficiency and user satisfaction without observable degradation in factual accuracy.
Social Proactive Intelligence (SPI) is an emerging research area, aiming to shift embodied agents from reactive assistance toward proactively understanding human needs and executing socially desirable actions. Prior work has largely centered on the average user. However, human expectations are inherently diverse, and p...
Shu-Fan Zhang, Xin-Yi Che, Kuo-Fei Fang et al.· 0 citations
AffAdapt is presented, a seamless interaction design framework for AI-personas, which coordinates streaming speech recognition, proactive turn-management, persona-grounded response generation, a persistent emotional state, and synchronized embodied output into a single interaction loop.
Nishanth Chidambaram, K. Paliwal, Kayla Hom et al.· 1 citation· ⚡1
Large language models can discuss a family recipe, a quiet street, crowding, and fatigue fluently, but those expressions are not organized by the unfolding conjunction of heat, sound, movement, hesitation, physiological load, and later consequence in one person’s life. We propose relationship-scoped multimodal answerab...
Minoru Kegasawa· Companion Publication of the...· 1 citation
An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive datasets capture remote verbal instruction or undifferentiated co-working. They...
Akhil Ajikumar, Mahya Qorbani, Sakib Reza et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.