Skip to content
Preprint

Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding

Aug 2026 · 1 citation · 50 references
Computer Science

TL;DR

This work presents the training-free, prompt-only system that won first place in the fine-grained understanding track (MA-Bench) of the MAC~2026 Micro-Action Challenge, where both fine-tuning and ground-truth supervision are disallowed.

Abstract

Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little conscious intent yet that reliably leak emotional and psychological state. Understanding them goes beyond assigning a label: a model must also describe which body parts move and reason, faithfully, about why a clip warrants a particular fine-grained category. We present the training-free, prompt-only system that won first place in the fine-grained understanding track (MA-Bench) of the MAC~2026 Micro-Action Challenge, where both fine-tuning and ground-truth supervision are disallowed. Built entirely upon frozen multimodal large language models (MLLMs), the system dynamically routes each of the eight sub-tasks to the MLLM empirically best suited for that task: a discriminative MLLM for closed-ended recognition tasks and a generative MLLM for open-ended description and reasoning tasks. This architecture achieves a statistically significant performance advantage on open-ended tasks, attaining an average score of 2.68 (on a five-point scale) compared to 1.44 for the second-best approach.

View source

Similar papers

Preprint Aug 2026

Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

A novel test-time evidence calibration framework that improves both reasoning details and prediction reliability by introducing a tree search mechanism to progressively acquire localized, fine-grained visual evidence, coupled with a test-time calibration module to mitigate score biases.

Chengyan Wang, Hanliang Xie, Yueyi Yang et al. · 0 citations
Preprint Jul 2026

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

The results suggest that competitive MLLM can emerge from alignment alone, reducing multimodal extension to a lightweight projector-training problem that generalizes across modalities and adapts rapidly to each new LLM release.

Xuanru Zhou, Yiwen Shao, Jiahong Li et al. · 1 citation
Jul 2026

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

This work introduces Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability, and proposes Mixture-of-Thought-Tokens, a new free-form multimodal grounding method that bridges the perception-reasoning gap.

Tianyi Gao, Han Fang, Tianyi Ding et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models

We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting...

Gautam Rajendrakumar Gare, Si-Ying Li, He-Wei Wang et al. · 0 citations
Preprint Sep 2026

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

AudioICL-Bench is introduced, a diagnostic benchmark whose per-episode rules are resampled so that no correct answer is recoverable from prior knowledge, and its nine tasks are organized along two axes that separate what must be learned from demonstrations from what must be perceived in the signal, enabling failures to...

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.