Skip to content

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

Jul 2026 · arXiv.org · Vol abs/2607.13881 · 1 citation · 61 references
Computer Science

TL;DR

Two key mechanisms are introduced: Context-aware Multi-round Reasoning, which progressively refines interaction hypotheses to ensure exhaustive and compositional HOI discovery, and Multifaceted Interaction Localization, which enhances grounding precision by generating instance-specific descriptions that integrate semantic, spatial, and appearance cues.

Abstract

Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories. While such paradigms achieve strong performance on closed-set benchmarks, they fundamentally entangle interaction understanding with dataset-specific supervision, limiting their ability to generalize to open-world and compositional scenarios. Recent HOI detectors attempt to leverage MLLMs through prompting strategies to transfer interaction-specific knowledge. However, such prompt-based approaches primarily focus on extracting discriminative representations from pretrained models, while underexploring their inherent multimodal reasoning capabilities. As a result, they struggle to provide informative contextual reasoning for ambiguous and open-world interaction scenarios. In this work, we present AgentHOI, a training-free, agentic framework that transfers the generalist multimodal reasoning capabilities of foundation models to HOI detection in the wild. Instead of learning interaction classifiers, AgentHOI modularly orchestrates complementary vision foundation modules to perform open-ended semantic reasoning and spatial grounding in a coordinated manner. To address the challenges of incomplete interaction discovery and ambiguous localization in complex scenes, we introduce two key mechanisms: (1) Context-aware Multi-round Reasoning, which progressively refines interaction hypotheses to ensure exhaustive and compositional HOI discovery, and (2) Multifaceted Interaction Localization, which enhances grounding precision by generating instance-specific descriptions that integrate semantic, spatial, and appearance cues. Extensive experiments demonstrate that AgentHOI achieves superior performance over state-of-the-art supervised and weakly supervised methods in real-world settings, despite requiring no HOID data for training.

View source

Similar papers

Open access 2026

UniRS-Instruct: A Principle-Guided Unified Instruction-Following Dataset for Remote Sensing Understanding

UniRS-Instruct is presented, a high-quality, diversified, and unified multimodal instruction-following dataset for RSI understanding that unifies diverse tasks, including image captioning, visual question answering, visual grounding, and region-level captioning, into a consistent format.

Lin-Rui Xu, Yuhan Wang, Ling Zhao et al. · 0 citations
Preprint Sep 2026

FineHOI: Part-Aware Dense Representations for Zero-Shot Human-Object Interaction Detection

This work proposes FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features, and introduces an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interactio...

Francesco Tonini, Lorenzo Vaquero, Mohammad Mahdi Derakhshani et al. · 0 citations
#machine learning Preprint Sep 2026

Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

This work proposes Vague2Detect, a hybrid pipeline in which a fine-tuned Sentence-BERT retrieves candidates from a structured household Knowledge Base, and YOLO-World verifies their presence in the image, and improves performance to 61% VPSR with high precision, and up to 85% when augmented with GPT fallback.

Ibrohimjon Muminov, Jihie Kim Dongguk University, Seoul et al. · 0 citations
Preprint Aug 2026

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...

Jongsuk Kim, Qi-Yu Wu, Zhuoyuan Mao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

To disentangle visual evidence at both semantic and spatial levels, ProtoLIP is introduced, a lightweight prototype-mediated evidence layer that organizes reusable visual prototypes into text-derived semantic families and uses coarse-to-fine evidence routing, where semantic families constrain prototype eligibility and...

Yan Zhu, Yong-Bo Chen, Zheng-Ming Ding et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.