Skip to content

ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization

Jul 2026 · arXiv.org · Vol abs/2607.26553 · 0 citations · 56 references
Computer Science

TL;DR

This work proposes ThinkOmni, a reasoning-driven omni-modal large language model that jointly performs explicit forensic reasoning, spoofing detection, and temporal manipulation localization and proposes Forensic-Consistent Multi-task Loss (FCML), which combines weighted cross-entropy with an adaptive localization loss to coordinate reasoning generation, spoofing detection, and temporal localization.

Abstract

Existing audio forgery detection and localization (AFDL) methods often overfit dataset-specific low-level artifacts, limiting their generalization to subtle, localized, and unseen manipulations. Recent audio large language model (ALLM)-based approaches cast AFDL as question answering but still model forensic evidence implicitly, without linking manipulation cues to predictions. To bridge this gap, we propose ThinkOmni, a reasoning-driven omni-modal large language model that jointly performs explicit forensic reasoning, spoofing detection, and temporal manipulation localization. To enable explicit reasoning supervision, we construct Forensic-Aware Chain-of-Thought (FACoT), a 100K-sample dataset with structured forensic evidence and reasoning annotations. Leveraging FACoT, we introduce Forensic-Aware Modality-Incremental Learning (FMIL), which progressively aligns semantic, acoustic, and spectral-visual representations with the LLM backbone to capture complementary forensic cues. We further propose Forensic-Consistent Multi-task Loss (FCML), which combines weighted cross-entropy with an adaptive localization loss to coordinate reasoning generation, spoofing detection, and temporal localization. Extensive experiments show that ThinkOmni achieves strong cross-dataset generalization in both detection and localization. Code, models, data, and inference examples are available at https://beyond0814.github.io/ThinkOmni/.

View source

Similar papers

Preprint Aug 2026

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

This work introduces meta-detection into AI-generated video detection, enabling reliable forgery detection by jointly optimizing predicted labels and supporting evidence within reinforcement learning, and introduces evidence-aware credit assignment, which preserves reliable label supervision while encouraging detectors...

Bo-Wei Liu, Zheng Lu, Yuhan Bian et al. · 1 citation
Jul 2026

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

LaP-Forensics is presented, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence that supports the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.

Can Wang, Yuhao Wang, Yu-She Cao et al. · 1 citation
Preprint Aug 2026

Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

This work proposes a compositional forensic visual prompt learning framework that operates entirely in the visual feature space and employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches.

Fangling Jiang, Qi Li, Bing Liu et al. · 0 citations
Jul 2026

Uncertainty-Aware Deepfake Detection via Multi-View Structural Learning

An uncertainty-aware deepfake detection framework that identifies manipulations through inconsistencies across complementary evidence sources by introducing Inter-Branch Disagreement Calibration (IBDC), a disagreement-aware uncertainty modeling mechanism that links predictive uncertainty to conflicts among evidence str...

Muhammad Umar Farooq, Kutub Uddin, Awais Khan et al. · 1 citation
Preprint Aug 2026

PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs

A Perception-as-Tool paradigm is introduced and instantiated as PATE-Forensics, which architecturally decouples detection and localization from explanation generation while coupling detection and localization as tightly as possible within a forensic perception tool.

Yaqi Li, Jie Peng, Yabin Wang et al. · 0 citations
Preprint Aug 2026

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained text...

Xuechao Zou, Shun Zhang, Kai Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.