Skip to content

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Jul 2026

Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech

We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores. On a controlled English benchmark, StreamFraudNet achieves a ROC--AUC of \(0.9953\), significantly outperforming acoustic and mean-pooling baselines while remaining competitive with strong global temporal models. The model produces its first prediction after 10 seconds of audio, updates every 2 seconds, and operates faster than real time on the evaluated server hardware. Ablations identify recurrent temporal context as the principal contributor to performance. These results demonstrate that fraud risk can be scored incrementally from raw speech without transcripts or temporal annotations, while highlighting the need for latency-aware training to improve early prediction.

K. Vo, A. T. D. Dinh, Tien-Ta Tai et al. · 0 citations
Preprint Aug 2026

Causal Episodic Memory for Feedback-Driven Agent Repair

Results clarify when causal cross-query memory improves repair and when broader memory representations remain preferable, and show that negative memory contributes modestly, the value of type conditioning and lexical-dense ranking is dataset dependent, and schema-local experience provides the most consistent benefit.

K. Vo, Tam Minh Chu, A. T. D. Dinh et al. · 0 citations
Preprint Jul 2026

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models

A tree-structured framework is proposed that organizes failures in knowledge-intensive visual question answering into model-specific operational outcomes that support attribution-guided routing to targeted interventions, including image repair, entity support, question rewriting, and factual evidence.

K. Vo, Artem Vazhentsev, Artem Shelmanov et al. · 0 citations
Jul 2026

The JEPA Paradox in Language: The Geometry of Linguistic Alternatives

Text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point, showing that text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point.

A. T. D. Dinh, K. Vo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.