Skip to content
Preprint

MINT: A Universal Zero-Shot Predictor for Transaction Data

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

The Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning, achieves state-of-the-art predictive question-answering performance in both in-distribution and out-of-distribution questions.

Abstract

Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task-specific models as features. However, these Foundation Models are not designed for flexible zero-shot reasoning across novel downstream prediction tasks, limiting their adaptability and utility. Existing LLM-based approaches to zero-shot prediction often fail to fully exploit the predictive signal within transaction data, while relying on costly text serialization or task-specific architectures that scale poorly. To address these limitations, we present the Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning. We find that MINT achieves state-of-the-art predictive question-answering performance in both in-distribution and out-of-distribution questions, while substantially reducing input tokens, latency, and memory consumption compared to text-serialization baselines. Through comprehensive analyses of representations, alignment strategies, training data, and history length, we establish that compact transaction embeddings are a superior approach to transaction representation than text serialization for multimodal reasoning and zero-shot prediction tasks.

View source

Similar papers

Book Open access Aug 2026

Think-like-LSTM: Memory-Augmented Large Language Models via Dynamic Fine-Tuning for Financial Risk Assessment

FraLLM is a novel LLM fine-tuning framework that seamlessly internalizes transaction-oriented knowledge for FRA and introduces the Memory Token Mechanism, which recurrently aggregates historical text prototypes into a compact, continuously updated memory token that allows LLMs to effectively synthesize long-term transa...

Si-Wei Zhang, Yun Xiong, Xi Chen et al. · 1 citation
Open access Sep 2026

Parameter-Efficient Adaptation of Modern Pretrained Encoders for Low-Resource Entity-Level Financial Sentiment Classification

Financial news may express different sentiments towards the entities it mentions, while labelled examples for entity-level classification are often limited. This study compares the adaptation of three pretrained encoders to this task on FinEntity: FinBERT, a model based on Bidirectional Encoder Representations from Tra...

Xian-Hua Peng · 0 citations
#natural language process... Preprint Sep 2026

Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning

Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable answers, while textbooks must first be transformed into synthetic training exam...

Zhirayr Hayrapetyan, Andrei Kalmykov, Denis V. Kokosinskii et al. · 0 citations
Preprint Aug 2026

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

AdaMTP is proposed, an adaptive training paradigm that dynamically aligns the prediction horizon with the intrinsic predictability of the sequence, and consistently outperforms standard MTP in both task performance and inference speedup.

Ziqiang Cui, Han Shi, Bowei He et al. · 0 citations
#natural language process... Preprint Sep 2026

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure...

Cheng-Xian Hu, Zhi-Ming Ma, Ming-Jun Pan et al. · 0 citations
Conference Aug 2026

CGL-FED: A Continual Graph-Based Learning Framework for Fraudulent Email Detection

Email communication remains the primary vector for sophisticated cyber threats, including phishing and spam, resulting in billions of dollars in annual financial losses. While state-of-the-art deep learning (DL) models have demonstrated high precision in static environments, they frequently suffer from performance degr...

Mohammed Abdulwahab, Muneer Almekhlafi, Raeed Al-Sabri · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.