The Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning, achieves state-of-the-art predictive question-answering performance in both in-distribution and out-of-distribution questions.
Abstract
Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task-specific models as features. However, these Foundation Models are not designed for flexible zero-shot reasoning across novel downstream prediction tasks, limiting their adaptability and utility. Existing LLM-based approaches to zero-shot prediction often fail to fully exploit the predictive signal within transaction data, while relying on costly text serialization or task-specific architectures that scale poorly. To address these limitations, we present the Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning. We find that MINT achieves state-of-the-art predictive question-answering performance in both in-distribution and out-of-distribution questions, while substantially reducing input tokens, latency, and memory consumption compared to text-serialization baselines. Through comprehensive analyses of representations, alignment strategies, training data, and history length, we establish that compact transaction embeddings are a superior approach to transaction representation than text serialization for multimodal reasoning and zero-shot prediction tasks.
FraLLM is a novel LLM fine-tuning framework that seamlessly internalizes transaction-oriented knowledge for FRA and introduces the Memory Token Mechanism, which recurrently aggregates historical text prototypes into a compact, continuously updated memory token that allows LLMs to effectively synthesize long-term transa...
Si-Wei Zhang, Yun Xiong, Xi Chen et al.· Proceedings of the 32nd ACM...· 1 citation
Financial news may express different sentiments towards the entities it mentions, while labelled examples for entity-level classification are often limited. This study compares the adaptation of three pretrained encoders to this task on FinEntity: FinBERT, a model based on Bidirectional Encoder Representations from Tra...
Xian-Hua Peng· Applied and Computational En...· 0 citations
Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable answers, while textbooks must first be transformed into synthetic training exam...
Zhirayr Hayrapetyan, Andrei Kalmykov, Denis V. Kokosinskii et al.· 0 citations
AdaMTP is proposed, an adaptive training paradigm that dynamically aligns the prediction horizon with the intrinsic predictability of the sequence, and consistently outperforms standard MTP in both task performance and inference speedup.
Ziqiang Cui, Han Shi, Bowei He et al.· 0 citations
FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure...
Cheng-Xian Hu, Zhi-Ming Ma, Ming-Jun Pan et al.· 0 citations
Email communication remains the primary vector for sophisticated cyber threats, including phishing and spam, resulting in billions of dollars in annual financial losses. While state-of-the-art deep learning (DL) models have demonstrated high precision in static environments, they frequently suffer from performance degr...
Mohammed Abdulwahab, Muneer Almekhlafi, Raeed Al-Sabri· 2026 6th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.