Skip to content
#small language model Review Open access

Grounding language models with deterministic verifiers for fraud detection

Oct 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 31 references
Imbalanced Data Classification Techniques

TL;DR

VR-FraudNet, a five-stage framework combining a time-conditioned spectral graph encoder, LightGBM triage and threshold routing, a schema-constrained rationale language model, a fixed deterministic verifier, and an isotonic probability mixer with split-conformal calibration, does not establish universal adversarial safety, temporal robustness, regulatory sufficiency, or deployment readiness.

Abstract

Financial fraud detection is complicated by severe class imbalance, temporal change, adversarial manipulation, and the need for evidence-grounded explanations. Although large language models can generate readable rationales, their claims may be unsupported when they are not independently checked against structured evidence. We propose VR-FraudNet, a five-stage framework combining a time-conditioned spectral graph encoder, LightGBM triage and threshold routing, a schema-constrained rationale language model, a fixed deterministic verifier, and an isotonic probability mixer with split-conformal calibration. The language model proposes machine-checkable claims for review-band transactions, while the verifier evaluates each claim against the available transaction, temporal, and graph evidence. Verifier-rejected rationales are excluded from automated use, and the corresponding transactions are assigned to human-review escalation. D1 BAF, D2 AMLworld HI-Small, and D3 IEEE-CIS are used as the primary predictive benchmarks, whereas D4 Elliptic++ and D5 DGraph-Fin support complementary temporal-graph, strict-inductive, temporal-robustness, and message-passing stress tests. VR-FraudNet improves AUPRC over TabTransformer from 0.5012 to 0.5247 on D1, from 0.5137 to 0.5824 on D2, and from 0.7234 to 0.7456 on D3. The predeclared 5-percentage-point temporal-degradation target is satisfied in the evaluated D1 month-drift, D3 late-fold, and D4 pre-shock comparisons. However, D4 post-shock AUPRC decreases from 0.6823 to 0.6087, an absolute reduction of 7.36 percentage points that exceeds the target. Under the four-family structured predictive edit set, mean evasion rates are 0.0667, 0.1016, and 0.1672 for budgets 1, 2, and 4, respectively. Under the stated operational cost assumptions, VR-FraudNet produces lower total loss than TabTransformer, with the largest difference on D2. Measured p99 latency is 10.23 ms on the common path and 175.23 ms on the escalation path under the reported hardware configuration. Empirical split-conformal miscoverage remains below the stated bounds in the reported calibration settings, while the formal marginal-coverage statement applies only under the stated exchangeability assumptions. These findings support verifier-grounded fraud detection under the evaluated public datasets and protocols but do not establish universal adversarial safety, temporal robustness, regulatory sufficiency, or deployment readiness.

Read PDF

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.