This work proposes a four-layer framework (Policy, Engineering, Composition, Systemic) grounded in two distinct kinds of evidence, kept explicitly separate, and provides a 90-day implementation sequence spanning trading and payments/customer-facing systems.
Abstract
Agentic AI is gaining acceptance in asset management, but governance has not kept pace: 88\% of surveyed finance professionals report no operational governance framework for agentic AI, and only 24 of 75 large U.S. money managers disclosing AI use in Form ADV filings report a formal governance policy. We argue this gap is architectural: governance built for static validation does not survive continuously retrained agentic policies. We propose a four-layer framework (Policy, Engineering, Composition, Systemic) grounded in two distinct kinds of evidence, kept explicitly separate: two calibrated synthetic illustrations (a regret-covariance drift monitor; a crowding simulation showing joint drawdown risk rising from 39.2\% to 79.3\%), and three real, documented cases (a deployed LLM-embedding trading strategy, a \$45 billion discretionary fund's forced-deleveraging blowup, and a tribunal ruling holding an airline liable for its chatbot). The synthetic examples demonstrate computability from observable data; the cases demonstrate that the failure modes are not hypothetical. We provide a 90-day implementation sequence spanning trading and payments/customer-facing systems.
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.
It is argued that targeted amendments to the PA Master Directions recognising Agentic AI are necessary to keep up with the industry, and a dedicated working group on Agentic AI in financial services under the newly formed AI Governance and Economic Group (AIGEG).
Empirical measures of AI exposure ask language models to score O*NET tasks for technical feasibility. In finance, technically feasible tasks must still pass through review, documentation, supervision, confidentiality controls, and accountable human sign-off before entering production. We measure the gap between feasibility and institutional deployability using 2,199 O*NET tasks across 99 finance-and-insurance occupations. We score each task with eight frontier models and a prompt ladder that moves from bare capability to finance-industry context and named regulatory regimes. The within-model institutional markdown is about one-fifth of the mean feasibility score, and positive for all eight models. The markdown is largest for regulated, client-facing credit and advice roles and smallest for marketing, software, and support roles. Cross-model agreement also declines as finance context is added: models agree more about what AI can do than about what financial institutions can deploy. Mapping exposure to publicly traded firms through pre-ChatGPT staffing shares, we find that the pricing content resides in the institutional layer: firms in the top half of the markdown distribution underperform the bottom half by roughly 25 percentage points in market-adjusted cumulative abnormal returns over the three years after ChatGPT, while sorting on technical exposure alone produces no gap. The differential lies outside the range the same design produces over every pre-ChatGPT window of equal length, though with one event window and few subsector clusters we read it as evidence on where return information resides rather than as a causal estimate. Especially in regulated industries, deployable exposure rather than technical feasibility is the more relevant measure of AI exposure.
Claes Backman, Christos A. Makridis· CESifo working papers· 0 citations
A design-science framework for institutional legacy: the durable capacity of a decision system to continue producing beneficial, lawful, explainable, and adaptable outcomes after its original designers have stepped away is developed.
As decentralized finance (DeFi) ecosystems continue to expand, yield aggregators play an important role in automating capital allocation across lending and liquidity protocols. However, most existing aggregators still rely on static strategies and governance-driven update cycles, limiting their ability to respond to rapidly changing market conditions. This paper proposes a conceptual Agentic AI approach for yield aggregation that introduces adaptive, policy-constrained autonomy into decentralized financial systems. The proposed approach adopts a modular three-layer architecture consisting of a Perception Module for contextual data collection, an Agentic AI Core for reasoning and strategy formulation, and an Action Execution Module for controlled on-chain interaction. The study follows a conceptual exploratory design combining architectural modeling, policy-aware pseudocode, and scenario-based behavioral evaluation using historical DeFi data drawn from Aave V3 and Compound V3 lending markets. Rather than benchmarking yield performance, the evaluation focuses on behavioral properties such as decision timing, responsiveness to market signals, and compliance with governance and risk constraints. The results indicate that the proposed approach reduces response delays associated with governance update cycles while maintaining controlled and selective decision behavior under volatile conditions, without compromising policy compliance or introducing unnecessary transaction overhead. This work contributes a reference architecture and a behavioral evaluation approach for integrating Agentic AI into DeFi yield aggregation, offering practical design insights for adaptive, governance-aligned decision systems, and provides a foundation for future empirical validation using live on-chain deployment and learning-based extensions.
Alvan Nauval, A. Alamsyah· International Conference on...· 0 citations
This conceptual study analyzes the accountability gap that opens when strategic goals are delegated to algorithmic agents and develops the Dynamic Authority Delegation Model (DADM), which distributes responsibility among human strategic intent, algorithmic operational execution, and institutional oversight.
Mustafa Kaya· Kamu Yönetimi ve Teknoloji D...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.