Skip to content
Preprint

OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

Aug 2026 · 1 citation · 28 references
Computer Science

TL;DR

This work presents OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents, and builds a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility.

Abstract

Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that are described but not enforced. We present OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents. In OpenPM, an agent manages a \$1M long-only book over the S\&P 500 universe using market data at five-minute intervals. Every record visible to the agent must be available at the decision time. Natural-language risk mandates are converted into typed constraints and enforced on the executed portfolio. Each run produces audit artifacts, including a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report. We also build a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility. We isolate constructor behavior by capturing analyst evidence once and replaying it across constructor models. In our short-window case study, stronger constructors show modest and model-dependent gains over equal weighting on the same pool, but analyst quality matters more than constructor choice, and turnover is the main cost driver. All returns are upper bounds on a single frozen window without market impact, not validated alpha.

View source

Similar papers

DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion

This work introduces Search-to-Record, a database-assurance task in which search-enabled large language models reconstruct institution-defined event records from public sources for a known security universe and historical cutoff, and DelistBench, a 1,200-record benchmark for security-level delisting announcements.

Xuan Yao, Shu-Ping Li, Yang Dai et al. · 0 citations
Preprint Aug 2026

Point-in-Time Audit Before Alpha: Public-Archive Availability and a Negative Matched-Budget Study on BTC Perpetual Futures

Public cryptocurrency archives may appear usable when files exist, although factor research requires observations available and executable at each decision time. We audit public Binance BTCUSDT USD-M perpetual-futures data using event, publication, and availability times and separate proposal from deterministic auditin...

Baocheng Zeng, Jin-Hao Yang, Pei-Lin Han et al. · 0 citations
Preprint Aug 2026

ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, are introduced to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence to frame financial compliance evaluation as an audit of rule-grounded actions and eviden...

Yiyan Luo, Yihang Jiang, Qijun Xie et al. · 1 citation
Preprint Aug 2026

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

Yijun Pan, Yu-Kun Lian, Kun-Yu Shi et al. · 2 citations
Review Aug 2026

Context Is Not Authority: Structured Runtime Governance for Financial Market Agents

SAGE-Fin is presented, a finance-specific authority-handoff contract that makes the proposed effect, not merely its text, the object of runtime control, and its results establish executable conformance, not independent safety accuracy.

Rui Tang, Qiang Liu, Yi-Chi Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.