This work presents OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents, and builds a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility.
Abstract
Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that are described but not enforced. We present OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents. In OpenPM, an agent manages a \$1M long-only book over the S\&P 500 universe using market data at five-minute intervals. Every record visible to the agent must be available at the decision time. Natural-language risk mandates are converted into typed constraints and enforced on the executed portfolio. Each run produces audit artifacts, including a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report. We also build a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility. We isolate constructor behavior by capturing analyst evidence once and replaying it across constructor models. In our short-window case study, stronger constructors show modest and model-dependent gains over equal weighting on the same pool, but analyst quality matters more than constructor choice, and turnover is the main cost driver. All returns are upper bounds on a single frozen window without market impact, not validated alpha.
This work introduces Search-to-Record, a database-assurance task in which search-enabled large language models reconstruct institution-defined event records from public sources for a known security universe and historical cutoff, and DelistBench, a 1,200-record benchmark for security-level delisting announcements.
Xuan Yao, Shu-Ping Li, Yang Dai et al.· 0 citations
Public cryptocurrency archives may appear usable when files exist, although factor research requires observations available and executable at each decision time. We audit public Binance BTCUSDT USD-M perpetual-futures data using event, publication, and availability times and separate proposal from deterministic auditin...
Baocheng Zeng, Jin-Hao Yang, Pei-Lin Han et al.· 0 citations
ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, are introduced to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence to frame financial compliance evaluation as an audit of rule-grounded actions and eviden...
Yiyan Luo, Yihang Jiang, Qijun Xie et al.· 1 citation
This work presents TRACE-RealWorld (TRW), to their knowledge the first commitment-level consistency contract for world models, which makes a world model an auditable predictive interface rather than a self-validating source of truth.
Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.
Yijun Pan, Yu-Kun Lian, Kun-Yu Shi et al.· 2 citations
SAGE-Fin is presented, a finance-specific authority-handoff contract that makes the proposed effect, not merely its text, the object of runtime control, and its results establish executable conformance, not independent safety accuracy.
Rui Tang, Qiang Liu, Yi-Chi Zhang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.