Skip to content

Do LLM Trading Agents Herd? A Market Microstructure Stress Test

Sep 2026 · IEEE Conference on Computational Intelligence for Financial Engineering & Economics · pp. 61-66 · 0 citations · 14 references

Abstract

Large language models (LLMs) are increasingly used as decision agents in financial workflows, raising questions about aggregate behavior when multiple inference instances receive similar prompts and market signals. We introduce an event-replay stress test that measures prompt- and model-induced decision concentration among LLM trading-agent inference instances. Using 18 market events with real intraday OHLCV data from Databento and timestamp-clean event descriptions audited to remove hindsight language, we evaluate GPT-4o-mini and GPT-4o under three prompt conditions: shared (identical prompts), diversified (persona-based prompts), and risk guardrail (explicit risk warnings). In this experiment, shared prompts produce higher action-concentration indices than diversified prompts, with paired event-level gaps of 0.498 for GPT-4o-mini and 0.634 for GPT-4o (p < 0.05, Wilcoxon signed-rank), computed over the seven LLM inference instances per condition. Risk guardrails increase hold rates by 47.6 and 53.2 percentage points, respectively, raising concentration on HOLD rather than directional BUY or SELL imbalance. We use "herding" in a reduced-form sense for action concentration under common inputs; the design does not identify classical imitation or informational cascades. A supplementary GPT-4o-mini run at temperature 0.2 preserved the same qualitative shared-versus-diversified pattern. On an eight-event raw-source subset for GPT-4o-mini, the shared-versus-diversified gap was 0.652 (p = 0.016), directionally consistent with that model’s main estimate of 0.498. This is a controlled simulation of intended actions, not an estimate of live-market impact.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.