Bazaar is introduced, a dynamic sealed-bid benchmark for multi-attribute auction under multi-attribute auction under these conditions, grounded in closed-form customer utilities, enabling exact evaluation.
Abstract
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.
Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers. However, with catalogs spanning millions of active items, manual governance of price relationships is infeasible. Inconsistent pricing across item variants distorts customer value perception and cannibalizes sa...
Ravi Teja Chunduri, Srikaran Reddy Boya, Dr. Deepanshu Mishra et al.· 0 citations
As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment...
Intent-based decentralized exchanges delegate execution to a competitive class of agents -- solvers -- whose behavior is shaped by protocol-designed reward rules. We measure how a change to those rules reshapes who captures value, using a governance-dated natural experiment: CoW Protocol CIP-74 (effective 8 December 20...
The proposed Multi-Agent Regime Intelligence Framework is an interpretable, extensible architecture unifying signal domains that are normally treated separately in Indian derivatives literature, rather than a claim of validated trading performance.
Deepanshu Lamba, Neelam Srivastav· Journal of Frontiers in Mult...· 0 citations
The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks, and contributes a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting.
Pei-Ke Zhu, Si-Di Chang· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.