TRACE, a novel framework for automated catalog attribute enrichment using agentic Large Language Models (LLMs), triangulates multimodal evidence across merchant catalogs, syndicated feeds, and identity-matched web search to propose candidate attribute values with supporting evidence, and verifies the proposed value against its supporting evidence.
Abstract
Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often attribute-sparse: the attributes shoppers and downstream systems rely on are either buried in unstructured content such as titles and images or missing from the catalog altogether. Manually enriching e-commerce catalogs is impractical given their scale and rapid growth. This paper introduces TRACE, a novel framework for automated catalog attribute enrichment using agentic Large Language Models (LLMs). A ScoutAgent triangulates multimodal evidence across merchant catalogs, syndicated feeds, and identity-matched web search to propose candidate attribute values with supporting evidence, while a JudgeAgent verifies the proposed value for each attribute value against its supporting evidence and decides whether to publish it or route it to human review. On an offline human evaluation dataset, TRACE's proposed attribute values were 98.2% accurate at 74.7% attribute coverage. Deployed in production on an industry-scale catalog, TRACE increased impression-weighted enrichment coverage across four business verticals by 90.4%. An online experiment subsequently showed that surfacing the enriched attributes on the product detail page increased checkout conversion by 0.48%.
For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highly variable, and their large-scale analysis is currently restricted by...
Mathias Zinnen, Alisha Mund, Sabine Lang et al.· 0 citations
CascadeAgent is introduced, a multi-agent framework that automates prompt adaptation and specialization through error-driven semantic refinement and discusses the evaluation and deployment lessons from this setting, including precision– coverage trade-offs, negative labels for abstention, semantic verification of ambig...
Peng Gao, A. Nikolakopoulos, Zhuo Cheng et al.· 1 citation
FinFIRST is the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics, retaining final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.
Wen-Qing Wang, Hai-Tao Xiang, Xin-Yi Zhao et al.· 0 citations
Customers rely on specific product attributes to compare products and make purchasing decisions, but e-commerce catalogs are messy and unstructured, making it difficult to identify which attributes matter most and extract them at scale. Standard Attribute Value Extraction (AVE) systems treat all attributes equally, pro...
Nikhita Vedula, Dushyanta Dhyani, Bryan Wang et al.· 0 citations
Multi-merchant e-commerce catalogs contain equivalent and related products under different merchant-scoped identifiers, fragmenting behavioral evidence across merchants. Expert-defined taxonomies, meanwhile, are often too coarse for fine-grained discovery. We investigate whether a single hierarchical Semantic ID (\sid{...
Steven Xu, Sanjyot Thete, Saathvik Dirisala et al.· 0 citations
Finding a relevant paper is not the same as producing a verifiable answer from it. LitTraceQA requires canonical paper identifiers, exact evidence at the page or object level, and typed answers that match the evaluator. We call the separation between source access and scorer-visible correctness the grounding contract g...
Sachin Gupta, Divya Godara· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.