Book
Open access
Aug 2026
CEComBench: Benchmarking Large Language Models' performance on Chinese E-commerce tasks
A fundamental gap between generation fluency and reasoning ability is uncovered, a pronounced ''inverse scaling effect'' where larger models can underperform in domain-specific reasoning, and systemic bottlenecks across all SOTA models are identified, exposing fundamental limitations of current architectures.
Guangtao Nie, Huimu Wang, Gewei Lu et al.
· Proceedings of the 32nd ACM... · 0 citations