Preprint
Jul 2026
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
DataGovBench is introduced, a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios that includes two tasks: Table QA that requires solving complex decomposable questions and producing textual answers or visualizations, and Table Insight that evaluates the ability of models to generate expert-level findings through exploratory data analysis.
So Hasegawa, Shailaja Keyur Sampat, Lei Liu et al.
· 0 citations