Skip to content

Author

Hung-Nghiep Tran

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Evaluating AutoEDA Without Human Raters: A Diagnostic Framework and Case Study of Structured Question Generation

Human evaluation remains the dominant way to assess automated exploratory data analysis (AutoEDA) systems, but it is expensive, subjective, and hard to reproduce. We introduce a reproducible automatic diagnostic framework that surfaces structural and statistical failures in AutoEDA—failure modes often under-emphasized when reviewers prioritize fluent prose over breakdown validity and statistical form. The framework uses 9 metrics covering correctness, structural validity, statistical quality, subspace exploration, and question-insight alignment. We apply it to three representative pipelines: a QUIS-style structured question-guided system, ONLYSTATS (a statistics-based ablation), and a free-form agentic LLM. Across Adidas US Sales, IBM Employee Attrition, and Online Sales, structured intermediate representations show fewer column-type errors and broader subspace coverage in our runs. The QUIS-style pipeline reaches 94.0% average structural validity versus 40.0% for the agentic LLM, and discovers more subspace insights on average (84.4% vs. 37.4%). It also attains the highest Subspace Score (x=1.067), while the agentic LLM leads on Uplift Win Rate (66.7%, subspace beats global in 2 of 3 datasets). The QUIS-style pipeline discovers significant paradoxes in 2 of 3 datasets. Overall, schema-aware question guidance tracks higher structural validity and subgroup coverage than free-form prompting in our runs. More broadly, the paper contributes a reusable diagnostic framework for reproducible comparison of AutoEDA systems. Code and data are available at https://github.com/sunnydovision/EDA_Evaluation.

B. N. Do, Phu-Vinh Nguyen, Hung-Nghiep Tran · 0 citations
Conference Aug 2026

Operationalization Matters: When Graph-Aware Learning Adds Value for IC-Based Influence Approximation

Approximating Monte Carlo Independent Cascade (MC-IC) influence rankings with a learned surrogate is only meaningful if the target itself is clearly defined. We study this dependency on the Twitch Gamers graph under two IC operationalizations, keeping the graph, labeled-node sample, split, Monte Carlo budget, and evaluation protocol fixed while using regime-aligned feature policies. Under the structural weighted-cascade regime, binary top-k labels are too unstable for the primary target; in this setting, the best raw-attribute GNN, GCN, remains below degree centrality (ρ = 0.808 vs. 0.826). Under the source-community operationalization, degree centrality becomes uninformative (ρ = −0.006), and raw-attribute GraphSAGE outperforms the pre-specified flat linear-regression (LR) baseline (ρ = 0.915 vs. 0.884). Trained surrogates provide sub-second full-graph inference, whereas MC label generation takes hundreds to about two thousand seconds on the frozen labeled subset. The central finding is conditional: on this graph, neighborhood-aware surrogates add value beyond strong structural, flat, and shallow-embedding alternatives when the target construction moves influence signal away from local degree structure. A supplementary 1-hop neighbor-aggregation diagnostic reaches a Spearman correlation comparable to that of standard Graph-SAGE in the source-community regime, suggesting that part of the observed gain may be attributable to local neighborhood aggregation rather than to learned message passing specifically. The source code and dataset are available at https://github.com/qvinhx89/operationalization-matters-influence.

Dinh-Duy Tran, Q. Pham, Q. Tran et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.