Skip to content

Author

Wei-Jung Huang

We have 4 of 7 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

What Does an LLM-Agent Leaderboard Rank Actually Compare?

An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.

Wei-Jung Huang · 0 citations
Review Jul 2026

LLM-Driven CI-CD Workflow Intelligence for Cyber Systems Engineering

An LLM-based CI/CD analysis pipeline that combines repository enrichment, anti-pattern detection, stage mining, and recommendation generation over a large GitHub corpus is presented, arguing for CI/CD observability that combines diagnosis, context, and human review.

Bo-Nan Shen, Jiazhou Gao, Tao Ning et al. · 0 citations
Preprint Aug 2026

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision

ParEvalLayer, a decision layer that reads paired outcomes for two agent systems and a comparison policy chosen in advance, is introduced, a decision layer that reads paired outcomes for two agent systems and a comparison policy chosen in advance.

Wei-Jung Huang, Bo-Nan Shen · 0 citations
Jul 2026

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

This work asks whether LLM-generated skills offer a useful low-curation alternative: do they improve performance over the task prompt alone, and whether any part of the skill is useful by ablating different skill components, and finds no reliable improvement from full generated skills over No-Skill prompting.

Wei-Jung Huang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.