Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

CoLSM: Collaborative Large and Small Models for Automatic Software Generation

Large language models (LLMs) support human-in-the-loop code development by rapidly generating high-quality code snippets. However, they still face prominent challenges in fast and efficient deployment on edge environments. Such challenges mainly involve heavy computation costs, poor domain accuracy, unbalanced collaboration efficiency and inconsistent cross-model knowledge. This study proposes CoLSM, a new collaboration mechanism guided by mixture experts for automatic software generation. It establishes a hierarchical and iterative working pipeline. A mixture-expert router assigns tasks dynamically. Large models take charge of system architecture and complex logic design. Domain-adapted small models refine code details, optimize resource usage and ensure security compliance. This mechanism integrates an abstract syntax tree based synchronization module to resolve cross-model conflicts and embeds a quality feedback loop to support adaptive iterative optimization. We evaluate the proposed CoLSM on a self-built multi-scenario software generation dataset. Experimental results demonstrate that CoLSM improves software generation accuracy and functional consistency by 4.3% and 5.7%, respectively. It also reduces inference latency by 20.3% and energy consumption by 14.8%. CoLSM effectively combines the respective advantages of large and small models. It realizes accurate and low-cost automatic software generation and provides reliable technical support for agile automated software development.

Quan Wen, Liu-Shun Zhao, Xiongtao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?

Chao Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.