Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

General-Purpose vs. Domain-Specific Large Language Models in Antibiotic Clinical Decision-Making: A Double-Blind Evaluation with a 2X2 Factorial Design

Background: Antimicrobial resistance poses a major threat to global public health. Large language models (LLMs) offer new possibilities for optimizing antibiotic prescribing decisions, but the capabilities of general-purpose versus domain-specific medical LLMs under different prompting strategies remain to be clarified. Methods: This double-blind, randomized-sequence evaluation used a 2X2 factorial design comparing four AI conditions-the domain-specific model MedGo and the general-purpose model DeepSeek V3.5, each under standard direct prompting and chain-of-thought (CoT) prompting-alongside real physician prescriptions across 59 complex inpatient infection cases. Five parallel regimens were generated per case and independently evaluated by three senior clinicians (1-5 comprehensive score and five domain sub-scores). ChatGPT 5.2 was additionally assessed as an automated evaluation tool. Results: Score ranking: real physicians > MedGo-CoT > DeepSeek-CoT > MedGo> DeepSeek (Friedman test, p<0.001). In base mode, MedGo significantly outperformed DeepSeek (Holm-adjusted p=0.040). CoT improved both models (Holm-adjusted p<0.001 for DeepSeek; p=0.024 for MedGo) and reduced score dispersion. MedGo-CoT significantly outperformed DeepSeek-CoT in individualized adjustment (adjusted p<0.001) and dosing precision (adjusted p=0.005). ChatGPT-expert correlation was negligible (overall Kendall {tau}=0.153, p=0.003; subgroup {tau}=0.06-0.20, all p>0.05). Conclusions: Domain-specific medical LLMs enhanced by CoT approach the antibiotic decision-making level of real physicians, with advantages in individualization and dosing precision. However, notable deficiencies persist in antimicrobial stewardship ecological awareness and automated evaluation reliability, underscoring the continued indispensability of senior clinical expertise.

Y. Liu, C. Zhang, F. Wang et al. · 0 citations
#small language model Preprint Aug 2026

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.

Apodex Team B. An, B. Li, B. Wang et al. · 1 citation
#small language model Preprint Aug 2026

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.

Apodex Team B. An, B. Li, B. Wang et al. · 0 citations