Skip to content

Category

machine learning

3,367 papers

#artificial intelligence Preprint Open access Aug 2026

SkillNet: Create, Evaluate, and Connect AI Skills

Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. Without a unified mechanism for skill consolidation, agents frequently ``reinvent the wheel'', rediscovering solutions in isolated contexts without leveraging prior strategies. To address this challenge, we introduce SkillNet, an open infrastructure for creating, evaluating, and organizing AI skills at scale. SkillNet structures skills within a unified ontology that supports creating skills from heterogeneous sources, establishing rich relational connections, and performing multi-dimensional evaluation across Safety, Completeness, Executability, Maintainability, and Cost-awareness. Our infrastructure integrates a repository of over 600,000 skills, an interactive platform, and a versatile Python toolkit. Experiments on ALFWorld, WebShop, and ScienceWorld show 40% higher average rewards and 30% fewer execution steps across multiple backbone models. Furthermore, SkillNet-Gym benchmarks skill retrieval, utilization, and composition, while SkillNet-Fabric enables task-specific skill routing through lightweight Wikis. By formalizing skills as evolving, composable assets, SkillNet provides a robust foundation for agents to move from transient experience to durable mastery.

Yuan Liang, Ruobin Zhong, Haoming Xu et al. · 0 citations
#artificial intelligence Preprint Open access Aug 2026

Conformal Policy Control

An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause harm and must be taken offline, curtailing any future interaction. Imitating old behavior is safe, but excessive conservatism discourages exploration. How much behavior change is too much? We show how to use any safe reference policy as a probabilistic regulator for any optimized but untested policy. Conformal calibration on data from the safe policy determines how aggressively the new policy can act, while provably enforcing the user's declared risk tolerance. Unlike conservative optimization methods, we do not assume the user has identified the correct model class nor tuned any hyperparameters. Unlike previous conformal methods, our theory provides finite-sample guarantees even for non-monotonic bounded loss functions, and it introduces a new policy control setting. Our experiments on applications ranging from natural language question answering to biomolecular engineering show that safe exploration is not only possible from the first moment of deployment, but can also improve performance.

Drew Prinster, Clara Fannjiang, Ji Won Park et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

Group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance, and demonstrates that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance.

Zhu Zhang, Ji-Xun Wang, Xiaoan Xu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

PGFS++ is introduced, a synthesis-aware reinforcement learning framework for input-specific molecular improvement that improves target properties while preserving high output diversity, and experiments show that PGFS++ improves target properties while preserving high output diversity.

Boqiao Zhang, Godbless James, S. Gottipati et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

The Masked Diffusion Time-series Imputation Model (MDTIM) is proposed, which leverages the training paradigm of masked diffusion model for imputation tasks, and introduces Stochastic Discretization, which maps continuous values to ordinal-aware tokens while preserving continuous dynamics.

Dongbin Kim, Seungyun Lee, Geonwoo Shin et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh, systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student.

Huan Gao, Haohan Chi, Yong Yan et al. · 0 citations
#artificial intelligence Preprint Open access Aug 2026

Bernstein-Vazirani Networks: Quantum Machine Learning by Interference

We introduce Bernstein-Vazirani Networks (BVNs), a non-variational quantum machine learning framework that leverages quantum interference for supervised learning, demonstrated on vision and representation learning tasks. In their standard form, BVNs follow the principle of quantum Fourier sampling: labelled data are placed in superposition and interfered in the Fourier basis to extract globally informative features. We then define generalised BVNs that enable interference in problem-adapted bases, yielding more expressive models under the same measurement budget as in the standard setting. BVNs achieve universal function approximation through (over)complete interference bases, while training of BVNs is gradient-free. Experiments on synthetic and real-world classification tasks, as well as implicit image representation, show strong generalisation capabilities and competitive performance with classical and quantum baselines.

Natacha Kuete Meli, Tolga Birdal, Prayag Tiwari et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.

B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Graphical Design of Interpretable Architectures

A graphical notation for designing interpretable AI architectures, adapted from Penrose tensor notation is introduced, which gives a global view of an architecture and maps one to one onto PyTorch einsum code.

Pietro Barbiero · 0 citations
#artificial intelligence Preprint Aug 2026

MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

The proposed Module Level Reward Evolution Framework integrates three mechanisms: reflection-based refinement, hybrid credit assignment, and a merge strategy with rollback, which together improve the effectiveness and robustness of reward optimization.

Chenglin Liu, Xun Wang, Ruishuo Chen et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.