A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the right trade-off changes with the workload. In the context of information retrieval (IR), a transformer-based model can be made smaller in three ways---using fewer layers,...
Yu Wang, Shengyao Zhuang, Xue-Guang Ma et al.· 1 citation
Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inferen...
Roy Xie, Dan Friedman, Feng Nan et al.· 0 citations
CLEAR first employs a reflection agent to perform contrastive analysis over past execution trajectories and summarize useful context for each observed task to train a context augmentation model (CAM), which further optimize CAM using reinforcement learning, where the reward signal is obtained by running the task execut...
Lin-Bo Liu, Guan Wu, Han Ding et al.· arXiv.org· 0 citations
A novel simulation approach that combines categorical judgments with evaluator-specific auxiliary data--retrospective reasoning traces and interface telemetry--to enable LLM-based simulation of individual evaluators via in-context learning is proposed.
Zeyu He, Xuan Qi, Subramanian Chidambaram et al.· SIGDIAL Conferences· 0 citations
Tevatron 3.0 is presented, which integrates a Megatron-Core training backend into Tevatron while preserving its data pipeline, evaluation workflow, and Hugging Face-compatible checkpoints, and supports both LoRA and full-parameter fine-tuning.
Zhichao Xu, Xueguang Ma, Shengyao Zhuang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.