Results show that long-horizon reflective data is an effective route toward self-improving agents, and synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration.
Hong-Jin Qian, Chao-Fan Li, Kun Luo et al.· 0 citations
Dense retrieval has become a cornerstone of modern local-lifestyle e-commerce search by encoding queries and items into semantic embedding spaces. While recent advancements have transitioned from BERT-based embedding models to Large Language Models (LLMs), most approaches still treat LLMs as static text encoders, negle...
Ang-Qing Jiang, Gao-Ming Zhang, Jian-Chun Song et al.· 2 citations
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge,...
Jianlyu Chen, Yuyang Hu, Hong-Jin Qian et al.· 1 citation
This work proposes Dynamic Index-based RECommendation with Transport-Optimized Retrieval with Transport-Optimized Retrieval (DIRECTOR), a transport-guided parallel reranking framework that consistently outperforms strong reranking baselines, achieving significant improvement in large-scale industrial recommendation sce...
A Hierarchical Semantic Alignment module to align query's latent space with item's quantization path and synchronize multi-granular semantics, and a personalized GR framework that models user behavior by synergizing discrete SIDs for structural guidance and continuous representations for fine-grained semantic refinemen...
Gao-Ming Zhang, Ang-Qing Jiang, Jian-Chun Song et al.· 0 citations
A novel method, namely AnDPro, is proposed, which introduces a projection-based scoring function to more accurately measure token importance and guide more accurate token selection in key-Value cache eviction.
Zijie Geng, Jie Wang, Ziqi Liu et al.· Neural Information Processin...· 6 citations
It is found that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful.
Zhaoyi Li, Deyang Kong, Yuan Wei et al.· 1 citation
The exact selection time for an isolated cycle of NOTEARS and DAGMA is derived, and a truth-free separation statistic predicts selection time on 320 official NOTEARS/DAGMA trajectories.
Rui Wu, Zongyuan Chen, Hong Xie et al.· 0 citations
The resulting lesson is task-specific: a first stage for generated controls should be judged by control fidelity, downstream relevance, and graph compatibility together.
Rui Wu, Zongyuan Chen, Hong Xie et al.· 0 citations
The results show that simple parameter averaging, when paired with lightweight dimensional adaptation and carefully controlled ratios, is a surprisingly strong baseline for heterogeneous LLM merging, suggesting that the limits of direct weighted fusion may also bound what more complex heterogeneous merging methods can...
Jiahe Fan, Yinghao Hou, Sixiang Chen et al.· 0 citations
SWIM (Step-Wise Integrated Measure), a list-level evaluator that models user behaviors as a finite-horizon prefix session-level survival process, and efficiently estimates continuation probabilities and utilities in parallel, satisfying strict industrial latency constraints.
Yuan Pu, Cheng-Hao Zhang, Chao Feng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.