A three-role Monte Carlo Tree Search (MCTS) framework that treats the Lean 4 compiler purely as a reward oracle using compiler output as a scalar signal for UCB-guided tree updates without feeding error content into the generation context is proposed.
AutoScientist-Quant, a self evolving search process that regards quantitative research as one budgeted search problem, is presented, a self evolving search process that regards quantitative research as one budgeted search problem.
Zong-Qian Li, Yaoyiran Li, Yao-Hui Guo et al.· 0 citations
The proposed learning-assisted Tabu Search notably reduces computation time while consistently producing higher-quality solutions than the standard algorithm, highlighting the potential of combining machine learning with metaheuristics by leveraging the implicit knowledge embedded in search trajectories.
Wissem Ahmed Zaid, Alain Hertz, Denny Liu· 0 citations
A novel preference elicitation algorithm for linear utilities that outperforms prior techniques in practice and is applied to heart transplant allocation where a policy must balance competing objectives such as post-transplant outcomes, waitlist mortality, geographic ease, and equity.
Itai Zilberstein, I. Anagnostides, Zachary W. Sollie et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work introduces a conservative, fully auditable spell-correction reliability layer conceived as a safety-oriented preprocessing module rather than a maximal-accuracy corrector: under conditions of uncertainty, the system abstains from editing, in accordance with a medical do-no-harm philosophy.
Moustafa Mohamed Hassan, Sharon Wong, Woh Kai Xuan· International Conference on...· 0 citations
Results provide initial evidence that KV states can serve as transferable computational representations rather than strictly model-local caches, and motivate context mobility as a systems abstraction for reducing redundant prefill across heterogeneous LLM and multi-agent inference workflows.
Yi Li, Dongming Jiang, Yi Zhao et al.· 0 citations
This work calls on the stream-learning community to make bounded resource usage a first-class design objective alongside drift adaptation, and proposes concrete steps toward this goal, including an API through which stream learners can explicitly expose and respect resource budgets.
Sebastian Buschjäger, N. Gunasekara, H. Gomes· 0 citations
This work releases BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set, and releases an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-to-head.
Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al.· 1 citation
DyTrim is proposed, a principle-based dynamic pruning framework that reallocates gradient budget through class-aware pruning on labeled data and confidence-based soft pruning on unlabeled data and provides theoretical guarantees that DyTrim reduces class bias and improves generalization.
Yue Cheng, Jia-Jun Zhang, Xiao-Hui Gao et al.· 1 citation
Learning materials properties from scarce labels and unlabeled crystals is a central challenge for data-driven materials discovery. We present SemiMat, a controlled benchmark for semi-supervised materials property regression, and MatRank, a reliability-weighted objective for continuous pseudo-label uncertainty. SemiMat fixes labeled and unlabeled crystal inputs, graph-backbone interfaces, validation-only checkpoint selection, held-out test reporting, normalized MAE (NMAE), and method-rank summaries across six scarce-label tasks, four graph backbones, and five predefined split runs. MatRank builds pseudo-targets from labeled anchors, weights them by local reliability and weak-prediction agreement, trains weak and strong graph views consistently, and adds ranking signals so that unlabeled crystals shape both values and candidate order. Across the retained 24 backbone-task blocks, one fixed MatRank objective gives the lowest aggregate held-out test NMAE (0.896) and best average method rank (2.208). The component, OOD, and generated-pool diagnostics identify where the gain is reliable and where further screening evaluation remains necessary. Code is available at https://github.com/littlepeachs/SemiMat.
Wen-Tao Li, Yi-Zhe Chen, Jiang-Jie Qiu et al.· 0 citations
CoMPASS is presented, a retrieval-calibrated framework for small-large model collaboration that retains a graph attention network as the predictive anchor, retrieves locally relevant training molecules, provides attention-grounded evidence to an LLM, and converts its proposal into a bounded correction through an agreement-aware gate.
Wen-Tao Li, Jiang-Jie Qiu, Yi-Jun Li et al.· 0 citations
The AUC is shown to be non-collapsible because it decomposes into within- and cross-group AUC terms when subpopulations coexist, such that its overall value may fall outside the range of subgroup specific AUCs.
João Matos, B. van Calster, Richard D. Riley et al.· 0 citations