Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge this gap, we propose MIMIC, a framework that leverages executable code as a rigorous mediu...
Jin-Yang Zhang, Wei-Bin Liao, Ke-Qin Bao et al.· 0 citations
Quantile-Guided Density Estimation (QGDE), which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates and suggests that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference...
Qingjie Zhang, Xing-Zhang Ren, Zi-Xuan Chen et al.· 0 citations
The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests.
Zi-Han Qiu, Zekun Wang, Xiao Li et al.· 20 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.