It is shown decoder-only MDMs, despite a larger modeling space, can achieve significant inference speedups and comparable perplexity with techniques like temperature annealing, offering a path to reduced inference compute.
Shuchen Xue, Tianyu Xie, Tianyang Hu et al.· arXiv.org· 17 citations· ⚡1
Overall, same-model draft-and-refine provides evidence that bidirectional refinement is a useful decoding primitive for DLMs, while speculative correction demonstrates a training-free route to fast DLM generation.
Brian K. Chen, Chongwu Wu, Kenji Kawaguchi· 0 citations
This work systematically investigates the impact of the projection unit on LoRP methods, and extends existing LoRP approaches by introducing an additional degree of freedom, projection granularity, beyond the traditional rank hyperparameter, which enables a framework capable of performing fine-grained projections, which is named VLoRP.
Yezhen Wang, Zhou-Hao Yang, Fanyi Pu et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.