2026· Swiss Text Analytics Conference· pp. 63-74· 1 citation· 39 references
Computer Science
TL;DR
This work compares Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Odds Ratio Preference Optimization (ORPO) using a novel reward modeling approach based on execution and semantic principles, revealing that while standard PPO suffers from reward sparsity and catastrophic collapse on 7B models, monolithic alignment via ORPO scales efficiently to 20B parameter models.
SPOC-SQL is proposed, which decomposes Text-to-SQL into four sequential subtasks following standard SQL execution logic and designs stage-specific optimization strategies for the model to learn key decisions, with the objective of enhancing structured decision-making during query construction.
Yingnan Chen, Chun Ding, Tianshi Xu et al.· 0 citations
AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa et al.· International Journal of Dat...· 1 citation
A novel reward function is introduced, designed to guide LLMs toward a targeted simplification style with Group Relative Policy Optimization (GRPO), that combines the SARI metric with specific penalty components.
Arnau Ayguadé Domingo, Stefan Bott, Horacio Saggion· 0 citations
This work finds that likelihood-trained TPMs can result in failed generations due to overly large corrections to the LM’s logits, and trains TPMs with LM-aligned objectives that better align with the LM token-probability space.
Hanzhang Liu, William Zhao, Zi-Lei Shao et al.· 0 citations
This work studies how to make LLMs natively generate TTS-friendly text, which is frame as a preference alignment problem: instead of relying on downstream rewriting modules, this work directly align LLMs to generate text optimized for spoken delivery.
Thibaut Thonet, Jos Rozen, Laurent Besacier· 0 citations
These results provide exploratory evidence that LLM-assisted rewriting can make some moderate-complexity inputs usable within the evaluated DisCoCat configuration, while highlighting prompt design, filtering, and circuit-aware preprocessing as considerations for more scalable QNLP-based financial sentiment analysis.
Brian Llinás, Nikos Chrisochoides· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.