The hugang11 system addresses a practical trade-off in creative text generation: models that produce sharper and more stylized jokes often become less stable in output format, and builds a three-stage pipeline that combines chain-of-thought-augmented supervised fine-tuning (CoT-SFT), teacher-constructed direct preference optimization (DPO), and deterministic post-processing.
Error analysis shows that a non-trivial fraction of failures are placeholder strings caused by API errors rather than incorrect generations, and that surface-level mismatches (verbosity, ortho-graphic variation) account for many of the remaining errors.
Our team, VANGUARD, presents IROH (Insightful Ranking of Humor), a three-stage retrieval system for JOKER Task 1 English at CLEF 2026, achieving first place on the leaderboard with 0.6347 MAP. Our pipeline combines hybrid sparse-dense retrieval, cross-encoder reranking, and a LoRA-adapted Large Language Model judge ens...
A. Mocanu, Sebastian Mocanu, Ciprian-Octavian Truică et al.· 0 citations
Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety ali...
A Context-Aware Direct Preference Optimization (CA-DPO) method that significantly outperforms the baseline across all objective and subjective metrics, and establishes an evaluation method featuring a 500-sample test set and an LLM-as-Judge framework to independently assess reasoning and execution fidelity.
Jing-Bin Hu, Lu-Yu Wang, Wen-Jie Tian et al.· 0 citations
We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation)...
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist...
Tsung-Han Wu, Heekyung Lee, An-Ya Ji et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.