Skip to content

Team hugang11 at SemEval-2026 Task 1: A CoT-SFT, Teacher-Constructed DPO, and Deterministic Post-processing Pipeline for Chinese Humor Generation

· 1 citation · 11 references

TL;DR

The hugang11 system addresses a practical trade-off in creative text generation: models that produce sharper and more stylized jokes often become less stable in output format, and builds a three-stage pipeline that combines chain-of-thought-augmented supervised fine-tuning (CoT-SFT), teacher-constructed direct preference optimization (DPO), and deterministic post-processing.

View source

Similar papers

#natural language process... Preprint Sep 2026

IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English

Our team, VANGUARD, presents IROH (Insightful Ranking of Humor), a three-stage retrieval system for JOKER Task 1 English at CLEF 2026, achieving first place on the leaderboard with 0.6347 MAP. Our pipeline combines hybrid sparse-dense retrieval, cross-encoder reranking, and a LoRA-adapted Large Language Model judge ens...

A. Mocanu, Sebastian Mocanu, Ciprian-Octavian Truică et al. · 0 citations
#natural language process... Preprint Sep 2026

Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning

Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety ali...

Tejasvi C. Addagada · 0 citations
Preprint Sep 2026

Bridging Data, Reasoning, and Alignment: A Unified Framework for Context-Aware Instruction-Following TTS

A Context-Aware Direct Preference Optimization (CA-DPO) method that significantly outperforms the baseline across all objective and subjective metrics, and establishes an evaluation method featuring a 500-sample test set and an LLM-as-Judge framework to independently assess reasoning and execution fidelity.

Jing-Bin Hu, Lu-Yu Wang, Wen-Jie Tian et al. · 0 citations
#natural language process... Preprint Sep 2026

Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

We submit M\'eTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation)...

A. Wasserman, David Beauchemin · 0 citations
#computer vision Preprint Aug 2026

CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist...

Tsung-Han Wu, Heekyung Lee, An-Ya Ji et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.