PingAn-NLP at SemEval-2026 Task 9: Multi-Stage Alignment via GRPO and Tiered Ensemble Voting for Multilingual Polarization Detection
Abstract
This paper presents the PingAn-NLP system developed for SemEval-2026 Task 9, focusing on multilingual online polarization identification across 18 languages. We pro-pose a multi-stage optimization framework combining Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To overcome label imbalance and linguistic nuances, we utilized synthetic reasoning chain augmentation via a high-capacity teacher model (Qwen3-235B) and developed a Smart-Tradeoff reward mechanism to balance precision and recall during reinforcement learning. A language-aware tiered ensemble voting strategy was further implemented to optimize inference performance across diverse linguistic tracks. Our 8B-GRPO-Vote configuration achieved the highest Macro-F1 scores among our experimental variants in 7 out of 18 languages. Officially, our system secured second place in the Bengali, English, Odia, and Turkish tracks.