Preprint
Aug 2026
Falcon Perception-HD: High Density Perception via Reinforcement Learning
This paper explores post-training reinforcement learning (RL), specifically GRPO, to directly align autoregressive perception models with their evaluation metrics, and designs an RL framework that addresses perception-specific challenges: reward design for set-structured outputs and multi-head sampling control.
Sofian Chaybouti, Yasser Dahou, Ngoc Dung Huynh et al.
· 0 citations