Skip to content

Author

Chi-Min Chan

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Open access 2026

Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities

The rise of Omni-modality Large Language Models (OLLMs) capable of jointly processing text, audio, and visual inputs marks a major step toward general intelligence. Ensuring their alignment with human preferences requires effective Omni-modality Reward Models (ORMs), which serve as surrogates for human judgment to guide OLLMs behavior. However, ORMs evaluation remains under-developed in the previous literature. Existing benchmarks are largely text-centric or limited to bimodal tasks, restricting comprehensive assessment for ORMs. To bridge this gap, we introduce Omni-RewardBench , the first benchmark for comprehensive evaluation of ORMs across modalities. In short, our contributions are threefold: (1) a hybrid automatic-annotation and human-verification pipeline to construct high-quality evaluation data; (2) extensive experiments on 20+ models, including inherently omni-modal and modality-bridged systems. Our experimental results demonstrate that current OLLMs fall short as reward models, revealing several common failure modes such as perception failure , modality dominance failure , and cross-modal fusion failure . and (3) strong correlations between Omni-RewardBench scores and downstream performance (IID r = 0.94, OOD r = 0.72), validating its reliability as a predictor of real-world capability and alignment quality.

Chi-Min Chan, Yujin Zhou, Pengcheng Wen et al. · 0 citations
Preprint Aug 2026

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

FISA is proposed, a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases that generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion.

Chunyang Jiang, Pingping Zhang, Yuzhi Zhao et al. · 0 citations