Multi-Modal Language Models as Text-to-Image Model Evaluators
MT2IE is presented, an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images, showing that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics pre...