Skip to content
Review

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

Poplar is presented, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis and is released together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.

Abstract

Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography-oriented prompts. Render uses a realism-adapted image generator across composition-aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision--language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar-9K: 9,401 curated human-centric image--text pairs retained from 11,765 reviewed candidates (79.9\% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.

View source

Similar papers

Preprint Sep 2026

Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation

Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting...

Shangzhe Di, Zhaokai Wang, Wei-Di Xie · 1 citation
Preprint Sep 2026

GenScale: A Benchmark for Relative Object Scale in Image Generation and Editing

This work introduces GenScale, a benchmark and evaluation protocol for real-world relative object scale in image generation and editing, and introduces Rescale, a model-agnostic post-processing agent for localized scale correction without modifying the source generator.

Ling-Xiao Li, Max Whitton, Ledell Yu Wu et al. · 1 citation · ⚡1
Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations
Preprint Sep 2026

GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models

Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solv...

Chang-Peng Zhao, Yi-Ren Song, Jin-Peng Wang · 0 citations
Review Aug 2026

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text, and Ideogram 3.0 also frequently omits requested elements.

Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad · 1 citation
Preprint Sep 2026

Human-Centric Image Captioning with Subject-Centered Spatial Understanding

While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, they frequently suffer from structural hallucinations in human-centric scenarios. Accurately modeling human subjects is foundational for critical downstream applications, such as accurate avatar/video/image genera...

Bozhou Li, Jia-Hang Zhang, Yue Ding et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.