What Portrait Benchmarks Actually Measure: 22,000 Human Ratings Explained
Charts from Rapidata's 35-prompt face benchmark, compared with DreamBench++ identity testing and demographic identity-drift research.

Rapidata tests how closely generated faces follow a written description. DreamBench++ tests reference-subject preservation and prompt adherence. A 2026 portrait-editing study tests identity drift across race, gender, and age. These benchmarks answer different questions.
Executive Summary
- In Rapidata's historical 12-model face-generation benchmark, Seedream 3 had the highest raw win rate at 58.2%, followed by Qwen Image at 57.5% and Recraft v3 at 57.1%. Unequal matchup counts and older model versions mean this is not a current universal model ranking.
- The benchmark is large enough to reveal broad preference patterns: 2,310 paired rows, 35 portrait prompts, 12 models, and about 22,000 human responses. Its vote asks which image follows the description of the human better.
- The 35 prompts explicitly cover children, teens, adults, middle age, and older adults. That breadth helps evaluate described-person fidelity, but it does not test whether a known person's identity survives an edit.
- DreamBench++ and demographic portrait-editing research fill different gaps. They should be reported beside Rapidata, not blended into its scores.
The largest open set measures portrait-description fidelity
Rapidata presented paired outputs and asked raters, “Which Image follows the description of the human better?” The chart below divides each model's published wins by its published total matches. It is a transparent raw preference rate, not an inferred Elo rating and not Rapidata's separate weighted leaderboard score.
Raw preference rate in Rapidata's face benchmark
Published wins divided by published matches; unequal exposure counts are shown for every model.
- Seedream 358.2%2,604 wins / 4,472 matches
- Qwen Image57.5%1,654 / 2,877
- Recraft v357.1%1,157 / 2,028
- MiniMax Image-0153.2%2,726 / 5,125
- Imagen 4 Fast51.9%3,205 / 6,180
- Flux 1.1 Pro48.0%1,272 / 2,650
- Wan 2.2 Image47.8%1,121 / 2,344
- Luma Photon46.7%1,839 / 3,934
- Bria Image 3.246.7%1,507 / 3,226
- Ideogram v3 Balanced46.5%2,868 / 6,170
- Luma Photon Flash42.5%1,499 / 3,529
- Flux Schnell42.4%1,024 / 2,417
The leading trio is tightly grouped: only 1.1 percentage points separate Seedream 3, Qwen Image, and Recraft v3. That supports a historical finding about this benchmark's pairwise votes. It does not establish that any of the three is the best portrait model today, nor does it predict reference-photo identity preservation.
Prompt coverage spans life stages, not identities
We classified all 35 published prompts using only explicit age and life-stage wording. Seven prompts name children or teens, seven name young adults, five name middle-aged or mature adults, and eight name older adults. Eight contain no explicit life-stage term. This is a reproducible coverage summary, not a demographic-quality score.
Explicit life-stage coverage in the 35 prompts
Keyword-based categorization of the complete published prompt list; bars show each category's share of all prompts.
- Older adults8 · 22.9%Prompts explicitly naming older or elderly adults
- No explicit life stage8 · 22.9%Portrait descriptions without an age or life-stage term
- Children and teens7 · 20.0%Prompts explicitly naming children, teenagers, or teen years
- Young adults7 · 20.0%Prompts explicitly naming young adults or ages in the twenties
- Middle-aged or mature5 · 14.3%Prompts explicitly naming middle age or mature adults
The prompt set is useful for testing whether models follow age, appearance, clothing, expression, and setting descriptions across a broad portrait mix. But there is no reference face to preserve. A model can score well by making a plausible described person while changing every identity-level facial feature.
Identity preservation needs a different benchmark
DreamBench++ starts with reference subjects rather than only text. Its repository describes 150 reference images, nine prompts per subject, complete samples from seven personalization methods, human ratings, and automated GPT, DINO, and CLIP measurements. That makes it much closer to the question users ask when they upload their own portrait: does this still look like the same person after the scene changes?
| Evidence set | Main question | Scale and signal | Reuse status |
|---|---|---|---|
| Rapidata Face Generation Benchmark | Which generated image better follows a written human description? | 35 prompts · 12 models · 2,310 pairs · about 22,000 human responses | Dataset: CDLA-Permissive-2.0 |
| DreamBench++ | Does a personalized output preserve the reference subject and follow the new prompt? | 150 references · 9 prompts each · 7 methods · human and automated ratings | Repository: Apache-2.0; source-image licenses require per-image review |
| Demographic portrait-editing study | Does the same edit cause unequal identity drift across race, gender, or age? | Controlled portrait edits · human and vision-language-model evaluation | Method is public; a reusable licensed output corpus was not confirmed |
| AI Photo controlled portrait benchmark | What changes when the exact same authorized portrait and prompt are sent to several models? | 6 unique rows · 27 outputs · byte-identical input within each row | First-party records; no universal winner assigned |
AI Photo's controlled portrait comparison compares 27 outputs using the same reference image within each row.
Caveats and assumptions
- Raw win rates were calculated from the wins and total-match counts published on Rapidata's dataset page; they are not the dataset's weighted score.
- Life-stage categories were assigned from explicit English prompt wording. They do not infer age from images.
- DreamBench++ repository licensing does not automatically clear every underlying reference image for commercial redistribution.
- The demographic portrait-editing paper is included as methodological evidence because a reusable licensed output dataset was not confirmed.