What Portrait Benchmarks Actually Measure: 22,000 Human Ratings Explained
Charts from Rapidata's 35-prompt face benchmark, compared with DreamBench++ identity testing and demographic identity-drift research.

The strongest open portrait-specific evidence we found does not answer one universal question. Rapidata measures whether a generated face follows a written human description. DreamBench++ measures whether a personalized generator preserves a reference subject while following a new prompt. A 2026 demographic portrait-editing study measures whether edits change identity unequally across race, gender, and age. Those are three different product decisions, so their results should not be collapsed into one leaderboard.
Executive Summary
- In Rapidata's historical 12-model face-generation benchmark, Seedream 3 had the highest raw win rate at 58.2%, followed by Qwen Image at 57.5% and Recraft v3 at 57.1%. Unequal matchup counts and older model versions mean this is not a current universal model ranking.
- The benchmark is large enough to reveal broad preference patterns: 2,310 paired rows, 35 portrait prompts, 12 models, and about 22,000 human responses. Its vote asks which image follows the description of the human better.
- The 35 prompts explicitly cover children, teens, adults, middle age, and older adults. That breadth helps evaluate described-person fidelity, but it does not test whether a known person's identity survives an edit.
- DreamBench++ and demographic portrait-editing research fill different gaps. They should be reported beside Rapidata, not blended into its scores.
The largest open set measures portrait-description fidelity
Rapidata presented paired outputs and asked raters, “Which Image follows the description of the human better?” The chart below divides each model's published wins by its published total matches. It is a transparent raw preference rate, not an inferred Elo rating and not Rapidata's separate weighted leaderboard score.
Raw preference rate in Rapidata's face benchmark
Published wins divided by published matches; unequal exposure counts are shown for every model.
- Seedream 358.2%2,604 wins / 4,472 matches
- Qwen Image57.5%1,654 / 2,877
- Recraft v357.1%1,157 / 2,028
- MiniMax Image-0153.2%2,726 / 5,125
- Imagen 4 Fast51.9%3,205 / 6,180
- Flux 1.1 Pro48.0%1,272 / 2,650
- Wan 2.2 Image47.8%1,121 / 2,344
- Luma Photon46.7%1,839 / 3,934
- Bria Image 3.246.7%1,507 / 3,226
- Ideogram v3 Balanced46.5%2,868 / 6,170
- Luma Photon Flash42.5%1,499 / 3,529
- Flux Schnell42.4%1,024 / 2,417
The leading trio is tightly grouped: only 1.1 percentage points separate Seedream 3, Qwen Image, and Recraft v3. That supports a historical finding about this benchmark's pairwise votes. It does not establish that any of the three is the best portrait model today, nor does it predict reference-photo identity preservation.
Prompt coverage spans life stages, not identities
We classified all 35 published prompts using only explicit age and life-stage wording. Seven prompts name children or teens, seven name young adults, five name middle-aged or mature adults, and eight name older adults. Eight contain no explicit life-stage term. This is a reproducible coverage summary, not a demographic-quality score.
Explicit life-stage coverage in the 35 prompts
Keyword-based categorization of the complete published prompt list; bars show each category's share of all prompts.
- Older adults8 · 22.9%Prompts explicitly naming older or elderly adults
- No explicit life stage8 · 22.9%Portrait descriptions without an age or life-stage term
- Children and teens7 · 20.0%Prompts explicitly naming children, teenagers, or teen years
- Young adults7 · 20.0%Prompts explicitly naming young adults or ages in the twenties
- Middle-aged or mature5 · 14.3%Prompts explicitly naming middle age or mature adults
The prompt set is useful for testing whether models follow age, appearance, clothing, expression, and setting descriptions across a broad portrait mix. But there is no reference face to preserve. A model can score well by making a plausible described person while changing every identity-level facial feature.
Identity preservation needs a different benchmark
DreamBench++ starts with reference subjects rather than only text. Its repository describes 150 reference images, nine prompts per subject, complete samples from seven personalization methods, human ratings, and automated GPT, DINO, and CLIP measurements. That makes it much closer to the question users ask when they upload their own portrait: does this still look like the same person after the scene changes?
| Evidence set | Main question | Scale and signal | Reuse status |
|---|---|---|---|
| Rapidata Face Generation Benchmark | Which generated image better follows a written human description? | 35 prompts · 12 models · 2,310 pairs · about 22,000 human responses | Dataset: CDLA-Permissive-2.0 |
| DreamBench++ | Does a personalized output preserve the reference subject and follow the new prompt? | 150 references · 9 prompts each · 7 methods · human and automated ratings | Repository: Apache-2.0; source-image licenses require per-image review |
| Demographic portrait-editing study | Does the same edit cause unequal identity drift across race, gender, or age? | Controlled portrait edits · human and vision-language-model evaluation | Method is public; a reusable licensed output corpus was not confirmed |
| AI Photo controlled portrait benchmark | What changes when the exact same authorized portrait and prompt are sent to several models? | 6 unique rows · 27 outputs · byte-identical input within each row | First-party records; no universal winner assigned |
AI Photo's controlled portrait comparison is a smaller but more product-like complement: it retains input, prompt, run, and output hashes so readers can inspect identity, anatomy, and scene retention row by row.
What the results support—and what they do not
- Supported: a historical raw-preference summary for Rapidata's named model versions on 35 text-to-image portrait prompts.
- Supported: a coverage audit showing how often the Rapidata prompts explicitly mention life stages.
- Not supported: calling the highest raw rate the current best portrait generator across products, editing modes, or reference-person workflows.
- Not supported: treating Rapidata's external responses as AI Photo Arena ballots or combining them with AI Photo's 25-ballot publication threshold.
- Not supported: inferring Elo ratings, demographic fairness, identity preservation, or statistical significance from the published win totals alone.
Recommended next steps
- Use Rapidata as the open evidence base for generic portrait-description fidelity and preserve its exact model-version labels.
- Create a separate reference-identity lane using licensed subjects and current production models; score likeness, edit fidelity, anatomy, and output completion independently.
- Add demographic slices only when rater counts, subject consent, category definitions, and uncertainty can be reported responsibly.
- Keep external research, AI Photo's controlled outputs, and community Arena ballots visibly separate in both charts and conclusions.
Further questions
- Would the Rapidata ordering hold if every model faced the same prompt and opponent mix?
- How much do portrait preferences change when raters see a reference identity rather than only a written description?
- Which current models preserve identity without flattening age, skin texture, or other subject-specific features?
Caveats and assumptions
- Raw win rates were calculated from the wins and total-match counts published on Rapidata's dataset page; they are not the dataset's weighted score.
- Life-stage categories were assigned from explicit English prompt wording. They do not infer age from images.
- DreamBench++ repository licensing does not automatically clear every underlying reference image for commercial redistribution.
- The demographic portrait-editing paper is included as methodological evidence because a reusable licensed output dataset was not confirmed.