Comparisons

What Portrait Benchmarks Actually Measure: 22,000 Human Ratings Explained

Charts from Rapidata's 35-prompt face benchmark, compared with DreamBench++ identity testing and demographic identity-drift research.

Reference portrait used in AI Photo's controlled portrait benchmark
Illustrative comparison input: Remix.Camera

The strongest open portrait-specific evidence we found does not answer one universal question. Rapidata measures whether a generated face follows a written human description. DreamBench++ measures whether a personalized generator preserves a reference subject while following a new prompt. A 2026 demographic portrait-editing study measures whether edits change identity unequally across race, gender, and age. Those are three different product decisions, so their results should not be collapsed into one leaderboard.

Executive Summary

  • In Rapidata's historical 12-model face-generation benchmark, Seedream 3 had the highest raw win rate at 58.2%, followed by Qwen Image at 57.5% and Recraft v3 at 57.1%. Unequal matchup counts and older model versions mean this is not a current universal model ranking.
  • The benchmark is large enough to reveal broad preference patterns: 2,310 paired rows, 35 portrait prompts, 12 models, and about 22,000 human responses. Its vote asks which image follows the description of the human better.
  • The 35 prompts explicitly cover children, teens, adults, middle age, and older adults. That breadth helps evaluate described-person fidelity, but it does not test whether a known person's identity survives an edit.
  • DreamBench++ and demographic portrait-editing research fill different gaps. They should be reported beside Rapidata, not blended into its scores.

The largest open set measures portrait-description fidelity

Rapidata presented paired outputs and asked raters, “Which Image follows the description of the human better?” The chart below divides each model's published wins by its published total matches. It is a transparent raw preference rate, not an inferred Elo rating and not Rapidata's separate weighted leaderboard score.

Raw preference rate in Rapidata's face benchmark

Published wins divided by published matches; unequal exposure counts are shown for every model.

  1. Seedream 358.2%2,604 wins / 4,472 matches
  2. Qwen Image57.5%1,654 / 2,877
  3. Recraft v357.1%1,157 / 2,028
  4. MiniMax Image-0153.2%2,726 / 5,125
  5. Imagen 4 Fast51.9%3,205 / 6,180
  6. Flux 1.1 Pro48.0%1,272 / 2,650
  7. Wan 2.2 Image47.8%1,121 / 2,344
  8. Luma Photon46.7%1,839 / 3,934
  9. Bria Image 3.246.7%1,507 / 3,226
  10. Ideogram v3 Balanced46.5%2,868 / 6,170
  11. Luma Photon Flash42.5%1,499 / 3,529
  12. Flux Schnell42.4%1,024 / 2,417
Rates are descriptive. Models faced different numbers and mixes of matchups, and the set represents the model versions named in the dataset.Rapidata Face Generation Benchmark

The leading trio is tightly grouped: only 1.1 percentage points separate Seedream 3, Qwen Image, and Recraft v3. That supports a historical finding about this benchmark's pairwise votes. It does not establish that any of the three is the best portrait model today, nor does it predict reference-photo identity preservation.

Prompt coverage spans life stages, not identities

We classified all 35 published prompts using only explicit age and life-stage wording. Seven prompts name children or teens, seven name young adults, five name middle-aged or mature adults, and eight name older adults. Eight contain no explicit life-stage term. This is a reproducible coverage summary, not a demographic-quality score.

Explicit life-stage coverage in the 35 prompts

Keyword-based categorization of the complete published prompt list; bars show each category's share of all prompts.

  1. Older adults8 · 22.9%Prompts explicitly naming older or elderly adults
  2. No explicit life stage8 · 22.9%Portrait descriptions without an age or life-stage term
  3. Children and teens7 · 20.0%Prompts explicitly naming children, teenagers, or teen years
  4. Young adults7 · 20.0%Prompts explicitly naming young adults or ages in the twenties
  5. Middle-aged or mature5 · 14.3%Prompts explicitly naming middle age or mature adults
Categories are mutually exclusive and sum to 35. They describe prompt wording, not the people depicted or raters' demographics.Rapidata dataset prompt field

The prompt set is useful for testing whether models follow age, appearance, clothing, expression, and setting descriptions across a broad portrait mix. But there is no reference face to preserve. A model can score well by making a plausible described person while changing every identity-level facial feature.

Identity preservation needs a different benchmark

DreamBench++ starts with reference subjects rather than only text. Its repository describes 150 reference images, nine prompts per subject, complete samples from seven personalization methods, human ratings, and automated GPT, DINO, and CLIP measurements. That makes it much closer to the question users ask when they upload their own portrait: does this still look like the same person after the scene changes?

Evidence setMain questionScale and signalReuse status
Rapidata Face Generation BenchmarkWhich generated image better follows a written human description?35 prompts · 12 models · 2,310 pairs · about 22,000 human responsesDataset: CDLA-Permissive-2.0
DreamBench++Does a personalized output preserve the reference subject and follow the new prompt?150 references · 9 prompts each · 7 methods · human and automated ratingsRepository: Apache-2.0; source-image licenses require per-image review
Demographic portrait-editing studyDoes the same edit cause unequal identity drift across race, gender, or age?Controlled portrait edits · human and vision-language-model evaluationMethod is public; a reusable licensed output corpus was not confirmed
AI Photo controlled portrait benchmarkWhat changes when the exact same authorized portrait and prompt are sent to several models?6 unique rows · 27 outputs · byte-identical input within each rowFirst-party records; no universal winner assigned

AI Photo's controlled portrait comparison is a smaller but more product-like complement: it retains input, prompt, run, and output hashes so readers can inspect identity, anatomy, and scene retention row by row.

What the results support—and what they do not

  • Supported: a historical raw-preference summary for Rapidata's named model versions on 35 text-to-image portrait prompts.
  • Supported: a coverage audit showing how often the Rapidata prompts explicitly mention life stages.
  • Not supported: calling the highest raw rate the current best portrait generator across products, editing modes, or reference-person workflows.
  • Not supported: treating Rapidata's external responses as AI Photo Arena ballots or combining them with AI Photo's 25-ballot publication threshold.
  • Not supported: inferring Elo ratings, demographic fairness, identity preservation, or statistical significance from the published win totals alone.

Recommended next steps

  • Use Rapidata as the open evidence base for generic portrait-description fidelity and preserve its exact model-version labels.
  • Create a separate reference-identity lane using licensed subjects and current production models; score likeness, edit fidelity, anatomy, and output completion independently.
  • Add demographic slices only when rater counts, subject consent, category definitions, and uncertainty can be reported responsibly.
  • Keep external research, AI Photo's controlled outputs, and community Arena ballots visibly separate in both charts and conclusions.

Further questions

  • Would the Rapidata ordering hold if every model faced the same prompt and opponent mix?
  • How much do portrait preferences change when raters see a reference identity rather than only a written description?
  • Which current models preserve identity without flattening age, skin texture, or other subject-specific features?

Caveats and assumptions

  • Raw win rates were calculated from the wins and total-match counts published on Rapidata's dataset page; they are not the dataset's weighted score.
  • Life-stage categories were assigned from explicit English prompt wording. They do not infer age from images.
  • DreamBench++ repository licensing does not automatically clear every underlying reference image for commercial redistribution.
  • The demographic portrait-editing paper is included as methodological evidence because a reusable licensed output dataset was not confirmed.
Your turn

Share your opinion — enter the Arena

Vote blind on real AI image outputs, or continue with a related guide.

Share your opinionEnter the ArenaVote blind on real AI image outputs and see which models readers prefer.Related guideAI Portrait RealismCompare skin texture, anatomy, camera logic, lighting, and the difference between a polished image and a believable photograph.Related guideCharacter ConsistencyEvaluate whether the same authorized person or fictional character remains recognizable across changes in pose, wardrobe, crop, and environment.