AI Photo evidence library

AI image model benchmarks and datasets

Every first-party result keeps the outputs and failures visible. Exact-input tests are separated from saved output reviews, and focused views never count twice.

43 unique comparison rows144 model outputs shown6 downloadable datasets

Deduplication rule: the Seedream three-model and Lite/Pro pages are useful focused views of the controlled portrait dataset, but they reuse the same rows and do not inflate these totals.

Controlled benchmarks

Controlled same-input model comparisons

These tests preserve the resolved input fingerprint, prompt hash, run ID, model route, and output lineage. A block remains a block; a focused subset remains part of its parent dataset.

Focused views and output reviews

Focused comparisons and output reviews

The Seedream view below uses the controlled portrait dataset. Broader saved-output reviews remain clearly separated when every historical reference file is not byte-certified.

Portrait output from the five-model comparison
Five models14 prompts

Nano Banana 2 vs Seedream vs GPT Image 2 vs Grok

Fourteen portrait rows across five model columns. Missing and blocked outputs remain in the grids instead of being silently replaced.

Rows
14 prompts
Outputs
56 images shown
Use
Failure-mode inspection
Independent benchmark watch

Independent image-model comparisons

These are original tests from individuals and smaller research or API teams—not broad arena scores. AI Photo links to the source, records the useful finding, and keeps the caveat beside it.

Rapidata

Human portrait-description preferences across 12 image models

12 models · 35 portrait prompts · 2,310 paired rows · about 22,000 human responses

Method
Raters selected which paired output followed the written description of the human better; the dataset publishes weighted responses and per-model win totals.
Useful signal
Seedream 3, Qwen Image, and Recraft v3 had the three highest raw win rates in the published historical comparison.
Limit
Unequal matchup exposure, older named model versions, and no reference identity; this is not a current universal portrait-model ranking.
Human preferencePortraitsOpen dataset
Contra Labs

Brief fidelity across four frontier image models

4 models · 320 prompts across four criteria · 6,400 blind pairwise ratings

Method
Five professional designers scored color, spatial accuracy, typography, and overall brief fidelity without model labels.
Useful signal
Nano Banana 2 led all four fidelity criteria; the study also separated brief-following from hallucination severity.
Limit
Five raters per prompt and English-only briefs; each criterion used a different 80-prompt set.
Blind panelDesign briefsLarge sample
PiAPI

Seedream 5 Pro vs Lite on exact text and multi-object control

2 models · 3 prompt pairs · 6 supplied outputs

Method
Identical prompt text within each pair, with the first supplied Pro and Lite output inspected for visible errors.
Useful signal
Pro made fewer exact-data and directional errors; Lite followed one cabin-composition detail more closely.
Limit
One output per prompt, unequal resolutions and one unequal aspect ratio; no latency or failure records.
First outputTypographySpatial control
PromptFrenzy

The Duck Transform Ladder

6 models · 8 edits · 3 runs each · 144 visible edit grades

Method
One source duck and one exact eight-part edit prompt, with every run published and each edit scored pass, half, or fail.
Useful signal
GPT Image 2 scored 24/24 and was the only entrant to render the requested mirror correctly in all three runs.
Limit
One deliberately synthetic composite-edit task; defaults differ by provider and Meta Muse was run manually.
Objective rubricImage editingEvery run visible
Remix.Camera

Seedream V5 Lite vs V5 Pro: same-input portrait test

2 models · 5 controlled prompts · 10 visible outputs

Method
One frozen Mila input, with the prompt, profile, portrait aspect ratio, and model order held fixed within every row.
Useful signal
The side-by-sides expose pose, lighting, wardrobe, and identity tradeoffs without changing the subject or request; the article leaves the quality call to the reader.
Limit
One trained subject and one retained output per model and prompt; the comparison intentionally does not declare an overall winner.
Exact inputEvery output visible
Vidguru AI Lab

Seedream V5 Pro vs GPT Image 2 in ten first-take tests

2 models · 10 prompts · 20 scored outputs

Method
Identical prompts and references, one generation per model, equal platform credit cost, and a published five-point task rubric.
Useful signal
GPT Image 2 Medium led 47–44, with the clearest separation in multi-reference fidelity, spatial constraints, and commercial layout.
Limit
One output per task and editorial rather than blind scoring; the test equalized Vidguru credit cost instead of maximum model quality.
First outputEvery prompt visibleEqual platform cost
Himanshu Bamoria on X

Five prompts across GPT Image 2, Nano Banana Pro, and Seedream 5 Pro

3 models · 5 prompts · 15 outputs

Method
The same prompt was sent to all three models for typography, photorealism, product, illustration, and prompt-adherence cases.
Useful signal
GPT Image 2 produced the author's favorite single image; Seedream 5 Pro was the most consistent across all five tasks.
Limit
One output per model and prompt, shown in a stitched video; no blinded scoring or retained request metadata.
Individual testSame promptCurrent models
AIReiter

Nano Banana 2 vs Seedream 5 Lite: three first-output tests

2 models · 3 prompts · 6 first outputs

Method
One API platform, identical prompt pairs, each model's first output, and no best-of-N selection.
Useful signal
Nano Banana 2 won the author's three cases on finished layout and real-world accuracy; Seedream kept cleaner glyphs in the simplest design.
Limit
Single runs at model-default resolutions; three hand-picked tasks cannot estimate reliability.
No cherry-pickingCJK textReal-world grounding
u/glusphere on Reddit

Boogu vs Flux Klein vs Qwen Image Edit in a story pipeline

3 systems · 9 edit scenarios · same inputs, prompts, and seed

Method
A working story-video pipeline tested compositing, chained edits, and multi-angle synthesis at native 1280×720 without cherry-picking.
Useful signal
The useful signal is task-specific: models traded wins on screen placement, scene retention, multi-character action, and camera rotation.
Limit
The three systems were not all tested on every scenario, and the post uses qualitative rather than blind scoring.
RedditImage editingProduction workflow
u/Reasonable_Bear_6258 on Reddit

Eleven local image models across seven prompt families

11 models · 7 prompts · 77 full-resolution outputs

Method
One seed per model, a fixed 896×1152 frame, published grids, timing notes, and downloadable workflows for most models.
Useful signal
The post is most useful as a failure atlas for text, realism, posters, and detail—not as a single winner ranking.
Limit
The author was new to the tested models and warns that some default workflows or settings may be suboptimal.
RedditOpen weightsWorkflows published
u/Winter_unmuted on Reddit

Advanced prompt-adherence test for eight local models

8 models · 12 scored prompt tests · 60-point rubric

Method
Prompts were published, settings held constant, and the same seed was used across models within each prompt.
Useful signal
Flux 2 Dev led overall at 51/60, while Qwen 2512 was the more balanced runner-up and anatomy changed the ordering materially.
Limit
One author's 1–5 judgments and local workflows; the author stopped after 12 written tests despite exploring a larger set.
RedditPrompt adherenceLocal models

Reddit AI girlfriend app review survey

2026 review survey

AI Girlfriend App Reviews on Reddit

29,146 collected posts became 35 eligible firsthand reviews from 34 authors after affiliate signals, coordinated promotion, copied templates, product affiliations, and non-firsthand material were removed.

Raw posts
29,146
Eligible reviews
35
Distinct authors
34
Evidence labels

How each evidence type is classified

01

Controlled benchmark

Exact input checksum, prompt, route, settings, run conditions, and every outcome are retained.

02

Blind evaluation

Model identities are hidden while a named rubric is applied; panel size and agreement stay visible.

03

Output review

Saved prompts and outputs are inspectable, but at least one laboratory control is missing.

04

Review survey

Published user records are cleaned and deduplicated. The result describes the retained evidence set.

For writers and researchers

Citation and image reuse guidance

Link to the full article or dataset so readers can inspect the rows. Comparison images reproduced on AI Photo are credited to Remix.Camera; external tests remain on their publishers' pages.

Open the full research and media kit →

Suggested citation

AI Photo Editorial. “Five AI Image Models, One Input: 27-Output Controlled Portrait Test.” July 18, 2026. https://aiphoto.ai/blog/controlled-ai-image-model-comparison

Summary of the controlled five-model portrait comparison

Controlled five-model portrait test

Exact-input portrait rows with prompts, routes, hashes, outputs, and the one blocked or missing outcome retained.

Download image ↓Download data ↓
Seedream 4.5, V5 Lite, and V5 Pro portrait comparison grid

Seedream 4.5 vs V5 Lite vs V5 Pro

A broad saved-output portrait review paired with a separate exact-input controlled subset.

Download image ↓Download data ↓
Mirror-selfie portrait used for AI model comparison coverage

Realistic mirror-selfie benchmark

Same-input mirror-selfie prompts with returned images and safety blocks kept in the row-level record.

Download image ↓Download data ↓
Night portrait representing the adult-fashion moderation comparison

Adult-fashion moderation boundary test

Permitted adult-fashion prompts with accepted outputs and blocks shown without retries or substitutions.

Download image ↓Download data ↓
AI Photo chart of screened Reddit AI companion reviews

Reddit AI companion review survey

A cleaned, deduplicated index of eligible firsthand reviews with commercial-risk exclusions recorded.

Download image ↓Download data ↓