Ask for the same thing three times
Eight architectural briefs, three public models, three runs each. Not one of the 72 images repeats another: floor plans vary a third less than photorealistic scenes, and a tightly written brief steadies the framing.
Open run.jsonNothing here is a repeat.
- No model repeats itself. Composition overlap between two runs of the same brief averages 0.07 out of 1.0: you do not get the same picture back with small differences, you get a different picture. If a frame is worth keeping, keep it; you cannot re-roll your way back to it.
- The mode matters more than the model. Floor plans vary a third less than photorealistic work (34.77 against 52.73 for exteriors, 55.62 for interiors and 52.84 for masterplans). Line drawings converge on a similar answer; a furnished, weathered scene has far more room to differ.
- Writing a tighter brief is the one lever that works. Briefs carrying six to eight explicit requirements came back at 46.65 against 51.33 for open one-liners, with composition agreement 0.078 against 0.066.
- The gap between models is smaller than the gap between a floor plan and an interior (45.8 to 50.8 against 34.8 to 55.6). Pick a model for its character; none of them offers repeatability.






Gemini 3.1 Flash Lite Image
























Gemini 3.1 Flash Image
























FLUX.2 pro
























- Question
We generated the same eight briefs three times each, through three public models, and measured how much the repeats differ from one another. The question is narrow and practical: if you generate something and it is nearly right, can you ask again and expect a near-miss variation, or something else entirely?
- Models and access
-
- Gemini 3.1 Flash Lite Image — Google, called through the Vercel AI Gateway
ai-gateway.vercel.sh/v1/chat/completions· google/gemini-3.1-flash-lite-image - Gemini 3.1 Flash Image — Google, called through the Vercel AI Gateway
ai-gateway.vercel.sh/v1/chat/completions· google/gemini-3.1-flash-image - FLUX.2 pro — Black Forest Labs, called through fal
fal.run/fal-ai/flux-2-pro· flux-2-pro, image_size preset
- Gemini 3.1 Flash Lite Image — Google, called through the Vercel AI Gateway
- Measures
- For every brief-and-model triple all three pairs of runs were compared two ways: the mean absolute pixel difference (variance) and the overlap of their edge maps (composition). Both are computed by script; nothing on this page is judged by eye.
- System prompt
- Every request carried the same fixed system prompt (architectural visualisation, photographic style, follow the user’s text). It is not shipped in the kit; the harness reads one from a file you supply.
- Reproduce
-
- run.json — every measurement on this page
- model-selection-test.mjs — the harness that produced the frames (bring your own gateway and fal keys)
- bench-repeatability.py — computes variance and composition
- raw-run.json — the untouched run output
- Scope
A dated snapshot of three public models, executed on 28 July 2026. Providers update models under the same name; figures are left exactly as the run produced them. Providers update models without renaming them, so the date, endpoint and version of everything tested are stated on the page.
The eight briefs, verbatim
Can I regenerate an AI image and get the same result again?
Not from the same prompt with these models. Across 24 brief-and-model combinations, composition overlap between repeat runs was 0.07 out of 1.0: effectively a different photograph each time, not a variation of the first. A generated image you like is a one-off; save it and branch from it rather than trying to reproduce it from the prompt.
Do more detailed prompts make AI image output more consistent?
Measurably, yes. It is the one thing that reliably narrows the spread. Briefs carrying six to eight explicit requirements came back at 46.65 variance against 51.33 for open one-line briefs, with composition agreement 0.078 against 0.066. They still do not make two runs the same picture.
Which AI images are most reproducible: plans or renderings?
Plans, by a wide margin. Floor-plan briefs varied at 34.77 against 52.73 for exteriors, 55.62 for interiors and 52.84 for masterplans. A black-line plan has a narrow space of plausible answers, so repeat runs converge; a photorealistic scene has to re-decide materials, vegetation and weather each time.
How was repeatability measured?
Eight briefs across four modes were each generated three times by each of three models, 72 images. For every brief-and-model triple we compared all three pairs two ways: the mean absolute pixel difference, and the intersection-over-union of their edge maps, which measures whether the composition is the same regardless of colour. Frames were resampled to a common long edge first so resolution is not a factor. The script and the raw run are published.