Three image models, one set of briefs
Eight architectural briefs through Gemini 3.1 Flash Lite Image, Gemini 3.1 Flash Image and FLUX.2 pro, three runs each. Speed, brief adherence, aspect ratio and the character of the frames, with no total score.
Open run.jsonThree jobs, not three tiers.
- Gemini 3.1 Flash Lite Image is about three times faster than the other two: a median of 5.0 seconds against 15.2 and 15.8. It returns the same pixel count as Gemini 3.1 Flash Image.
- It also followed the briefs best, which we did not expect: 4.71 out of 5 against 4.62 for Gemini 3.1 Flash Image and 4.08 for FLUX.2 pro, scored blind against a written checklist. Being the quick option does not mean being the careless one.
- Gemini 3.1 Flash Image carries the most detail and the most colour (sharpness 3162, colourfulness 29.0, against 2046 and 26.8 for Flash Lite). It is not a better answer, it is a denser one.
- FLUX.2 pro is the only one that misses the requested shape: fal accepts preset sizes and has no 3:2 preset, so every 3:2 brief came back 4:3. Both Gemini models hit every ratio to within 0.15%.
- FLUX.2 pro is weakest where requirements are countable: floor plans 3.50 and masterplans 3.67, against 4.17 and 4.67 for Flash Lite. It draws handsomely and loses track of three bedrooms or twelve dwellings. It is also the least colourful by a wide margin (20.0 against 26.8 and 29.0).
- This run measures a single generation from a written brief and nothing else. How a model behaves when its result is edited again and again is a separate question, so nothing here is a verdict on a model overall.












Floor plans and masterplans












- Question
We generated the same eight briefs three times through each of three public image models and measured what the choice gets you: how long you wait, the kind of file you get back, how closely the brief is followed, and how the frames differ in character.
There is no winner at the bottom of this page. The three suit different jobs, and the numbers are for deciding which job you are doing.
- Models and access
-
- Gemini 3.1 Flash Lite Image — Google, called through the Vercel AI Gateway
ai-gateway.vercel.sh/v1/chat/completions· google/gemini-3.1-flash-lite-image - Gemini 3.1 Flash Image — Google, called through the Vercel AI Gateway
ai-gateway.vercel.sh/v1/chat/completions· google/gemini-3.1-flash-image - FLUX.2 pro — Black Forest Labs, called through fal
fal.run/fal-ai/flux-2-pro· flux-2-pro, image_size preset
- Gemini 3.1 Flash Lite Image — Google, called through the Vercel AI Gateway
- Scoring
- Every requirement in each brief was written down first (5 to 12 per brief, prohibitions such as “no cars” counting the same as positive ones). Frames were then scored against that list without knowing which model made them. Adherence is the only figure that is not computed by script.
- System prompt
- Every request carried the same fixed system prompt (architectural visualisation, photographic style, follow the user’s text). It is not shipped in the kit; the harness reads one from a file you supply.
- Reproduce
-
- run.json — every measurement on this page
- model-selection-test.mjs — the harness (bring your own gateway and fal keys)
- bench-model-metrics.py — computes every figure above
- raw-run.json — the untouched run output
- Scope
A dated snapshot of three public models, executed on 28 July 2026. Providers update models under the same name; figures are left exactly as the run produced them. Providers update models without renaming them, so the date, endpoint and version of everything tested are stated on the page.
The eight briefs, verbatim
Which AI model is best for architectural visualisation?
There is no single answer, which is why this benchmark publishes no total score. Across eight briefs, Gemini 3.1 Flash Lite Image returned an image in 5.0 seconds, about three times faster than the other two, and followed the briefs most closely. Gemini 3.1 Flash Image returned the most detailed and most colourful frames. FLUX.2 pro was the most restrained in colour, weakest where a brief specifies countable things, and cannot deliver a 3:2 frame through fal's preset sizes. Pick by the job: quick exploration, dense final frames, or a particular character of image.
Which AI model follows an architectural brief most closely?
On this run, the fastest one. Scored blind against a written checklist of every requirement in each brief, Gemini 3.1 Flash Lite Image averaged 4.71 out of 5, Gemini 3.1 Flash Image 4.62 and FLUX.2 pro 4.08. The gap opens where requirements are countable rather than atmospheric: on floor plans and masterplans, where briefs specify three bedrooms or twelve dwellings, FLUX.2 pro averaged 3.50 and 3.67. Open-ended briefs barely separate the three.
How long does an AI model take to generate an architectural image?
The median was 5.0 seconds for Gemini 3.1 Flash Lite Image, 15.2 for Gemini 3.1 Flash Image and 15.8 for FLUX.2 pro, with 90th-percentile times of 6.3, 17.0 and 20.6 seconds. At three images per idea, the fast model returns a set in about fifteen seconds where the others take closer to a minute.
Do AI image models respect the aspect ratio you ask for?
Only if the request carries one and the provider supports it. Both Gemini models hit every requested ratio to within 0.15% once the parameter was passed correctly. FLUX.2 pro through fal accepts preset sizes rather than arbitrary ratios and has no 3:2 preset, so every 3:2 brief came back 4:3. The first two passes of our harness did not send the ratio to the Gemini models at all, which measured the opposite; the run was repeated with that fixed.
Does this benchmark say which model is best overall?
No, and it cannot. It measures one operation: a single image generated from a written brief. It says nothing about how a model behaves when the result is edited repeatedly, which is most of real concept work. A model that scores lower here may be the one you want for iterating; this run does not cover that.
How was this benchmark run?
Eight briefs across four modes (exteriors, interiors, floor plans and masterplans, one open and one heavily constrained per mode) were each generated three times by each of three models, for 72 images. Every request carried the same fixed system prompt and the requested aspect ratio. Latency comes from the run, image statistics from the frames, and brief adherence from blind scoring against a written checklist. The harness, the metrics script and the raw output are published.