How it worksShowcaseBenchmarksWhy NuitPricing Start free
Benchmark

Three models, one set of briefs

Eight architectural briefs through the three models Nuit ships, three runs each. Speed, what each one hands back, how closely each follows the brief, and how the frames differ in character — with no total score, because they are not three tiers.

Run executed
28 July 2026
Briefs
8 across 4 modes
Repeats
3 per brief, per model
Images
72
Status
Current
Models tested
ModelProviderEndpoint
Lightgoogle/gemini-3.1-flash-lite-image Google (via Vercel AI Gateway) ai-gateway.vercel.sh/v1/chat/completions
Nuit Ngoogle/gemini-3.1-flash-image Google (via Vercel AI Gateway) ai-gateway.vercel.sh/v1/chat/completions
Nuit Fflux-2-pro, image_size preset Black Forest Labs (via fal) fal.run/fal-ai/flux-2-pro

We generated the same eight briefs three times through each of the three models Nuit ships, and measured what the choice actually gets you: how long you wait, the kind of file you get back, how closely the brief is followed, and how the frames differ in character.

There is no winner at the bottom of this page. The three exist for different jobs, and the numbers below are for deciding which job you are doing.

How long each one takes

Wall-clock time per image, median and 90th percentile, measured over 24 images each. The gap matters most while you are iterating: at three images per idea, the fast model returns a set in about fifteen seconds where the others take closer to a minute.

Model Median p90 Failures
Light 5.0 s 6.3 s 0%
Nuit N 15.2 s 17.0 s 0%
Nuit F 15.8 s 20.6 s 0%

What each one hands back

The three do not return the same kind of file, and that is part of the difference — a smaller lossy frame moves faster through everything downstream.

Model Container Delivered size Aspect accuracy
Light jpeg 4 sizes 1.002
Nuit N png 4 sizes 1.002
Nuit F png 3 sizes 0.972

Aspect accuracy is the delivered ratio divided by the requested one — 1.000 is exact.

How closely each followed the brief

The only figures on this page that are not computed. Every requirement in each brief was written down first — 5 to 12 per brief, with prohibitions like "no cars" counting the same as positive requirements — then the frames were scored against that list while blind to which model produced them, and the mapping revealed afterwards.

ModelBrief adherenceFrames scored
Light 4.71 / 5 24
Nuit N 4.62 / 5 24
Nuit F 4.08 / 5 24

Open briefs barely separate the three — almost everything scores full marks when the brief only asks for a mood. The spread comes from briefs with countable requirements.

What was actually missed

Recorded verbatim per frame in run.json, because a list of misses is more use than a score. The recurring ones:

Character, not quality

These three describe how a frame looks, and none of them says which is better. They exist because "N is drier, F is more contrasty" should be a number, not an adjective.

Detail

Laplacian variance — how much fine structure the frame carries. Higher is not better; a soft dusk scene and a crisp noon one are both correct answers.

ModelDetail
Light2045.6
Nuit N3161.9
Nuit F2421.2

Contrast

Standard deviation of luminance across the frame.

ModelContrast
Light57.5
Nuit N62.5
Nuit F55.6

Colourfulness

Hasler–Süsstrunk metric. Separates "more saturated" from "sharper".

ModelColourfulness
Light26.8
Nuit N29.0
Nuit F20.0

Three jobs, not three tiers.

There is no total on this page and no ranking, because the run does not support one. Light is the quick option for deciding what to look at. Nuit N and Nuit F give a different character of result. The right one depends on what you are doing, not on which scores higher — and this page deliberately measures speed and output, not price.

The same brief, three ways

Run 1 of each model on every brief. All 72 frames are downloadable in the reproduction kit below.

Exterior, open brief exterior · 16:9
Exterior, open brief, generated by Light Light
Exterior, open brief, generated by Nuit N Nuit N
Exterior, open brief, generated by Nuit F Nuit F
Exterior, eight constraints exterior · 16:9
Exterior, eight constraints, generated by Light Light
Exterior, eight constraints, generated by Nuit N Nuit N
Exterior, eight constraints, generated by Nuit F Nuit F
Interior, open brief interiors · 3:2
Interior, open brief, generated by Light Light
Interior, open brief, generated by Nuit N Nuit N
Interior, open brief, generated by Nuit F Nuit F
Interior, material discipline interiors · 3:2
Interior, material discipline, generated by Light Light
Interior, material discipline, generated by Nuit N Nuit N
Interior, material discipline, generated by Nuit F Nuit F
Floor plan, open brief floor-plan · 4:3
Floor plan, open brief, generated by Light Light
Floor plan, open brief, generated by Nuit N Nuit N
Floor plan, open brief, generated by Nuit F Nuit F
Floor plan, programme given floor-plan · 4:3
Floor plan, programme given, generated by Light Light
Floor plan, programme given, generated by Nuit N Nuit N
Floor plan, programme given, generated by Nuit F Nuit F
Masterplan, open brief masterplan · 1:1
Masterplan, open brief, generated by Light Light
Masterplan, open brief, generated by Nuit N Nuit N
Masterplan, open brief, generated by Nuit F Nuit F
Masterplan, site rules masterplan · 1:1
Masterplan, site rules, generated by Light Light
Masterplan, site rules, generated by Nuit N Nuit N
Masterplan, site rules, generated by Nuit F Nuit F

The eight briefs, verbatim

Published in full so the run can be repeated.

ModeBriefFormat
exterior A two-storey family house on a gentle slope, timber and white render, large windows facing the valley, late afternoon light.Exterior, open brief 16:9
exterior A single-storey courtyard house in a dry mediterranean landscape. Flat roof with a deep overhang on the south side. Walls in warm lime render, window frames in dark bronze. One olive tree in the courtyard. Gravel ground, no lawn. Morning light from the left. No cars, no people. Photographed at eye level from the entrance side.Exterior, eight constraints 16:9
interiors A calm living room with a plaster fireplace, oak floor, linen sofa, and tall windows with sheer curtains.Interior, open brief 3:2
interiors A kitchen in a converted stone barn. Exactly three materials: pale oak, honed black stone, and lime plaster. A single long island with no upper cabinets. One pendant light over the island. Original timber roof trusses visible. Overcast daylight from a window on the right. No decoration on the surfaces.Interior, material discipline 3:2
floor-plan A clean architectural floor plan of a three-bedroom single-storey house, black lines on white, room labels, furniture indicated.Floor plan, open brief 4:3
floor-plan An architectural floor plan, black line drawing on white. Programme: entrance hall, open kitchen and dining, separate living room, three bedrooms, one bathroom plus one WC, utility room, and a covered terrace on the south side. Single storey, roughly rectangular footprint. Label every room. Show wall thickness and door swings. No colour, no shading.Floor plan, programme given 4:3
masterplan An aerial masterplan of a small residential cluster of eight houses around a shared green, with parking at the edge.Masterplan, open brief 1:1
masterplan An aerial site plan of a hillside development. Twelve dwellings in four terraced rows following the contours. One access road entering from the north-west, no through traffic. Communal parking in two pockets at the road edge. A pedestrian path linking all rows to a viewpoint at the south end. Existing trees kept along the eastern boundary. Top-down orthographic view.Masterplan, site rules 1:1

What we concluded

  1. Light is three times faster than the other two — 5.0 seconds against 15.2 and 15.8 — and it returns the same pixel count as Nuit N. If you are still deciding what to look at, the slower models are buying you very little.
  2. Light also followed the briefs best, which we did not expect: 4.71 out of 5 against 4.62 for Nuit N and 4.08 for Nuit F, scored blind against a written checklist. Being the quick option does not mean being the careless one.
  3. Nuit N carries the most detail and the most colour of the three (sharpness 3162, colourfulness 29.0 against Light's 2046 and 26.8). That is what the extra ten seconds are for. It is not a better answer, it is a denser one.
  4. Nuit F is the only one that misses the requested shape, and for a specific reason: fal accepts preset sizes, and there is no 3:2 preset — every 3:2 brief came back 4:3. The Gemini models hit every ratio to within 0.15%. This is the exact opposite of what we measured before fixing a bug in our own product, which is why the run was repeated.
  5. This run measures a single generation from a written brief, and nothing else. How a model behaves when you edit the result over and over is a separate question that this run does not answer, so nothing here should be read as a verdict on a model overall.
  6. Nuit F is weakest where requirements are countable — floor plans 3.50 and masterplans 3.67 against 4.17 and 4.67 for Light. It draws handsomely and loses track of "three bedrooms" and "twelve dwellings". It is also the least colourful by a wide margin (20.0 against 26.8 and 29.0); whether that is restraint or flatness depends on the project.

Reproducing this

Downloadable here, not a list of paths in a repository you cannot open. Bring your own gateway and fal keys; the harness reads them from the environment.

Measurements, scripts and briefs are published under CC BY 4.0. Model versions change, so this page carries the date the run was executed.

Questions

Which AI model is best for architectural visualisation?

There is no single answer, which is why this benchmark publishes no total score. Measured across eight briefs: Light returns an image in 5.0 seconds, roughly three times faster than the other two at the same pixel count as Nuit N, and it followed the briefs most closely of the three. Nuit N returns the most detailed and most colourful frames. Nuit F is the most restrained in colour and the weakest where a brief specifies countable things, and it cannot deliver a 3:2 frame because its provider only accepts preset sizes. Pick by the job: quick exploration, dense final frames, or a particular character of image.

Which AI model follows an architectural brief most closely?

On this run, the fastest one. Scored blind against a written checklist of every requirement in each brief, Light averaged 4.71 out of 5, Nuit N 4.62 and Nuit F 4.08. The gap opens where requirements are countable rather than atmospheric: on floor plans and masterplans, where briefs specify three bedrooms or twelve dwellings, Nuit F averaged 3.50 and 3.67. Open-ended briefs barely separate the three at all — almost everything scores full marks when the brief only asks for a mood.

How long does an AI model take to generate an architectural image?

On this run the median was 5.0 seconds for Light, 15.2 for Nuit N and 15.8 for Nuit F, with 90th-percentile times of 6.3, 17.0 and 20.6 seconds. The gap matters most when you are iterating: at three images per idea, the fast model returns a set in about fifteen seconds where the others take closer to a minute.

Do AI image models respect the aspect ratio you ask for?

Only if the request actually carries one, and only if the provider supports it. Both Gemini models hit every requested ratio to within 0.15% once the parameter was passed correctly. Flux accepts preset sizes rather than arbitrary ratios, and has no 3:2 preset, so every 3:2 brief came back as 4:3. We found this out the hard way: an earlier version of this run measured the opposite, because our own code was not sending the aspect ratio to Gemini at all.

Does this benchmark say which model is best overall?

No, and it cannot. It measures one operation: a single image generated from a written brief. It says nothing about how a model behaves when you edit that image repeatedly, which is most of real concept work — and the sequential-editing run published here tested third-party editing engines, not the three models compared on this page. So a model that scores lower here may well be the one you want for iterating; this run simply does not cover that.

How was this benchmark run?

Eight briefs across four modes — exteriors, interiors, floor plans and masterplans, one open and one heavily constrained per mode — were each generated three times by each of three models, for 72 images. The harness calls the same endpoints with the same request bodies the product uses, including the production system prompt, which is read from the source rather than copied so it cannot drift. Latency comes from the run, image statistics from the frames, and brief adherence from blind scoring against a written checklist. The harness, the metrics script and the raw output are all published.

Related