How it worksShowcaseBenchmarksWhy NuitPricing Start free
Benchmarks

What these tools actually do to your image

We run AI image models on real architectural tasks and publish what we measure — including the parts that make our own product look worse. Every run carries the date it was executed, the exact version of everything tested, and the prompts needed to repeat it.

How we run these

Dated and versioned

Models change under their own names. An undated benchmark becomes false without anyone touching it, so every run states when it was executed and against which endpoint and parameters.

Reproducible

Seed images, prompts verbatim, masks, the harness and the measurement script are all published. If you cannot repeat it, it is not a benchmark — it is an advertisement.

Honest about losing

Where another model beats the one we ship, we say so and show the number. That is the only reason to trust the runs where it does not.

Published runs

Seed image for the Three models, one set of briefs benchmark
Run 28 July 2026 3 engines steps

Three models, one set of briefs

Eight architectural briefs through the three models Nuit ships, three runs each. Speed, what each one hands back, how closely each follows the brief, and how the frames differ in character — with no total score, because they are not three tiers.

Light is three times faster than the other two — 5.0 seconds against 15.2 and 15.8 — and it returns the same pixel count as Nuit N. If you are still deciding what to look at, the slower models are buying you very little.

Read the run
Seed image for the Ask for the same thing three times benchmark
Run 28 July 2026 3 engines steps

Ask for the same thing three times

Eight architectural briefs, three models, three runs each. Not one of the 72 images is a repeat of another — floor plans vary a third as much as photorealistic scenes, and a tightly written brief steadies the framing where an open one does not.

No model repeats itself. Composition overlap between two runs of the same brief averages 0.07 out of 1.0 — you do not get the same picture back with small differences, you get a different picture. If a frame is worth keeping, keep it; you cannot re-roll your way back to it.

Read the run
Seed image for the Sequential editing: what eight edits do to one image benchmark
Run 27 July 2026 3 engines 8 steps

Sequential editing: what eight edits do to one image

Eight architectural edits applied in sequence to a single generated exterior, three times — once through each of three editing engines. Every intermediate frame saved, every metric computed by script.

There is no winner. The two approaches fail in opposite directions, and which failure you can live with depends on what you are doing.

Read the run

Planned

Listed before they exist so it is clear what is measured and what is still opinion. Nothing here is claimed as a result yet.

How we chose the three models

The same set of briefs through Light, Nuit N and Nuit F — latency, cost, aspect-ratio adherence and brief adherence. Not a ranking: the three exist for different jobs, and the run is designed to show which job each one is for.

Run not yet executed

Style consistency across modes

Whether an exterior, its floor plan and its interiors read as the same building when generated from one brief.

Not started

Try the workflow these runs were built to test.

Nuit routes a marked-up region to mask-based editing and a written instruction to full-frame editing, because this benchmark is what those two paths do differently. Ten generations free, no card.

Start free