# AI Will Draw Your Building. It Will Not Count Your Bedrooms.

**We generated eight architectural briefs three times each, through three different
models, and scored all 72 images blind against checklists written before anyone
looked at a picture. The atmosphere was almost always right. The counting was
almost always the problem.**

This is not the failure mode most people expect. The worry about AI in architecture
is usually that it produces generic, plausible mush — pretty renders with no
substance. What we measured was the opposite: the images are specific, materially
convincing and often beautiful. They just cannot be trusted to contain the number
of bedrooms you asked for.

That distinction matters, because the two failures need completely different
responses.

## How the scoring worked

Eight briefs, two per mode — exteriors, interiors, floor plans and masterplans.
One open brief per mode and one loaded with six to eight explicit requirements, so
the run could separate *can it draw a house* from *can it follow instructions*.

Every requirement in every brief was written down as a checklist **before any
frame was viewed**. Prohibitions counted the same as positive requirements: "no
cars, no people" is a requirement, and ignoring it is a miss.

Then the frames were shuffled into contact sheets, nine at a time, and scored
against that checklist while blind to which model produced them. The mapping was
revealed only after every score was written down. Scores run 0 to 5, where 5 means
every requirement met including the prohibitions.

Blinding matters here more than it might seem. We had expectations going in — a
faster, cheaper model *ought* to be the sloppier one — and those expectations turned
out to be wrong. Knowing which model you are looking at while you score is exactly
how expectations become findings.

## What came back right

Nearly everything atmospheric.

The mediterranean courtyard brief asked for warm lime render, dark bronze window
frames, gravel rather than lawn, morning light from the left, and no cars or people
in the frame. Across nine frames, the render was warm and lime-coloured every time,
the frames were dark every time, the ground was gravel every time, and not one
frame contained a car or a person.

The converted-barn kitchen brief asked for exactly three materials — pale oak,
honed black stone, lime plaster — with no upper cabinets and a single pendant light.
Most frames delivered precisely that: three materials, a long island, one light,
original trusses overhead. The failures were small and specific: one frame left
objects on the worktop after the brief said no decoration on the surfaces; another
put open shelves with crockery on the wall, which is both upper storage and
decoration in one move.

Exteriors scored **5.00 out of 5** on open briefs and **4.56** on constrained ones.
Interiors scored **4.89** and **4.56**. This is a high standard being met
consistently, and it is worth saying plainly before the criticism starts.

## What came back wrong

Anything you could count.

Floor plans and masterplans both averaged **4.11 out of 5**, and the misses cluster
so tightly that they read like a single defect:

- **Duplicated room numbering — the single most common failure in the entire run.**
  "BEDROOM 3" appearing twice on the same plan, or "Bedroom 1" labelling two
  different rooms. It occurred in seven of the nine floor plans in the constrained
  set.
- **A requested count silently short.** A masterplan brief asking for twelve
  dwellings in four terraced rows came back with eight. A three-bedroom plan came
  back with two, twice.
- **Programme items quietly dropped.** A brief specifying a bathroom *plus* a
  separate WC came back with the two merged; a utility room asked for by name
  simply was not drawn.

And in the worst cases, the labels themselves dissolved: "Kitcheng", "Kitchem 2",
"circulation en kliving", "Opparad Room". Text rendering in generated images is a
known weakness, but on a floor plan the labels are not decoration — they are the
content.

The thing that makes this dangerous is that none of it is visible at a glance. A
plan with two rooms both labelled "Bedroom 3" looks completely correct in a
thumbnail. It looks correct in a client presentation viewed from across a table. It
stops looking correct at exactly the moment someone tries to use it.

## Why the split exists

Atmosphere is a statistical property. A dusk render resembles other dusk renders;
warm lime render resembles other warm lime render. The model has seen an enormous
number of examples and reproduces the distribution well. Asking for "morning light
from the left" is asking it to land inside a region of a space it knows intimately.

A count is not a statistical property. It is a discrete fact that has to survive
the trip from your sentence into the picture, and there is no mechanism in image
generation that enforces it. The model has seen many floor plans with bedrooms
labelled 1, 2 and 3, so it produces something that *looks* like a numbered plan.
Whether the numbers are unique is not a thing it is checking, because it is not
checking anything — it is generating a plausible image.

This is why "the AI is getting better" does not straightforwardly fix it. Sharper
renders and more convincing materials are improvements along the axis the models
are already good at.

## Writing a tighter brief helps — but not the way you would hope

We measured this directly. Comparing repeat runs of the same brief, tightly
specified briefs produced measurably less variation than open ones: **46.65 against
51.33** in mean pixel difference, with composition agreement of **0.078 against
0.066**. Ask for more and you get a narrower range of answers, and the camera in
particular settles down.

But on adherence, constrained briefs scored slightly *lower*: **4.42 against 4.53**.
That is not a contradiction — it is arithmetic. A brief with eight requirements has
eight chances to lose one. An open brief asking for "a calm living room with a
plaster fireplace" is nearly impossible to fail.

So a longer brief buys you predictability, not obedience. Both are useful. Neither
is verification.

## You cannot re-roll your way out of it

The obvious workaround — generate again until the count is right — works less well
than instinct suggests, because a repeat is not a variation.

Across 24 brief-and-model combinations, composition overlap between two runs of the
same brief averaged **0.07 out of 1.0**. Not a near-miss version of the first image:
a different image, from the same words. Floor plans were the steadiest of the four
modes and still varied at 34.77 against 52.73 for exteriors — steadier, not stable.

Practically, this means a generated frame you like is a one-off. If it is nearly
right, do not throw it back; keep it and edit it. And if the count is wrong on a
plan, generating again gives you a different plan with a different wrong count.

## What to do instead

**Count it yourself, every time.** This is unglamorous and it is the whole answer.
Before a generated plan goes anywhere, read the labels and count the rooms. It takes
fifteen seconds and it catches the failure that survives every other check.

**Keep countable requirements few and bare.** Do not bury "three bedrooms" in the
middle of a paragraph about morning light and lime plaster. State the programme as
a short list. The model dilutes what it has to weigh against everything else.

**Use prohibitions — they work.** "No cars, no people" was honoured in every frame
that carried it. Negative constraints turned out to be among the most reliably
followed instructions in the run, which is the opposite of the common advice that
models ignore negations.

**Treat plans as studies, not documents.** A generated plan is worth having while
the programme is still moving and you want to see six layouts before lunch. The
moment the layout is settled, it needs drawing properly. Nothing in this run
suggests otherwise, and pretending otherwise is how a duplicated bedroom number
ends up in front of a client.

**Match the mode to the phase.** Exteriors and interiors are where these models are
strongest and where the atmospheric judgment you need is exactly what they provide.
Plans and masterplans are where you should slow down and verify.

## The method, in full

Everything above comes from two published runs, and both include the prompts
verbatim, the measurement scripts and the raw data:

- [Three models, one set of briefs](/benchmarks/model-selection/) — the 72-image
  run, blind-scored brief adherence, and what each model does differently
- [Ask for the same thing three times](/benchmarks/repeatability/) — how much two
  runs of the same brief differ, by mode and by brief style

The scoring rubric, the checklists written before viewing, and the verbatim list of
what each frame missed are all in the run data. We publish the misses rather than
only the scores, because a list of what went wrong is more use to you than a number.

For the related question of what happens when you edit an image over and over
rather than generating it fresh, see [what eight sequential edits do to an AI
architecture image](/blog/sequential-editing-quality-degradation/).

---

**Try Nuit free — 10 generations, no card required.** Write the brief, get the
concept, and keep every version in one project. [Start your project →](https://studio.nuit.archi)
