Benchmarks

Public image models, tested on architecture

We run public models on real architectural tasks and publish what we measure. Every run carries its date, the exact version and endpoint of everything tested, and the prompts needed to repeat it.

Latest run
Runs 3 published
A courtyard house with a deep flat roof overhang and an olive tree, one of the frames in the model comparison

Run 28 July 2026 · 3 models · 8 briefs

Three image models, one set of briefs

Gemini 3.1 Flash Lite Image is about three times faster than the other two: a median of 5.0 seconds against 15.2 and 15.8. It returns the same pixel count as Gemini 3.1 Flash Image.

A lime-rendered courtyard house, one of the frames in the repeatability run

Run 28 July 2026 · 3 models · 8 briefs

Ask for the same thing three times

No model repeats itself. Composition overlap between two runs of the same brief averages 0.07 out of 1.0: you do not get the same picture back with small differences, you get a different picture. If a frame is worth keeping, keep it; you cannot re-roll your way back to it.

A glass cabin on a stone plinth after eight sequential edits

Run 27 July 2026 · 3 engines · 8 edits

Sequential editing: what eight edits do to one image

There is no winner. The two approaches fail in opposite directions, and which failure you can live with depends on what you are doing.

How we run these

Dated and versioned

Models change under their own names. Every run states when it was executed and against which endpoint and parameters.

Reproducible

Seed images, prompts verbatim, masks, the harness and the measurement script are published. If you cannot repeat it, it is not a benchmark.

No total score

Models are good at different things. We publish the numbers, including where a model loses, and leave the ranking to the job you have.

Planned Not yet measured

Meta Muse Image 1.0 against the July line-up

The eight briefs of the model comparison put through Meta Muse Image 1.0, measured by the same scripts so the published July numbers stay comparable rather than being rewritten.

Run not yet executed

Style consistency across modes

Whether an exterior, its floor plan and its interiors read as the same building when generated from one brief.

Not started