Public image models, tested on architecture
We run public models on real architectural tasks and publish what we measure. Every run carries its date, the exact version and endpoint of everything tested, and the prompts needed to repeat it.
Latest run
Three image models, one set of briefs
Gemini 3.1 Flash Lite Image is about three times faster than the other two: a median of 5.0 seconds against 15.2 and 15.8. It returns the same pixel count as Gemini 3.1 Flash Image.

Ask for the same thing three times
No model repeats itself. Composition overlap between two runs of the same brief averages 0.07 out of 1.0: you do not get the same picture back with small differences, you get a different picture. If a frame is worth keeping, keep it; you cannot re-roll your way back to it.

Sequential editing: what eight edits do to one image
There is no winner. The two approaches fail in opposite directions, and which failure you can live with depends on what you are doing.
Dated and versioned
Models change under their own names. Every run states when it was executed and against which endpoint and parameters.
Reproducible
Seed images, prompts verbatim, masks, the harness and the measurement script are published. If you cannot repeat it, it is not a benchmark.
No total score
Models are good at different things. We publish the numbers, including where a model loses, and leave the ranking to the job you have.
Meta Muse Image 1.0 against the July line-up
The eight briefs of the model comparison put through Meta Muse Image 1.0, measured by the same scripts so the published July numbers stay comparable rather than being rewritten.
Run not yet executedStyle consistency across modes
Whether an exterior, its floor plan and its interiors read as the same building when generated from one brief.
Not started