How it worksShowcaseBenchmarksWhy NuitPricing Start free
Benchmark

Sequential editing: what eight edits do to one image

Eight architectural edits applied in sequence to a single generated exterior, three times — once through each of three editing engines. Every intermediate frame saved, every metric computed by script.

Run executed
27 July 2026
Published
28 July 2026
Seed
1408x768, one image
Chain length
8 sequential edits
Status
Current
Engines tested
EngineModeEndpoint and parameters
Flux Kontext ProBlack Forest Labs (via fal) Full-frame prompt fal.run/fal-ai/flux-pro/kontextflux-pro/kontext, aspect_ratio 16:9, safety_tolerance 5
SeedEdit 3.0ByteDance (via fal) Full-frame prompt fal.run/fal-ai/bytedance/seededit/v3/edit-imageseededit v3, default parameters
Recraft inpaint v3Recraft Mask-based inpaint external.api.recraft.ai/v1/images/inpaintrecraftv3, multipart image + mask

Disclosure: Flux Kontext Pro returned a different pixel grid than the input and was resampled to 1408x768 before measurement.

We took one exterior generated in Nuit and applied eight architectural edits to it, in order, three separate times — once through each of three editing engines. Each edit was applied to the previous output, so this measures what iterating actually does to an image rather than what a single edit does.

The point is not to crown a winner. It is that the two available approaches fail in opposite directions, and knowing which failure you are buying changes how you work.

Every frame

Three chains, eight edits each, applied to the previous output. Pick a step to compare it against the seed and the final frame.

Every frame of all three chains. Each column applies one more edit to the previous column of that row.
Engine Seed Step 1 roof Step 2 path Step 3 gable end Step 4 chimney Step 5 sconces Step 6 door Step 7 deck Step 8 base
Flux Kontext Pro prompt-only Seed The seed image before any edits: a glass cabin in a fir forest at dusk. Step 1 Flux Kontext Pro, after edit 1 (roof). Step 2 Flux Kontext Pro, after edit 2 (path). Step 3 Flux Kontext Pro, after edit 3 (gable end). Step 4 Flux Kontext Pro, after edit 4 (chimney). Step 5 Flux Kontext Pro, after edit 5 (sconces). Step 6 Flux Kontext Pro, after edit 6 (door). Step 7 Flux Kontext Pro, after edit 7 (deck). Step 8 Flux Kontext Pro, after edit 8 (base).
SeedEdit 3.0 prompt-only Seed The seed image before any edits: a glass cabin in a fir forest at dusk. Step 1 SeedEdit 3.0, after edit 1 (roof). Step 2 SeedEdit 3.0, after edit 2 (path). Step 3 SeedEdit 3.0, after edit 3 (gable end). Step 4 SeedEdit 3.0, after edit 4 (chimney). Step 5 SeedEdit 3.0, after edit 5 (sconces). Step 6 SeedEdit 3.0, after edit 6 (door). Step 7 SeedEdit 3.0, after edit 7 (deck). Step 8 SeedEdit 3.0, after edit 8 (base).
Recraft inpaint v3 masked Seed The seed image before any edits: a glass cabin in a fir forest at dusk. Step 1 Recraft inpaint v3, after edit 1 (roof). Step 2 Recraft inpaint v3, after edit 2 (path). Step 3 Recraft inpaint v3, after edit 3 (gable end). Step 4 Recraft inpaint v3, after edit 4 (chimney). Step 5 Recraft inpaint v3, after edit 5 (sconces). Step 6 Recraft inpaint v3, after edit 6 (door). Step 7 Recraft inpaint v3, after edit 7 (deck). Step 8 Recraft inpaint v3, after edit 8 (base).

What the measurements say

Computed by scripts/bench-metrics.py over the saved frames — not judged by eye. Seed value is the baseline each engine started from.

Drift from seed

Mean absolute pixel difference against the original. Higher = further from where you started.

Engine 12345678
Flux Kontext Pro 14.114.917.618.318.919.621.021.8
SeedEdit 3.0 13.720.222.325.027.729.234.134.6
Recraft inpaint v3 8.914.015.716.116.617.317.718.6

Detail retained

Laplacian variance relative to the seed. 1.00 = as sharp as the original; 0.02 = detail is gone.

Engine 12345678
Flux Kontext Pro 0.850.820.800.790.810.850.880.89
SeedEdit 3.0 0.420.190.100.040.020.010.020.02
Recraft inpaint v3 0.650.560.540.510.500.490.490.50

Edge-detail overlap

Overlap of the high-frequency edge map with the seed. A strict measure of fine detail, not of composition — a frame can hold its camera and massing exactly and still score low here. Read the ordering and the decay, not the absolute value.

Engine 12345678
Flux Kontext Pro 0.260.260.240.240.240.230.220.22
SeedEdit 3.0 0.260.190.170.150.130.120.090.08
Recraft inpaint v3 0.500.360.340.330.320.310.300.29

Tone

Mean red channel, seed to final. This is the drift you notice before any other: the image quietly stops being the light you designed for.

EngineSeedFinalChange
Flux Kontext Pro 36.5 21.5 -15.1 darker
SeedEdit 3.0 36.5 37.4 + 0.9 warmer
Recraft inpaint v3 36.5 48.6 + 12.0 warmer

The mask concentrates the change. It does not contain it.

Comparing the final masked frame to the seed outside the union of all eight masks (which covers 28.97% of the frame): 89.7% of those pixels moved by more than 2/255, and 41.8% moved by more than 8/255. Mean change outside the mask was 9.55, against 40.92 inside it — a ratio of about 4.3x. Our own internal design note claimed everything outside the mask stayed pixel-identical. It does not, and we only found out by measuring.

The eight edits, verbatim

Published in full so the run can be repeated. The Landed column is our own judgement by eye on the masked chain — the only figure on this page that is not computed. 4 of 8 landed.

# Instruction sent Landed (masked)
1 Replace the flat roof of the cabin with a pitched gable roof clad in dark standing-seam metal. Keep everything else in the scene identical. Mask fill prompt: a pitched gable roof with dark standing-seam metal cladding Yes
2 Replace the wooden boardwalk path with a flagstone stone path. Keep everything else in the scene identical. Mask fill prompt: a flagstone stone path of wet natural flat stones Yes
3 Change the left glass gable-end wall of the cabin into a solid wall clad in vertical timber boards. Keep everything else in the scene identical. Mask fill prompt: a solid exterior wall clad in vertical timber boards Came back as differently-detailed glass. The mask sat on a glazed wall and the surrounding context is glass, so the fill followed the context rather than the instruction. No
4 Add a slim dark metal chimney flue on the roof of the cabin. Keep everything else in the scene identical. Mask fill prompt: a slim dark metal chimney flue rising from the roof Yes
5 Add two exterior wall sconce lights on the frame beside the entrance door. Keep everything else in the scene identical. Mask fill prompt: exterior wall sconce lights mounted beside the entrance Nothing appeared. Adding a small object into a region that does not already contain one is the weakest case for inpainting. No
6 Replace the entrance door with an oversized pivoting timber door. Keep everything else in the scene identical. Mask fill prompt: an oversized pivoting timber entrance door The door region was re-rendered but not as a pivoting timber door. No
7 Add a wooden slat pergola canopy over the front deck of the cabin. Keep everything else in the scene identical. Mask fill prompt: a wooden slat pergola canopy over the deck No pergola appeared. Same failure mode as the sconces: a new structure in open space. No
8 Replace the base under the cabin with a solid natural stone plinth. Keep everything else in the scene identical. Mask fill prompt: a solid natural stone plinth base under the cabin Yes

What we concluded

  1. There is no winner. The two approaches fail in opposite directions, and which failure you can live with depends on what you are doing.
  2. Prompt-only editing always changes something, but never only what you asked for. SeedEdit lost 98% of the seed's high-frequency detail by step 8 — the image is still recognisable, but it is no longer photographic. Kontext held detail and instead walked the image 15 points darker.
  3. Masked inpainting held composition best of the three — camera, framing and massing are unchanged from the seed — and still did not deliver on its own promise: 89.7% of the pixels OUTSIDE the mask moved by more than 2/255 over the chain. The mask concentrates change roughly 4x; it does not isolate it.
  4. Recraft declined half the edits. Four of eight landed. Replacing a material that is already there works; adding an object into empty space does not.
  5. Kontext is the better tool if what you need is detail retention across a long chain — it beat our own routing choice on that axis, and we are not going to pretend otherwise.

Reproducing this

Everything needed to repeat the run is downloadable here — not a list of paths in a repository you cannot open. Bring your own fal and Recraft API keys; the scripts read them from the environment.

The measurements, scripts and seed image are published under CC BY 4.0 — quote them, chart them, disagree with them; just link back. The edited frames are outputs of third-party models and are reproduced here for comparison and commentary; we make no licensing claim over them. Model versions change, so this page carries the date the run was executed, not just when it was written.

Questions

Which AI image editor is best for architecture?

Neither category wins outright, which is why this benchmark has no leaderboard. Full-frame prompt editors (Flux Kontext, SeedEdit) always apply a change but shift the whole image while doing it. Mask-based inpainting (Recraft) preserves composition far better but refused half of our eight edits, because it fills a region toward your prompt rather than executing an instruction. Choose by what the edit is: material swaps favour masks, scene-wide changes favour prompt editing.

How much does an AI image degrade over repeated edits?

It depends entirely on the engine, and the failure is not uniform. Measured over eight sequential edits on one 1408x768 exterior: SeedEdit 3.0 retained 2% of the seed's Laplacian sharpness, a near-total loss of fine detail. Flux Kontext Pro retained 89% of sharpness but drifted 15 points darker in mean red channel. Recraft inpaint retained 50% of sharpness and drifted 12 points warmer. Every one of those numbers is computed by a published script, not judged by eye.

Does inpainting really leave the rest of the image untouched?

No. This is the claim we ourselves got wrong until we measured it. Comparing the final frame to the seed outside the union of all eight masks, 89.7% of pixels differ by more than 2/255 and 41.8% differ by more than 8/255, with a mean delta of 9.55. What inpainting does preserve is composition: camera position, framing and massing hold, and the structural edge overlap with the seed stayed roughly twice as high as either prompt-only engine. Geometry holds; tone does not.

Why did the inpainting model refuse some edits?

Inpainting fills the masked region toward your prompt using the surrounding pixels as context. When the surroundings already contain the thing being replaced — a roof, a path, a plinth — the fill follows the instruction. When you ask for a new object in open space, such as a pergola over a deck or sconces beside a door, context wins and nothing appears. Four of our eight edits failed for exactly this reason.

How can I reproduce this benchmark?

Everything needed is published. The seed image, the eight prompts verbatim, the masks, the harness that ran the chains (scripts/edit-chain-test.mjs) and the script that computed every metric (scripts/bench-metrics.py) are all in the repository, and the raw per-step measurements are served as run.json next to this page. The run date and the exact endpoint and parameters for each engine are recorded above, because model versions change and an undated benchmark becomes false without anyone editing it.

Related reading