Benchmark · run 27 July 2026

Sequential editing: what eight edits do to one image

Eight architectural edits applied in sequence to a single generated exterior, three times — once through each of three editing engines. Every intermediate frame saved, every metric computed by script.

Open run.json
Results 8 edits in a row · 1408x768
Metric
Flux Kontext ProBlack Forest Labs (via fal)
SeedEdit 3.0ByteDance (via fal)
Recraft inpaint v3Recraft
Detail retainedSharpness against the seed after all edits. 1.00 is as sharp as the seed
Flux Kontext Pro 0.89
SeedEdit 3.0 0.02
Recraft inpaint v3 0.50
Drift from seedMean pixel difference, 0 to 255. Higher: further from the start
Flux Kontext Pro 21.8
SeedEdit 3.0 34.6
Recraft inpaint v3 18.6
Tone shiftMean red channel, seed to final
Flux Kontext Pro -15.1 darker
SeedEdit 3.0 +0.9 warmer
Recraft inpaint v3 +12.0 warmer
Edge-detail overlapFine-edge overlap with the seed; read the ordering, not the value
Flux Kontext Pro 0.22
SeedEdit 3.0 0.08
Recraft inpaint v3 0.29
Edits that landedJudged by eye on the masked chain only
Flux Kontext Pro —
SeedEdit 3.0 —
Recraft inpaint v3 4 of 8
What we found

Two ways to fail.

  1. There is no winner. The two approaches fail in opposite directions, and which failure you can live with depends on what you are doing.
  2. Prompt-only editing always changes something, but never only what you asked for. SeedEdit lost 98% of the seed's high-frequency detail by step 8 — the image is still recognisable, but it is no longer photographic. Kontext held detail and instead walked the image 15 points darker.
  3. Masked inpainting held composition best of the three — camera, framing and massing are unchanged from the seed — and still did not deliver on its own promise: 89.7% of the pixels OUTSIDE the mask moved by more than 2/255 over the chain. The mask concentrates change roughly 4x; it does not isolate it.
  4. Recraft declined half the edits. Four of eight landed. Replacing a material that is already there works; adding an object into empty space does not.
  5. Kontext is the better tool if what you need is detail retention across a long chain: it beat the masked engine on that axis.
The mask Union of masks covers 28.97% of the frame

The mask concentrates the change. It does not contain it.

Comparing the final masked frame to the seed outside all eight masks: 89.7% of those pixels moved by more than 2/255 and 41.8% by more than 8/255. Mean change outside the mask was 9.55, against 40.92 inside it, a ratio of about 4.3x. Pixels outside a mask are not left identical; composition is what holds.

Every frame Each edit is applied to the previous output

Every edit is applied to the previous output. Swipe along a row to step through.

Flux Kontext Pro Full-frame prompt editing
The seed image before any edits: a glass cabin with a flat roof in a fir forest.
Seed
Flux Kontext Pro after edit 1: roof
After edit 1: roof
Flux Kontext Pro after edit 2: path
After edit 2: path
Flux Kontext Pro after edit 3: gable end
After edit 3: gable end
Flux Kontext Pro after edit 4: chimney
After edit 4: chimney
Flux Kontext Pro after edit 5: sconces
After edit 5: sconces
Flux Kontext Pro after edit 6: door
After edit 6: door
Flux Kontext Pro after edit 7: deck
After edit 7: deck
Flux Kontext Pro after edit 8: base
After edit 8: base
Flux Kontext Pro after all 8 edits
Final, after 8 edits
SeedEdit 3.0 Full-frame prompt editing
The seed image before any edits: a glass cabin with a flat roof in a fir forest.
Seed
SeedEdit 3.0 after edit 1: roof
After edit 1: roof
SeedEdit 3.0 after edit 2: path
After edit 2: path
SeedEdit 3.0 after edit 3: gable end
After edit 3: gable end
SeedEdit 3.0 after edit 4: chimney
After edit 4: chimney
SeedEdit 3.0 after edit 5: sconces
After edit 5: sconces
SeedEdit 3.0 after edit 6: door
After edit 6: door
SeedEdit 3.0 after edit 7: deck
After edit 7: deck
SeedEdit 3.0 after edit 8: base
After edit 8: base
SeedEdit 3.0 after all 8 edits
Final, after 8 edits
Recraft inpaint v3 Mask-based inpainting
The seed image before any edits: a glass cabin with a flat roof in a fir forest.
Seed
Recraft inpaint v3 after edit 1: roof
After edit 1: roof
Recraft inpaint v3 after edit 2: path
After edit 2: path
Recraft inpaint v3 after edit 3: gable end
After edit 3: gable end
Recraft inpaint v3 after edit 4: chimney
After edit 4: chimney
Recraft inpaint v3 after edit 5: sconces
After edit 5: sconces
Recraft inpaint v3 after edit 6: door
After edit 6: door
Recraft inpaint v3 after edit 7: deck
After edit 7: deck
Recraft inpaint v3 after edit 8: base
After edit 8: base
Recraft inpaint v3 after all 8 edits
Final, after 8 edits
How it was run
Question

We took one generated exterior and applied eight architectural edits to it, in order, three separate times — once through each of three editing engines. Each edit was applied to the previous output, so this measures what iterating actually does to an image rather than what a single edit does.

The point is not to crown a winner. It is that the two available approaches fail in opposite directions, and knowing which failure you are buying changes how you work.

Engines and access
  • Flux Kontext Pro — full-frame prompt editing, Black Forest Labs (via fal)fal.run/fal-ai/flux-pro/kontext · flux-pro/kontext, aspect_ratio 16:9, safety_tolerance 5
  • SeedEdit 3.0 — full-frame prompt editing, ByteDance (via fal)fal.run/fal-ai/bytedance/seededit/v3/edit-image · seededit v3, default parameters
  • Recraft inpaint v3 — mask-based inpainting, Recraftexternal.api.recraft.ai/v1/images/inpaint · recraftv3, multipart image + mask
Seed
One 1408x768 image, 8 sequential edits per engine, every intermediate frame saved.
Measures
Computed by scripts/bench-metrics.py over the saved frames. The only figure not computed by script is whether an edit landed.
Disclosure
Flux Kontext Pro returned a different pixel grid than the input and was resampled to 1408x768 before measurement.
Reproduce
Measurements, scripts and the seed image are published under CC BY 4.0. The edited frames are outputs of third-party models, shown for comparison and commentary.
Scope

A dated snapshot of three public editing engines, executed on 27 July 2026. Providers update models under the same name; figures are left exactly as the run produced them. Providers update models without renaming them, so the date, endpoint and version of everything tested are stated on the page.

The eight edits, verbatim
1
Replace the flat roof of the cabin with a pitched gable roof clad in dark standing-seam metal. Keep everything else in the scene identical.Mask fill prompt: a pitched gable roof with dark standing-seam metal cladding
Landed
2
Replace the wooden boardwalk path with a flagstone stone path. Keep everything else in the scene identical.Mask fill prompt: a flagstone stone path of wet natural flat stones
Landed
3
Change the left glass gable-end wall of the cabin into a solid wall clad in vertical timber boards. Keep everything else in the scene identical.Mask fill prompt: a solid exterior wall clad in vertical timber boardsCame back as differently-detailed glass. The mask sat on a glazed wall and the surrounding context is glass, so the fill followed the context rather than the instruction.
Did not land
4
Add a slim dark metal chimney flue on the roof of the cabin. Keep everything else in the scene identical.Mask fill prompt: a slim dark metal chimney flue rising from the roof
Landed
5
Add two exterior wall sconce lights on the frame beside the entrance door. Keep everything else in the scene identical.Mask fill prompt: exterior wall sconce lights mounted beside the entranceNothing appeared. Adding a small object into a region that does not already contain one is the weakest case for inpainting.
Did not land
6
Replace the entrance door with an oversized pivoting timber door. Keep everything else in the scene identical.Mask fill prompt: an oversized pivoting timber entrance doorThe door region was re-rendered but not as a pivoting timber door.
Did not land
7
Add a wooden slat pergola canopy over the front deck of the cabin. Keep everything else in the scene identical.Mask fill prompt: a wooden slat pergola canopy over the deckNo pergola appeared. Same failure mode as the sconces: a new structure in open space.
Did not land
8
Replace the base under the cabin with a solid natural stone plinth. Keep everything else in the scene identical.Mask fill prompt: a solid natural stone plinth base under the cabin
Landed
Questions
Which AI image editor is best for architecture?

Neither category wins outright, which is why this benchmark has no leaderboard. Full-frame prompt editors (Flux Kontext, SeedEdit) always apply a change but shift the whole image while doing it. Mask-based inpainting (Recraft) preserves composition far better but refused half of our eight edits, because it fills a region toward your prompt rather than executing an instruction. Choose by what the edit is: material swaps favour masks, scene-wide changes favour prompt editing.

How much does an AI image degrade over repeated edits?

It depends entirely on the engine, and the failure is not uniform. Measured over eight sequential edits on one 1408x768 exterior: SeedEdit 3.0 retained 2% of the seed's Laplacian sharpness, a near-total loss of fine detail. Flux Kontext Pro retained 89% of sharpness but drifted 15 points darker in mean red channel. Recraft inpaint retained 50% of sharpness and drifted 12 points warmer. Every one of those numbers is computed by a published script, not judged by eye.

Does inpainting really leave the rest of the image untouched?

No. It is a common assumption, and measuring it disproves it. Comparing the final frame to the seed outside the union of all eight masks, 89.7% of pixels differ by more than 2/255 and 41.8% differ by more than 8/255, with a mean delta of 9.55. What inpainting does preserve is composition: camera position, framing and massing hold, and the structural edge overlap with the seed stayed roughly twice as high as either prompt-only engine. Geometry holds; tone does not.

Why did the inpainting model refuse some edits?

Inpainting fills the masked region toward your prompt using the surrounding pixels as context. When the surroundings already contain the thing being replaced — a roof, a path, a plinth — the fill follows the instruction. When you ask for a new object in open space, such as a pergola over a deck or sconces beside a door, context wins and nothing appears. Four of our eight edits failed for exactly this reason.

How can I reproduce this benchmark?

Everything needed is published. The seed image, the eight prompts verbatim, the masks, the harness that ran the chains (scripts/edit-chain-test.mjs) and the script that computed every metric (scripts/bench-metrics.py) are all in the repository, and the raw per-step measurements are served as run.json next to this page. The run date and the exact endpoint and parameters for each engine are recorded above, because model versions change and an undated benchmark becomes false without anyone editing it.

More
Three image models, one set of briefsSpeed, brief adherence and aspect ratio.