ACM Multimedia 2026

Rethinking Where to Edit

Task-Aware Localization for Instruction-Based Image Editing

Two editing examples. Adding a lamp: the base model also beautifies the person; ours adds the lamp and keeps the face unchanged. Removing lotus flowers: the base model shifts the elephant statue; ours removes the flowers and keeps the statue in place.

A strong base editor (Qwen-Image-Edit) follows the instruction but also changes content it was never asked to touch, such as the person's appearance or the position of the statue. Our framework predicts where the edit should happen and keeps everything else anchored to the source image.

Abstract

Instruction-based image editing (IIE) aims to modify an image according to a natural language instruction. Despite recent advances in diffusion transformers, existing methods often introduce unintended changes to regions unrelated to the requested edit. We attribute this limitation to the absence of an explicit mechanism for edit localization. Different editing operations (e.g., subject addition, removal, and replacement) induce distinct spatial patterns, yet existing IIE models typically perform localization in a task-agnostic manner.

To address this limitation, we propose a training-free, task-aware edit localization framework that exploits the intrinsic source and target image streams of IIE models. For each image stream, we construct feature centroids from attention-based edit cues, and then partition tokens into edit and non-edit regions based on feature similarity. Observing that effective localization is inherently task-dependent, we introduce a unified mask construction strategy that selectively leverages the source and target streams according to the editing task. We also provide a systematic analysis of our underlying insights and design choices. Extensive experiments on EdiVal-Bench demonstrate that our framework consistently improves content consistency in non-edit regions while maintaining strong instruction-following performance on top of Step1X-Edit and Qwen-Image-Edit.

Key Insight: Localization Is Task-Dependent

Modern DiT-based editors process text, target image tokens and source image tokens in one joint attention sequence. Edit semantics do not emerge uniformly across these streams. Where to look depends on what the instruction asks for.

Subject addition

target stream

M̂ = M̂tgt

The new object does not exist in the source image, so its footprint appears where it emerges in the target stream.

Subject removal

source stream

M̂ = M̂src

The object to erase lives in the source image, so the source stream carries the localization signal.

Subject replacement

source∪target

M̂ = M̂tgt ∪ M̂src

Replacement couples a disappearance in the source with an emergence in the target, so both streams are combined.

Method

Our framework plugs into a pretrained dual-stream editor at inference time. It needs no training and no extra supervision, and its runtime and memory cost stay comparable to the base model.

Framework overview with four stages: attention-based semantic estimation, feature-based semantic assignment, task-aware mask construction, and mask-guided latent preservation inside a DiT image editor.
  1. Attention-based semantic estimation. We slice the joint attention matrix into text-to-image cross-attention and within-stream self-attention, then propagate the text signal one step through the self-attention graph. Aggregating over layers and instruction tokens gives a coarse attention mask for each image stream.
  2. Feature-based semantic assignment. Attention reflects relevance more than object structure, so we treat the coarse mask as a seed. Masked average pooling over deep DiT features yields an edit centroid and a non-edit centroid, and every token is assigned to the closer one by cosine similarity.
  3. Task-aware mask construction. The stream-wise masks are selected or merged according to the editing task, following the rules above, and lightly cleaned with connected-component filtering, hole filling and a small dilation.
  4. Mask-guided latent preservation. At a few early denoising steps, tokens outside the mask are replaced with an inverted source latent aligned to the current noise level, while tokens inside the mask evolve freely.

Where Do Edit Semantics Emerge?

Using SAM 3 pseudo ground-truth masks on Qwen-Image-Edit, we measure how well each signal localizes the edit across denoising timesteps. Two patterns hold across tasks: deep latent features (red) localize more accurately than attention maps with or without propagation, and the informative stream switches with the task.

IoU curves over denoising timesteps for subject addition, removal and replacement, comparing attention masks without propagation, with propagation, and feature-based masks for source and target streams.

IoU of attention-derived and feature-derived masks against pseudo ground truth. Feature-based masks lead in most settings; addition is best read from the target stream and removal from the source stream.

Qualitative comparison of attention-derived and feature-derived masks from source and target streams for an addition example (notebook) and a removal example (brick beige house).

Feature-derived masks have cleaner boundaries and fuller coverage than attention maps. The added “notebook” appears in the target stream, while the removed “brick beige house” is captured in the source stream.

Results on EdiVal-Bench

We apply the framework to Step1X-Edit and Qwen-Image-Edit and evaluate on EdiVal-Bench (572 real images, nine edit types). Content consistency (CC) improves on both backbones, and instruction following (IF) is maintained or slightly improved.

Method Base model EdiVal-IF ↑ EdiVal-CC ↑ EdiVal-O ↑ Perceptual
quality ↑
ObjectBackgroundOverall
InstructPix2PixSD 1.539.3477.7187.7982.7557.067.86
MagicBrushSD 1.542.6682.9093.6288.2661.367.89
UltraEditSD 351.5783.0493.5788.3167.487.91
ICEditFlux.1 Fill53.5088.2193.9291.0769.807.91
Step1X-EditStep1X-Edit59.0990.7397.3294.0374.547.86
+ GRAGStep1X-Edit59.62 +0.5391.63 +0.9097.58 +0.2694.60 +0.5775.10 +0.567.90 +0.04
+ OursStep1X-Edit60.84 +1.7591.77 +1.0497.80 +0.4894.79 +0.7675.94 +1.407.86 +0.00
Qwen-Image-EditQwen-Image-Edit70.8086.5193.7890.1479.897.96
+ GRAGQwen-Image-Edit67.31 −3.4991.17 +4.6695.49 +1.7193.33 +3.1979.26 −0.637.96 +0.00
+ OursQwen-Image-Edit71.15 +0.3591.78 +5.2796.96 +3.1894.37 +4.2381.94 +2.057.94 −0.02

EdiVal-IF, EdiVal-CC and EdiVal-O measure instruction following, content consistency and overall performance. Perceptual quality is a 0 to 10 naturalness score from a vision-language model. Differences are relative to each base model.

Qualitative Comparison

Each case shows the edited image next to its pixel-wise difference from the input. Brighter areas mean larger changes. Our difference maps concentrate on the requested region, while the base model also alters the fox, the woman's appearance and other untouched content.

Qualitative comparison of Qwen-Image-Edit, GRAG and ours on four instructions: add umbrella above the fox, remove chocolate cake, replace cat with decorative pumpkin, change background to garden, each with a difference map.

BibTeX

@article{he2026rethinking,
  title   = {Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing},
  author  = {He, Jingxuan and Wang, Xiyu and Zheng, Mengyu and Zeng, Xiangyu and Wang, Yunke and Xu, Chang},
  journal = {arXiv preprint arXiv:2604.20258},
  year    = {2026}
}