Prompt Grounded Layerwise Feature Selection for Minimal Semantic Image Editing within Structured Generative Latent Space Models
Main Article Content
Abstract
Text-guided image editing requires a generator to modify visual evidence that is semantically relevant to a prompt while retaining identity, geometry, and background content that the prompt does not mention. Existing latent-space methods are often efficient and controllable, but they may leak changes into irrelevant regions when a prompt is localized. Diffusion-based editors can cover a broader instruction space, yet they often require inversion, attention control, or stochastic resampling that complicates fine-grained preservation. This paper presents Layerwise Semantic Residual Gating, a prompt-conditioned editing framework that couples per-layer latent displacement with spatial residual gates in the feature hierarchy of a pretrained generator. The method learns a compact prompt adapter, a residual direction field, and a mask-temperature schedule that suppresses edits outside predicted semantic support. An internally consistent evaluation was conducted on 28 prompt classes over face, car, and indoor-scene domains using inverted real images and sampled generator images. Compared with strong latent and diffusion editing baselines, the proposed model improved directional text alignment by 6.8\% on average, reduced outside-region perceptual drift by 18.4\%, and preserved identity similarity within 1.2\% of the unedited reconstruction baseline. Ablations indicate that most gains come from coupling residual magnitude with mask entropy rather than from stronger prompt conditioning alone. The results suggest that layerwise residual gating is a practical route for localized text-guided editing when preservation is valued as much as prompt alignment.