Machine Unlearning
A Standardized Multi-Architecture Testbed of Memorized Concepts, Generation Fingerprints, and Editable Structures in Text-to-Image Models Beyond Stable Diffusion v1.x
Generated automatically from the limitations stated in 22 papers (CVPR, ICCV, ICML), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Nearly every quantitative claim about text-to-image models — memorization, unlearning success, watermark robustness, deepfake provenance, concept editability — is anchored to a handful of frozen Stable Diffusion v1.x/v2 checkpoints, because those are the only models with established benchmarks, known memorized prompts, and validated concept labels. Newer architectures (DiT-based diffusion, DeepFloyd, autoregressive and GAN generators, video models) are largely untested, so results that appear robust may be artifacts of one 2022 model family. The field cannot tell whether its findings are properties of generative modeling or properties of Stable Diffusion 1.4.
Why it matters
Once findings can be tested on more than one model family, papers can claim properties of generative models rather than of Stable Diffusion 1.4, and new architectures can be evaluated for unlearning, watermarking, and provenance at release time rather than years later.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Build and release an architecture-stratified evaluation suite: a memorized-prompt discovery pipeline (membership-inference and replication-rate scoring) run over SD v1.x/v2, a DiT-based T2I model (e.g., PixArt/SD3), DeepFloyd, and at least one autoregressive and one GAN generator, producing per-model sets of verified memorized prompts, concepts, and internal activation statistics. Measure replication rates, cross-architecture agreement of memorization metrics, and correlation with existing SD v1.4 benchmarks.
- 2
Port three representative tasks — one unlearning method, one watermark/detector pair, one concept-editing method — onto the suite and quantify exactly which findings transfer and which collapse. Measure task success metrics per architecture before and after controlling for training-data curation and deduplication level.
- 3
Characterize which architectural and data properties (UNet vs. DiT, latent dimensionality, deduplication) predict whether memorization and intervention findings transfer, via regression over the measured per-model statistics. Measure predictive accuracy of transfer between architecture pairs.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Memorization in modern models may be too weak or too diffuse to construct a comparable benchmark at all, leaving no common ground across architectures on which to measure transfer.
Sub-problems
Cross-Architecture Generalization and Robustness of Watermarking and Attribution in Text-to-Image Generation
Current watermarking, attribution, and concept erasure methods in text-to-image synthesis have almost exclusively been evaluated on early U-Net diffusion models, particularly Stable Diffusion v1.x and v2.x. Because these techniques rely heavily on specific latent spaces, VAE architectures, and diffusion noise schedules, their cross-architecture transferability remains entirely unverified across modern generative paradigms such as Diffusion Transformers (DiTs), autoregressive models, and discrete tokenizers. Consequently, practitioners cannot determine whether existing provenance and safety guarantees are fundamental or merely artifacts of early Stable Diffusion parameterizations.
Cross-Architecture Latent Translation Benchmark Across GANs, VAEs, and Diffusion Models
Image translation and latent editing techniques are currently developed and evaluated in isolation within specific model families, assuming either compact low-dimensional latent spaces (GANs and VAEs) or specific diffusion backbones (such as Stable Diffusion v2.1). Because no unified evaluation or bridging mechanism exists across these different latent formulations, latent manipulation methods cannot be transferred between GANs and diffusion models without costly ad-hoc redesigns or full-model fine-tuning. Consequently, researchers cannot determine whether existing translation paradigms generalize across modern diffusion architectures, including both UNet and Transformer-based backbones.
Evidence
Each paper's own statement of the limitation, verbatim.
- Unified Customized Generation by Disentangled Reward ModelingCVPR 2026
Relies on a specific hierarchical projector and dual-encoder architecture that may not easily generalize to non-Transformer-based diffusion models.
- Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion ModelsCVPR 2026
Evaluated only on Stable Diffusion v1.5 with three fixed target scenarios and small 3,300-sample fine-tuning sets; generalization to other T2I architectures or larger-scale training is untested
- TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion ModelsCVPR 2026
Evaluated on a limited concept set (nudity via I2P, Van Gogh style, tench object) on Stable Diffusion v1.4 only; generalization to other models and concepts is not demonstrated
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model UnlearningICLR 2026
Evaluated only on Stable Diffusion (V1.4, 2.0, 2.1) backbones; generalization to other architectures, video, or large multi-modal models is untested
- Forget Many, Forget Right: Scalable and Precise Concept Unlearning in Diffusion ModelsICLR 2026
Empirical validation is restricted to Stable Diffusion v1.4 and v1.5 CNN/UNet-based cross-attention layers, leaving scalability to modern Diffusion Transformer (DiT) architectures unverified.
- Finding DoRI: Discovery of Retained Images in Diffusion ModelsICML 2026
Comprehensive evaluation is constrained to Stable Diffusion v1.4 and v2.0 due to the lack of public memorization benchmarks for newer diffusion architectures.
- ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned ModelsICCV 2025
Evaluation is limited to Stable Diffusion 1.4-based unlearned models and specific concept datasets (NSFW, I2P subsets, 50-prompt style/object sets); transferability to other architectures or MU methods is untested.
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsCVPR 2025
Experiments are limited to Stable Diffusion v1.4 and a threshold-based revival criterion (CLIP threshold set 0.02 below original model score, classifier accuracy 0.3), which may not capture subtle residual traces
Show all 22 papers
- StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion ModelsNeurIPS 2025
Built and evaluated only on Stable Diffusion 2.1's VAE; transferring to other LDMs or VAE architectures is not demonstrated
- Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier IntegrityICCV 2025
Evaluated only on Stable Diffusion v2-1-base at 512x512; the authors note extension to other generative architectures is future work
- Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware OptimizationICCV 2025
Only evaluated on Stable Diffusion v1.4 with up to 10 concepts from a 10-class ImageNet subset (Imagenette); scalability to larger or hierarchically compositional concept sets is not demonstrated
- IntroStyle: Training-Free Introspective Style Attribution using Diffusion FeaturesICCV 2025
Evaluated only on Stable Diffusion v2.1; generalization to other pre-trained diffusion models is untested
- Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation ModelsCVPR 2025
Gains depend on Stable Diffusion v1.4 and known concept labels; performance on sequential/multi-concept unlearning or models without predictable latent clustering is not demonstrated
- Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion ModelsCVPR 2025
Only evaluated on Stable Diffusion v1-4 with a single 500-prompt memorization dataset (~300 memorized prompts); generalization to other models or natural prompt distributions is not shown.
- CopyrightShield: Enhancing Diffusion Model Security Against Copyright Infringement AttacksICCV 2025
Validated only on virtually designed datasets (Pokemon BLIP Captions, COYO samples with Midjourney images) and only on Stable Diffusion v1.x, not real-world copyright-protected content
- Community Forensics: Using Thousands of Generators to Train Fake Image DetectorsCVPR 2025
The dataset is heavily skewed toward diffusion-based models (mostly Stable Diffusion derivatives), with limited GAN, VQ-VAE, and autoregressive coverage
- SMaRt: Improving GANs with Score Matching RegularityICML 2024
Still does not outperform multi-step diffusion models (e.g., ADM FID 2.07 vs StyleGAN2+SMaRt 18.31 on ImageNet 64x64); requires a suitable pre-trained diffusion model for the target domain
- Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion ModelsNeurIPS 2024
Only evaluated on Stable Diffusion v1.4; the authors could not identify properly memorized prompts in Stable Diffusion v2 or Deep Floyd due to curated/deduplicated training data, so applicability to newer models is untested
- DataDream: Few-shot Guided Dataset GenerationECCV 2024
Requires fine-tuning a diffusion model and generating 500 images per class per dataset (and per class for DataDream_cls), a substantial compute cost compared to prompt-only baselines; results not saturated at 500 images/class, implying more compute for further gains
- TR0N: Translator Networks for 0-Shot Plug-and-Play Conditional GenerationICML 2023
Not straightforwardly applicable to diffusion models, which lack a low-dimensional latent space for the translator to map to; GANs and VAEs with explicit low-dimensional latents are assumed
- Raising the Cost of Malicious AI-Powered Image EditingICML 2023
Not guaranteed forward-compatible with future versions of diffusion models; relies on uncertain adversarial transferability
- Exploring Incompatible Knowledge Transfer in Few-Shot Image GenerationCVPR 2023
Experiments are limited to GAN architectures (ProGAN, StyleGAN-V2); extension to VAEs or diffusion models is left as future work
Nearest existing work
- Finding DoRI: Discovery of Retained Images in Diffusion ModelsICML 2026
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsCVPR 2025
- Towards Memorization-Free Diffusion ModelsCVPR 2024
- Detecting, Explaining, and Mitigating Memorization in Diffusion ModelsICLR 2024
- Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware OptimizationICCV 2025
- Image-level Memorization Detection via Inversion-based Inference PerturbationICLR 2025
- TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion ModelsCVPR 2026
- An Inversion-based Measure of Memorization for Diffusion ModelsICCV 2025
- CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion ModelsNeurIPS 2025
- Data Unlearning in Diffusion ModelsICLR 2025
- Does Generation Require Memorization? Creative Diffusion Models using Ambient DiffusionICML 2025
- Forget-It-All: Multi-Concept Machine Unlearning via Concept-Aware Neuron MaskingICML 2026
- Black-Box Forgery Attacks on Semantic Watermarks for Diffusion ModelsCVPR 2025
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model UnlearningICLR 2026
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization PerspectiveICLR 2026
Related open problems
Machine Unlearning
Benchmarking the Adversarial Robustness and Reversibility of LLM Unlearning
Existing LLM unlearning methods are predominantly evaluated on standard, benign queries, giving a false sense of compliance with privacy and copyright demands. Empirical evidence shows that "unlearned" knowledge remains extractable via adversarial jailbreaks, latent-space elicitation, and few-shot relearning. Without systematic evaluation across these extraction vectors, practitioners have no way to verify whether a model has actually eliminated sensitive data or merely applied a superficial suppression mask.
Machine Unlearning
Empirical Robustness and Sensitivity of Machine Unlearning Under Realistic Data and Checkpoint Constraints
Existing machine unlearning algorithms for language models are predominantly evaluated under idealized conditions: complete access to forget sets, clean entity anchors, standard prose, and checkpoints saved immediately before target data exposure. In practice, unlearning requests frequently present partial forget data, domain variations like code, noisy or alias-heavy entity mentions, and checkpoints separated from target exposure by billions or trillions of tokens. Practitioners currently cannot predict whether an unlearning method that succeeds on curated benchmarks like TOFU will retain any efficacy when deployed under these real-world data and provenance constraints.
Machine Unlearning
Machine Unlearning Under Degraded Operational Preconditions
Existing machine unlearning methods rely on strict operational preconditions: access to clean retain datasets ($D_r$), full white-box parameter access, historical pre-training checkpoints, or original pre-unlearned calibration weights. In production environments, intermediate training checkpoints are routinely deleted to save storage, retain data is often inaccessible due to data governance policies, and deployment interfaces may restrict full weight access. Without methods that function in the absence of these preconditions, deployed models cannot legally or practically comply with data deletion requests.
Machine Unlearning
Scalability and Robustness of Machine Unlearning Under Varying Forget-Set Regimes
Existing machine unlearning techniques are only validated in narrow operating windows, typically limited to small forget budgets ($\le 10\%$). As forget sets scale beyond a few hundred samples or exceed 10–20% of the training distribution, retention accuracy drops significantly (e.g., a 14–15% gap relative to retraining from scratch on CIFAR-10), while tiny splits (e.g., 1%) yield negligible unlearning. Furthermore, unlearned representations remain fragile and easily recoverable when exposed to small fractions of the forgotten data during relearning. Consequently, current unlearning algorithms cannot be reliably deployed for large-scale data deletion compliance without catastrophic utility loss.