Watermarking & Deepfake Detection
Benchmarking Watermarking and Deepfake Detection Across Non-Latent and Flow-Matching Generative Architectures
Generated automatically from the limitations stated in 3 papers (CVPR, NeurIPS, ICLR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current watermarking and detection methods in generative media are almost exclusively developed and evaluated on standard latent diffusion models relying on Gaussian noise sampling. Consequently, it is unknown whether these techniques transfer, degrade, or fail entirely when applied to pixel-space diffusion, non-iterative architectures, or modern rectified-flow and flow-matching models such as SD3 and FLUX. As production generative pipelines shift away from standard latent diffusion, safety and provenance mechanisms risk operating under untested assumptions. A systematic evaluation across these architectural families is necessary to establish the empirical boundary of existing detection and watermarking schemes.
Why it matters
Establishes the exact architectural conditions under which provenance schemes fail, enabling researchers to design model-agnostic watermarks and detectors rather than relying on latent-specific artifacts.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Multi-architecture benchmark: Implement a unified testbed evaluating state-of-the-art watermarking and attribution methods across three architectural families (pixel-space diffusion like DDPM/Pixel-DiT, latent diffusion, and flow-matching models like SD3 and FLUX). Measure bit-error rate, watermark tamper-resistance, and receiver operating characteristic (AUROC) under standard post-processing perturbations.
- 2
Cross-architecture attack transferability: Evaluate how post-hoc deepfake detectors and watermark removal attacks trained on latent diffusion models transfer to pixel-space and rectified-flow outputs. Measure cross-model detection drop-off and attack success rate across varying sampling steps and guidance scales.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If concurrent work rapidly standardizes plug-and-play evaluation benchmarks for SD3/FLUX provenance before the study is complete, or if existing watermarking schemes turn out to be trivial to port to flow-matching pipelines with zero performance delta.
Evidence
Each paper's own statement of the limitation, verbatim.
- Forensic-Friendly Image Manipulation via Controllable Latent DiffusionCVPR 2026
Cannot be applied to diffusion models that do not rely on random noise sampling in latent space, e.g., SD3 and FLUX, as acknowledged by the authors
- Breaking Latent Prior Bias in Detectors for Generalizable AIGC Image DetectionNeurIPS 2025
Demonstrated primarily on latent diffusion models; efficacy on non-latent or non-iterative generative architectures is not established.
- Toward effective protection against diffusion-based mimicry through score distillationICLR 2024
Only applicable to latent diffusion models; the authors note pixel-space diffusion models are not covered and may require different attacks
Nearest existing work
- Black-Box Forgery Attacks on Semantic Watermarks for Diffusion ModelsCVPR 2025
- WaTeRFlow: Watermark Temporal Robustness via Flow ConsistencyCVPR 2026
- RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable GuaranteesNeurIPS 2024
- Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion PerspectiveICML 2026
- The Stable Signature: Rooting Watermarks in Latent Diffusion ModelsICCV 2023
- WMAdapter: Adding WaterMark Control to Latent Diffusion ModelsICML 2025
- TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion SensitivityICCV 2025
- MaxMark: High-Capacity Diffusion-Native Watermarking via Robust and Invertible Latent EmbeddingCVPR 2026
- SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery AttacksICML 2026
- Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion ModelsCVPR 2024
- How to Trace Latent Generative Model Generated Images without Artificial Watermark?ICML 2024
- Tree-Rings Watermarks: Invisible Fingerprints for Diffusion ImagesNeurIPS 2023
- Artificial Fingerprinting for Generative Models: Rooting Deepfake Attribution in Training DataICCV 2021
- GaussMarker: Robust Dual-Domain Watermark for Diffusion ModelsICML 2025
- TrustMark: Robust Watermarking and Watermark Removal for Arbitrary Resolution ImagesICCV 2025
Related open problems
Watermarking & Deepfake Detection
Black-Box Verification and Trigger Generation for Model Watermarking and Detection
Current watermarking and deepfake detection methods require full white-box access to target model parameters and gradients during trigger construction or verification. When proprietary models are deployed solely behind inference APIs or distributed as encrypted binaries, these techniques cannot be applied at all. This leaves model owners and auditors unable to verify intellectual property theft, track provenance, or detect misuse across commercial black-box deployments.
Watermarking & Deepfake Detection
Black-Box and Model-Agnostic Deepfake and Watermark Verification for Closed-Source Generative Models
Current watermark verification and deepfake detection methods require white-box access to target model parameters, intermediate layer activations, or gradient information. Because leading generative models and vision-language systems are served exclusively behind closed commercial APIs, internal inspection methods cannot be executed by downstream verifiers or auditors. Consequently, defenders cannot detect deepfakes or verify watermarks when the generating model is proprietary, unknown, or inaccessible. Removing the precondition of internal weight and activation access is necessary for auditing real-world deployed models.
Watermarking & Deepfake Detection
Black-Box and Architecture-Agnostic Watermarking and Defense for Diffusion Models
Current proactive watermarking, attribution, and anti-customization defenses structurally assume white-box access to target latent diffusion models, requiring direct inspection of latent spaces, VAE fine-tuning, or gradient backpropagation through internal U-Net attention layers. In practical deployment scenarios, defenders and copyright holders must protect content against proprietary black-box APIs, pixel-space diffusion models, or non-VAE architectures where internal activations and weights are inaccessible. Because existing methods are tethered to specific LDM preconditions, there is currently no verified mechanism to provide provenance or anti-fine-tuning protection across arbitrary or black-box diffusion pipelines.
Watermarking & Deepfake Detection
Cross-Family Generalization and Robustness of Text Watermarking Attacks and Defenses
Existing evaluations of text watermark attacks, detection, and removal methods have almost exclusively tested narrow subsets of logit-based schemes (e.g., KGW and SIR) evaluated on fixed base models like OPT-1.3B. Because no single study has evaluated these attack and defense methods across fundamentally different watermark families—including sampling-based, dynamic-hash, content-adaptive, and error-correction-augmented schemes—practitioners cannot determine whether reported evasion and detection results reflect general properties or family-specific artifacts. Consequently, deployment decisions for provenance tracking in production language models rely on unverified theoretical compatibility claims rather than measured empirical robustness.