Open Problems

Watermarking & Deepfake Detection

Black-Box Verification and Trigger Generation for Model Watermarking and Detection

Barrier to removePartly addressed
Possible candidate · 2/5 runs8 papers report this38% from 2025+

Generated automatically from the limitations stated in 8 papers (ICML, CVPR, ICCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current watermarking and deepfake detection methods require full white-box access to target model parameters and gradients during trigger construction or verification. When proprietary models are deployed solely behind inference APIs or distributed as encrypted binaries, these techniques cannot be applied at all. This leaves model owners and auditors unable to verify intellectual property theft, track provenance, or detect misuse across commercial black-box deployments.

Why it matters

Model owners and regulators can audit, verify intellectual property infringement, and detect watermarks in commercial models accessible solely through query APIs without cooperation from model hosting providers.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Zeroth-order gradient estimation and query-based optimization: Construct watermarking triggers or verification probes using black-box output scores or generated samples, measuring detection accuracy and query efficiency against commercial API endpoints.

  2. 2

    Black-box surrogate distillation: Train local white-box surrogate models on input-output queries to optimize transferrable verification triggers, measuring trigger transfer success rates across diverse target model architectures.

  3. 3

    Feature-level query-response fingerprinting: Design input perturbations that elicit statistically distinct behavioral signatures from suspect black-box models without parameter access, evaluating false positive rates across independent non-watermarked models.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Black-box query costs and rate limits on commercial APIs may render query-based trigger optimization intractable, or API output defenses (such as output filtering and noise perturbation) may eliminate watermark signatures without parameter access.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Watermarking & Deepfake Detection

Barrier to removeOpen

Black-Box and Model-Agnostic Deepfake and Watermark Verification for Closed-Source Generative Models

Current watermark verification and deepfake detection methods require white-box access to target model parameters, intermediate layer activations, or gradient information. Because leading generative models and vision-language systems are served exclusively behind closed commercial APIs, internal inspection methods cannot be executed by downstream verifiers or auditors. Consequently, defenders cannot detect deepfakes or verify watermarks when the generating model is proprietary, unknown, or inaccessible. Removing the precondition of internal weight and activation access is necessary for auditing real-world deployed models.

Strong candidate · 4/5 runs4 papers report this100% from 2025+

Watermarking & Deepfake Detection

Barrier to removePartly addressed

Black-Box and Architecture-Agnostic Watermarking and Defense for Diffusion Models

Current proactive watermarking, attribution, and anti-customization defenses structurally assume white-box access to target latent diffusion models, requiring direct inspection of latent spaces, VAE fine-tuning, or gradient backpropagation through internal U-Net attention layers. In practical deployment scenarios, defenders and copyright holders must protect content against proprietary black-box APIs, pixel-space diffusion models, or non-VAE architectures where internal activations and weights are inaccessible. Because existing methods are tethered to specific LDM preconditions, there is currently no verified mechanism to provide provenance or anti-fine-tuning protection across arbitrary or black-box diffusion pipelines.

Possible candidate · 3/5 runs4 papers report this100% from 2025+

Watermarking & Deepfake Detection

Scope to testPartly addressed

Cross-Family Generalization and Robustness of Text Watermarking Attacks and Defenses

Existing evaluations of text watermark attacks, detection, and removal methods have almost exclusively tested narrow subsets of logit-based schemes (e.g., KGW and SIR) evaluated on fixed base models like OPT-1.3B. Because no single study has evaluated these attack and defense methods across fundamentally different watermark families—including sampling-based, dynamic-hash, content-adaptive, and error-correction-augmented schemes—practitioners cannot determine whether reported evasion and detection results reflect general properties or family-specific artifacts. Consequently, deployment decisions for provenance tracking in production language models rely on unverified theoretical compatibility claims rather than measured empirical robustness.

Possible candidate · 3/5 runs4 papers report this75% from 2025+

Watermarking & Deepfake Detection

Scope to testOpen

Benchmarking Watermarking and Deepfake Detection Across Non-Latent and Flow-Matching Generative Architectures

Current watermarking and detection methods in generative media are almost exclusively developed and evaluated on standard latent diffusion models relying on Gaussian noise sampling. Consequently, it is unknown whether these techniques transfer, degrade, or fail entirely when applied to pixel-space diffusion, non-iterative architectures, or modern rectified-flow and flow-matching models such as SD3 and FLUX. As production generative pipelines shift away from standard latent diffusion, safety and provenance mechanisms risk operating under untested assumptions. A systematic evaluation across these architectural families is necessary to establish the empirical boundary of existing detection and watermarking schemes.

Strong candidate · 5/5 runs3 papers report this67% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.