CSPaperby Scholar7

Open ProblemsβetaIdea CheckβetaAgentic ReviewReview RankingCorrectness CheckCode CheckReference CheckRebutlyβeta

OpenPrint

API Platform
BenchmarkArticlesnew
Sign In

Open Problems

178 research problems, each a limitation that at least three papers independently report. They come from an automated analysis of 128,730 papers at venues such as NeurIPS, ICML, ICLR, CVPR and ACL, not from papers submitted to CSPaper. Treat each one as a lead to investigate. Updated Sep 8, 2026. Data: priorwork.fyi

178 problems

Model Safety, Alignment & Jailbreaks

Barrier to removeOpen

A Framework for Deploying Activation- and Logit-Based Methods Against Closed-Source LLM APIs via Instrumented Proxy Models

A very large body of techniques — safety defenses, interpretability probes, evaluation metrics, decoding controls — requires hidden states, attention maps, gradients, or token-level probabilities, and is therefore silently restricted to open-weight models. As a result, published methods are never validated on, and cannot protect or audit, the models most people actually use (GPT-4, Claude, Gemini). This is the single most common self-reported scope limitation in the literature, yet each paper treats it as an isolated footnote rather than a solvable engineering and inference problem.

Strong candidate · 5/5 runs177 papers report this78% from 2025+

Continual Learning & Catastrophic Forgetting

Scope to testPartly addressed

A Standardized Transfer-Validation Benchmark for Methods Validated Only on Small-Scale Image Classification

Nearly every training-time method in the modern literature—continual learning, distillation, unlearning, robustness, federated aggregation—is validated exclusively on CIFAR/ImageNet-subset image classification with ResNet-18/50-class backbones, and each paper explicitly defers scale, architecture, and modality transfer to future work. Because every paper defers this simultaneously, the field has no evidence base at all for whether any of these methods survive the shift to transformers, higher resolutions, dense tasks, or non-vision data; practitioners choosing between methods are choosing based on measurements made in a regime none of them deploy in. The barrier is not one method's limitation but a collective, unpriced evaluation debt: nobody can tell which small-scale conclusions are regime-robust and which are artifacts of ResNet-18 on 32×32 images.

Possible candidate · 3/5 runs152 papers report this44% from 2025+

Multilingual & Cross-Lingual Transfer

Scope to testPartly addressed

A Resource-Floor Benchmark: Measuring the Minimum Data, Lexical, and Pretrained-Model Prerequisites Below Which Multilingual NLP Methods Fail

Across transfer, adaptation, evaluation, and data-construction work, methods repeatedly assume some minimal target-language resource — unlabeled text, a bilingual dictionary, a working MT system, fastText embeddings, a pretrained multilingual encoder, or native-speaker references — and each assumption silently excludes the world's least-served languages. Because every paper states this as a single-sentence caveat, nobody knows which prerequisite is binding, how much of it is actually needed, or where the failure floor lies. The consequence is a systematic blind spot: methods are reported as "multilingual" while being untested, and likely unusable, for languages with only tens of thousands of sentences, no dictionary, and no MT coverage. A quantified floor would turn these caveats into design constraints and tell funders and corpus-builders exactly which resources to create first.

Strong candidate · 5/5 runs119 papers report this54% from 2025+

Learning with Noisy & Partial Labels

Barrier to removePartly addressed

Eliminating the Auxiliary Clean-Label Anchor Required to Bootstrap Training under Imperfect Supervision

Across weak-supervision methods — noisy labels, partial labels, single positives, model merging, node classification — the same structural dependency recurs: the method only works if some external supervised anchor exists, whether a verified clean reference set, a small labeled validation split, class-frequency priors from a labeled pool, or a warm-up model trained on trustworthy data. In the regime where imperfect labeling actually occurs (annotation too expensive or impossible), that anchor is exactly what is missing, so these methods cannot be deployed there at all. The result is a circular dependency: robust learning from weak labels presupposes the labels one is trying to avoid collecting. Today there is no principled way to replace the anchor with signals derivable from the weak data itself.

Strong candidate · 4/5 runs63 papers report this35% from 2025+

Diffusion Models

Scope to testOpen

Extending Class-Conditional ImageNet Generation Methods to Text-Conditioned, Open-Domain Synthesis

Nearly every method validated on class-conditional ImageNet 256×256 stalls at that benchmark, and it is unknown whether any of them survive contact with natural language conditioning and open-caption distributions. Class labels provide weak, discrete conditioning with 1,000 modes; text requires cross-attention or caption-token conditioning, resolution diversity, and open-vocabulary semantics. Until the class-to-text transition is demonstrated, the large literature of ImageNet-benchmarked generative advances cannot be assumed to transfer to practical text-to-image systems.

Possible candidate · 3/5 runs53 papers report this83% from 2025+

Learning with Noisy & Partial Labels

Effect to explainPartly addressed

Validation-Free Hyperparameter Selection for Robust Learning Pipelines

Across these works, the reported gains hinge on hyperparameters whose optima shift per dataset, noise level, class prior, or architecture, and which are chosen by grid search on held-out clean validation data or expensive ablation. This blocks plug-and-play deployment: a practitioner cannot know whether a method's claimed advantage survives at untuned settings, and comparisons between methods are confounded by unequal tuning budgets. Until hyperparameter settings can be selected from the training data itself (or shown to be inert), robust-learning methods remain fragile recipes rather than tools.

Strong candidate · 4/5 runs49 papers report this39% from 2025+

Hand-Object Interaction & Affordance

Barrier to removeOpen

Robust Hand-Object Interaction Learning Beyond Cascaded Parametric Pose Estimators

Downstream hand-object interaction (HOI) and affordance models are universally bottlenecked by a strict reliance on upstream off-the-shelf 3D hand pose estimators (e.g., FrankMocap, HaMeR, MANO parameterizations) and bounding-box detectors. When these upstream tools encounter severe object occlusions, fast manipulation dynamics, or egocentric perspective distortion, their estimation errors cascade directly into downstream HOI training, causing catastrophic mis-projections and synthetic contact artifacts. Because downstream methods require clean parametric 3D meshes as preconditions, massive repositories of in-the-wild manipulation videos cannot be utilized without fragile filtering heuristics that discard over half the data.

Possible candidate · 2/5 runs42 papers report this45% from 2025+

3D Point Clouds & LiDAR Perception

Scope to testOpen

An Outdoor-Scale Benchmark for 3D Scene Understanding Methods Currently Validated Only on Indoor ScanNet-Derived Data

Nearly the entire 3D perception, scene reconstruction, grounding, and question-answering stack is trained and evaluated on ScanNet-family indoor scans, so claims of scene understanding hold only within bounded, static, room-scale environments with known camera poses. As a result, no one can tell whether reported architectural advances—reasoning modules, open-vocabulary grounding, generalization arguments—transfer to outdoor, unbounded, or unsegmented scenes at all, or whether they simply exploit ScanNet's statistics. Progress reports that look cumulative within this benchmark family may be incommensurable with outdoor work that cannot reuse the same evaluation. Any paper today claiming 3D scene understanding is doing so on a single environmental regime.

Possible candidate · 3/5 runs39 papers report this30% from 2025+

Bandits & Online Learning

Scope to testOpen

A Standard Benchmark Suite and Evaluation Protocol for Validating Bandit, Online Learning, and Sequential Decision Algorithms on Real-World Logged Data

Across dozens of recent papers in sequential decision making, the empirical section is a small synthetic simulation with hand-picked parameters, and authors explicitly acknowledge that real-world validation is missing. The consequence is that theoretical guarantees and synthetic results cannot be translated into practice: practitioners cannot tell which algorithms hold up on logged production data, and researchers cannot compare against baselines because no shared real-data evaluation standard exists. Structurally, the field lacks accessible, well-curated real-world logged datasets (recommendation, clinical, resource-allocation, contract design) with standardized protocols, so every paper improvises its own toy experiment.

Possible candidate · 3/5 runs35 papers report this50% from 2025+

Efficient Attention & State-Space Models

Scope to testOpen

A Matched-Scale Testbed and Extrapolation Protocol for Testing Whether Toy-Scale Evidence Survives at Realistic Scale

Across theory-driven method papers, the empirical section reliably ends where the claim begins: guarantees about scaling, efficiency, or sample complexity are validated only on instances orders of magnitude below the regime where they matter, because realistic-scale runs are unaffordable for a single lab and no shared, property-matched task suites exist for buying at-scale evidence cheaply. The consequence is that scalability claims are effectively unfalsifiable — reviewers can neither confirm nor reject them, and methods whose advantages appear only above ~10⁸ parameters or on real data are adopted or discarded on faith. Because every group builds its own private toys, no evidence accumulates: the field cannot answer the basic question of how often toy-scale method rankings survive at realistic scale, so nobody knows what the current synthetic-only evidence base is actually worth.

Possible candidate · 2/5 runs32 papers report this56% from 2025+

Camera Pose, SLAM & Structure-from-Motion

Barrier to removePartly addressed

Reducing Reliance on Pre-Built 3D Supervision for Training 3D-Aware Perception and Reconstruction Models

Nearly every system that connects 2D images to 3D representations — whether for localization, reconstruction, or generation — is trained against 3D artifacts that must be manufactured in advance: dense depth, camera poses, meshes, point clouds, scene graphs, or splat reconstructions from dense multi-view rigs. This makes training data a bottleneck that caps scalability to ordinary single-view RGB collections, blocks application to domains without depth sensors or calibrated rigs, and drives enormous compute costs for pseudo-ground-truth generation. Until models can learn the 2D-to-3D link without manufactured 3D supervision, the field's coverage is limited to the small slice of the visual world that has been 3D-digitized.

Possible candidate · 3/4 runs24 papers report this50% from 2025+

Learning with Noisy & Partial Labels

Effect to explainOpen

Hyperparameter-Free Generalization: Benchmarking and Predicting Sensitivity of New Hyperparameters Across Tasks

Nearly every new method in this space ships with two to five new hyperparameters whose optimal values are dataset- or architecture-dependent, and the reported results are only achievable after per-task grid search on held-out validation data. This inflates effective cost, undermines plug-and-play claims, and silently biases the literature toward methods whose authors had tuning budget. Worse, when validation data is scarce (OOD, few-shot, cross-domain), the reported optima may not exist at transfer time at all — yet no one measures how much of a method's claimed gain survives under zero-tuning conditions.

Possible candidate · 1/2 runs23 papers report this27% from 2025+

Machine Unlearning

Scope to testOpen

A Standardized Multi-Architecture Testbed of Memorized Concepts, Generation Fingerprints, and Editable Structures in Text-to-Image Models Beyond Stable Diffusion v1.x

Nearly every quantitative claim about text-to-image models — memorization, unlearning success, watermark robustness, deepfake provenance, concept editability — is anchored to a handful of frozen Stable Diffusion v1.x/v2 checkpoints, because those are the only models with established benchmarks, known memorized prompts, and validated concept labels. Newer architectures (DiT-based diffusion, DeepFloyd, autoregressive and GAN generators, video models) are largely untested, so results that appear robust may be artifacts of one 2022 model family. The field cannot tell whether its findings are properties of generative modeling or properties of Stable Diffusion 1.4.

Strong candidate · 5/5 runs22 papers report this88% from 2025+

Continual Learning & Catastrophic Forgetting

Scope to testOpen

Characterizing and Reducing Pre-Trained ViT Dependence in Transfer-Based Methods

A wide range of methods built on frozen ImageNet-pretrained ViT features—whether for adapting to new tasks, constraining representations, or detecting anomalies—inherit an unexamined dependency: their guarantees hold only when the backbone's feature space is already good. Today nobody knows whether these methods degrade gracefully with weaker pretraining, generalize to domains far from ImageNet, or transfer to non-transformer architectures, because the dependence is never ablated. The consequence is that reported gains may be properties of the backbone, not of the methods themselves, and the methods are silently inapplicable wherever a strong ImageNet ViT does not exist.

Possible candidate · 3/5 runs20 papers report this78% from 2025+

3D Point Clouds & LiDAR Perception

Scope to testOpen

Benchmarking and Adapting Indoor-Trained 3D Perception and Generation Models to Outdoor and LiDAR Point Cloud Scenes

Nearly every 3D foundation-ish model — for detection, segmentation, registration, scene generation, and novel view synthesis — is trained and validated exclusively on indoor RGB-D or synthetic benchmarks (ScanNet, SUN RGB-D, 3D-FRONT, Habitat simulators), with explicit author acknowledgments that outdoor and LiDAR data is untested. This means the field's core 3D representations are calibrated to room-scale geometry (bounded depth, dense views, flat priors like indoor depth predictors and category vocabularies), and nobody can currently deploy them on autonomous-driving or aerial scans without unquantified degradation. The consequence is a silent incompatibility: outdoor researchers must either rebuild models from scratch or trust unmeasured transfer.

Possible candidate · 3/5 runs20 papers report this53% from 2025+

Information Extraction & Entity Recognition

Barrier to removeOpen

Error-Aware Entity Linking Evaluation and Propagation-Bounded Pipelines for Downstream NLP Tasks

Virtually every system that touches entities — relation extraction, KBQA, semantic parsing, coreference, retrieval-augmented generation — treats entity linking as a solved preprocessing step and evaluates against gold links, so the systems' true field performance is unknown and unlinkable to their stated failure modes. When linking is wrong, downstream modules cannot detect, recover from, or bound the damage; reported failure analyses put linking errors at 15–35% of end-to-end errors. Today nobody can quantify how much of a downstream system's headroom is actually upstream, because no benchmark couples controlled link corruption with propagation measurement.

Strong candidate · 4/5 runs19 papers report this25% from 2025+

Time-Series Forecasting

Barrier to removeOpen

A Standard Benchmark and Protocol for Genuine Zero-Shot Transfer to Truly Unseen Domains in Structured Perception Tasks

Current zero-shot claims in structured prediction are not actually zero-shot: evaluations use datasets sharing variable structure, sensor type, or category overlap with training data; some protocols leak a fraction of test samples or target scenes to select prompts or instructions; and others substitute qualitative visualizations for scores because no clean mapping exists. The result is that reported "zero-shot" numbers cannot be compared across papers, and the field cannot measure how far true transfer to genuinely unseen conditions actually is. Until evaluation is tightened, progress on cross-domain generalization is unmeasurable and likely overstated.

Possible candidate · 2/5 runs19 papers report this57% from 2025+

Protein & Biomolecular Sequence Modeling

Scope to testPartly addressed

Benchmarking the Fidelity of In Silico Protein Design Proxies Against Experimental Biophysical Assays

Generative biomolecular modeling currently assesses sequence designs almost entirely through computational surrogates—such as AlphaFold confidence (pLDDT/pTM), ESMFold self-consistency, and Rosetta energy terms—rather than physical assay measurements. Because these proxies serve as both optimization objectives and evaluation metrics, models risk overfitting to artifacts of structure-prediction heuristics rather than true biophysical viability. Without systematic evaluation against physical assays across sequence design methods, it remains unknown whether improvements on in silico leaderboards translate to viable expression, foldability, or binding in physical systems.

Possible candidate · 2/5 runs19 papers report this95% from 2025+

Low-Resource Languages & Dialects

Effect to explainOpen

Reliability-Calibrated Annotation Protocols for Subjective and Low-Resource labeling Tasks

Across many NLP annotation efforts, gold standards rest on so few annotators—and on tasks so subjective—that reported inter-annotator agreement is often moderate at best (κ 0.18–0.46), and in some cases reliability fails outright (agreement too weak to include in evaluation at all). Because papers rarely distinguish "the task is ambiguous" from "the pool was too small," downstream users cannot judge how trustworthy the gold labels are, and resources with single-annotator or single-community provenance silently embed one worldview. Structurally, there is no standard way to report, audit, or correct for annotation reliability when it is low or unmeasured.

Possible candidate · 2/5 runs19 papers report this60% from 2025+

Model Safety, Alignment & Jailbreaks

Barrier to removePartly addressed

Decision-Only and Logit-Free Safety Auditing and Alignment Evaluation for Black-Box Models

Current safety auditing, jailbreak discovery, and representation-based alignment diagnostics overwhelmingly rely on white-box preconditions such as gradient access, intermediate activations, or output token logit distributions. When models are deployed behind production APIs that return only generated text, practitioners cannot run these diagnostic and red-teaming pipelines directly. Relying on surrogate transfer from open-source models produces high false-negative rates due to incomplete transferability across distinct architectures and tokenizers. Consequently, safety assessments for proprietary, API-governed models remain structurally disconnected from the methods developed in the literature.

Strong candidate · 4/5 runs17 papers report this59% from 2025+

Model Safety, Alignment & Jailbreaks

Barrier to removePartly addressed

Recovering Internal Attention and Activation Signals from API-Visible Model Behavior

The shared wall is a dependence on inference-time extraction of layer-wise internal states — attention weights, hidden activations, per-layer execution — which are unavailable outside open-weight deployments and costly and per-model-fragile even where weights are open. Every technique built on those signals stops working the moment the model is served through an API; the sanctioned fallbacks, transferring from an open proxy or degrading to logit-only features, demonstrably forfeit most of the signal (one published detector drops from its white-box ceiling to 0.66 AUROC). Even with open weights, each release re-breaks the tooling: layer indices must be re-selected per model and multi-pass decoding adds latency. The result is a method family whose reach shrinks exactly as the most capable models become less open, and whose findings nobody without weight access can verify or use.

Strong candidate · 5/5 runs15 papers report this100% from 2025+

Text-to-Image & Video Generation/Editing

Barrier to removePartly addressed

Systematically Benchmarking and Calibrating VLM-as-a-Judge Failure Modes in Image and Video Generation

Current text-to-image and video generation benchmarks overwhelmingly rely on Vision-Language Models (VLMs) like GPT-4o, Gemini, and Qwen-VL as automated judges for win rates and task success metrics. However, these VLM judges systematically exhibit low correlation with human raters on fine-grained visual details, suffer from position and same-family evaluator biases, and under-penalize visual artifacts such as boundary glitches or text corruption. Consequently, benchmark progress across generative models is structurally uninterpretable, as claimed state-of-the-art gains may reflect evaluator artifacts rather than genuine improvements in visual quality and prompt fidelity.

Possible candidate · 2/5 runs15 papers report this100% from 2025+

Uncertainty Quantification & Conformal Prediction

Scope to testOpen

Principled Multiclass Extensions of Binary-Only Theoretical Guarantees in Prediction Methodology

A large body of theory and tooling — covering uncertainty quantification, conformal methods, and attribution-style analyses — is derived and validated only for binary or single-label settings, with multiclass, multi-label, and open-set cases deferred to "future work" or handled via ad hoc one-versus-rest surrogates whose guarantees and sharpness are not established. As a result, practitioners applying these techniques to realistic multi-class problems either inherit unquantified approximation error or cannot use the methods at all (e.g., when only discrete predictions rather than continuous outputs are available). Each new paper re-derives binary results and re-defers the extension, so the field lacks a general recipe for lifting binary guarantees to K-class settings with provable validity and measured sharpness.

Strong candidate · 4/5 runs14 papers report this30% from 2025+

Model Safety, Alignment & Jailbreaks

Effect to explainPartly addressed

Robust Evaluation of Non-Binary and Phrasing-Agnostic Refusal in Safety-Aligned Language Models

Current safety benchmarks and verifiable reward mechanisms evaluate model refusals primarily through binary keyword matching, hand-crafted refusal templates, or uncalibrated model-based autoraters (which show inter-annotator agreement as low as $\alpha = 0.378$). This creates a blind spot where models that partially comply, use non-standard refusal vocabulary, or alter their stance under multi-turn pushback are misclassified as either fully compliant or safely refusing. Consequently, alignment interventions optimized against these brittle metrics suffer from artificial over-refusal on benign queries while leaking safety risks through stylistic variation and conversational pressure. Without a rigorous, non-binary evaluation standard across diverse phrasing distributions, researchers cannot reliably assess whether safety alignment generalizes beyond rigid heuristic templates.

Strong candidate · 4/5 runs14 papers report this71% from 2025+

Molecular Modeling & Drug Design

Scope to testOpen

Benchmarks and Training Protocols for Size-and-Structure Transfer in Molecular ML

Molecular machine-learning models are trained and evaluated within narrow slices of chemical space — small, drug-like, scaffold-homogeneous molecules averaging ~44 atoms — so any deployment beyond that slice degrades sharply or requires full retraining per system. This blocks the field's central promise: screening large protein-ligand complexes, extrapolating to novel chemistries, and reusing models across temperatures, states, or unseen systems. Without a measurement standard for out-of-distribution molecular generalization, papers cannot even quantify how far their models actually transfer.

Possible candidate · 3/5 runs14 papers report this75% from 2025+

3D Point Clouds & LiDAR Perception

Barrier to removeOpen

Robust Multimodal 3D Perception Under Sensor Extrinsic and Intrinsic Miscalibration

State-of-the-art multimodal 3D perception architectures strictly require accurate extrinsic and intrinsic calibration between LiDAR and camera frames to construct point-to-pixel projections, frustum cost volumes, and cross-attention neighbourhoods. In real-world deployments, mechanical vibration, thermal expansion, and mounting shifts cause calibration parameters to drift, leading to severe geometric misalignment. Because current multimodal models hardcode these spatial transformations into their feature lifting layers, minor calibration errors cause catastrophic performance degradation or complete pipeline failure.

Possible candidate · 2/5 runs13 papers report this38% from 2025+

Learning with Noisy & Partial Labels

Effect to explainPartly addressed

Robust Learning under Incomplete, Evolving, and Annotator-Heterogeneous Label Spaces

Nearly every label-learning method in circulation assumes a closed world: the class set is known a priori, fixed across training and deployment, fully covered by the annotation, and interpreted identically by every annotator. When any of these fail — a class missing from the taxonomy, new classes arriving in a stream, annotators with different accuracies or selection biases, or legacy datasets with only collapsed majority labels — the methods either fail silently or degrade without diagnostic. Today there is no principled way to even measure how much of a model's error comes from label-space misspecification versus model capacity, so practitioners cannot tell whether to fix the taxonomy, re-annotate, or retrain.

Possible candidate · 2/5 runs13 papers report this33% from 2025+

Deep Learning Theory & Optimization Dynamics

Barrier to removePartly addressed

Bridging the Activation Divide in Neural Network Optimization Dynamics Theory

Current theoretical analyses of neural network training dynamics are split across incompatible mathematical preconditions: one major body of literature strictly requires high-order smoothness ($C^2$, $C^\infty$, bounded derivatives, or odd symmetry), which explicitly excludes standard piecewise-linear activations like ReLU, while another body relies strictly on piecewise linearity and positive homogeneity, failing to transfer to smooth activations like GELU, SiLU, or tanh. As a result, foundational theoretical predictions—such as edge-of-stability behavior, width-independent convergence rates, and representation learning bounds—cannot be generalized across activation families. Without analytical tools that either handle subgradient boundary crossings or bound the divergence between smooth approximations and exact piecewise-linear trajectories during training, optimization theory remains fragmented into mutually exclusive activation regimes.

Possible candidate · 2/5 runs13 papers report this38% from 2025+

Quantization & Low-Precision Training

Effect to explainOpen

Characterizing and Mitigating Low-Bit Quantization Breakdown in Small-Scale Language Models

Post-training quantization consistently fails in the sub-3-bit and 2-bit regimes, suffering catastrophic perplexity degradation that disproportionately impacts small and mid-sized models (e.g., OPT-125M, LLaMA-7B/8B). Because smaller base models have less parameter redundancy, aggressive bit-width reductions push perplexity beyond functional utility (often exceeding 100+ PPL), preventing local deployment of compact models on edge hardware. Today, practitioners are forced to choose between unquantized models that exceed device memory budgets or severely degraded low-bit variants that produce gibberish.

Possible candidate · 3/5 runs12 papers report this58% from 2025+

Text-to-Image & Video Generation/Editing

Barrier to removePartly addressed

Benchmarking and Mitigating Upstream LLM Prompting Dependencies in Text-to-Visual Generation

Modern text-to-image and text-to-video generation and editing methods rely heavily on external large language models (LLMs) for prompt rewriting, sub-prompt decomposition, and attribute filtering. Consequently, visual generation quality is bottlenecked by upstream LLM biases, template artifacts, and hallucinated attributes, frequently failing on ambiguous or in-the-wild user instructions unless manually tuned. Because no systematic study has evaluated visual generation frameworks across diverse upstream prompt generators and decomposition strategies, the sensitivity and failure modes of these pipelines remain unquantified.

Possible candidate · 2/5 runs12 papers report this58% from 2025+

Self-Supervised & Contrastive Representation Learning

Effect to explainOpen

A Scaling-Law Characterization of Regularization and Transfer Benefits That Vanish at Scale

Many published training-time interventions show solid gains on small models, small batches, and small datasets, yet shrink toward zero as any of these scale up. Because each paper reports only its own endpoints, the community cannot tell whether a given benefit decays with scale or was an artifact of the small regime — leading to wasted adoption and redundant re-testing. A controlled characterization of this decay, with predictors of when a benefit will survive, is what is currently missing.

Strong candidate · 4/5 runs11 papers report this33% from 2025+

Digital Humans, Avatars & Virtual Try-On

UnclassifiedPartly addressed

Overcoming Parametric 3DMM Expressiveness and Mouth Interior Bottlenecks in Monocular Facial Avatars

Neural head avatar pipelines universally rely on parametric mesh models like FLAME as geometric proxies, inheriting their topological limitations. Because FLAME lacks internal oral anatomy (teeth and tongue) and fine-grained geometric detail, avatars suffer from severe projection artifacts such as teeth textures baked directly onto lip surfaces and distorted speech articulation. Furthermore, low-dimensional linear expression spaces cap the capture of dynamic micro-expressions, dynamic wrinkles, and non-parametric regions like hair.

Possible candidate · 3/5 runs11 papers report this91% from 2025+

3D Point Clouds & LiDAR Perception

Barrier to removePartly addressed

Benchmark and Training Protocols for 3D Perception from Unpaired or Calibration-Free 2D-3D Data

Nearly every current 2D-3D fusion method assumes registered multi-view imagery, known intrinsics/extrinsics, and per-frame paired data — assumptions that hold only in curated lab captures. The consequence is that point clouds collected standalone (LiDAR scans, crowdsourced reconstructions, legacy datasets) cannot benefit from image features at all, and methods cannot be deployed where a fixed calibrated rig with synchronized cameras is impractical. Progress is blocked because no established protocol or dataset exists for learning 2D-3D correspondences from unpaired, uncalibrated, or single-view inputs.

Possible candidate · 2/5 runs11 papers report this17% from 2025+

Algorithmic Fairness

Scope to testPartly addressed

A Unified Formulation for Multiclass and Regression Targets under Postprocessing-Based Guarantees

Ten independent lines of work — spanning fairness postprocessing, conformal prediction, and related certification-style methods — each stop at binary or, at best, categorical classification targets. Continuous outcomes and structured prediction are simply out of reach with current formulations, which means a large fraction of real prediction problems (prices, dosages, rankings, likelihoods) get either no guarantee at all or a crude discretization that distorts the guarantee. Practitioners must choose between losing the guarantee and losing the task, and every paper defers the same extension to "future work" rather than solving it jointly.

Strong candidate · 5/5 runs10 papers report this83% from 2025+

Syntactic & Semantic Parsing

Barrier to removeOpen

Benchmarking the Robustness of Syntax-Guided NLP Across Parser Degradation, Languages, and Formalisms

Syntax-guided models and parser-assisted NLP pipelines rely heavily on upstream parsers, yet their evaluations remain largely confined to high-resource English benchmarks using high-quality silver or gold parses. Current literature explicitly notes that methods have not been evaluated when parsers degrade under domain shift, when deployed on low-resource languages lacking specialized parsers, or when adapted across differing formalisms (e.g., constituency versus dependency trees). Consequently, practitioners cannot determine whether syntax-guided gains persist under realistic, noisy, or cross-lingual parsing conditions, or if parser error propagation negates structural benefits entirely.

Strong candidate · 5/5 runs10 papers report this20% from 2025+

Domain Adaptation & Generalization

Barrier to removePartly addressed

Domain-Invariant Learning Without Domain Labels or Target-Domain Data at Training Time

Current methods across these settings all condition on information that disappears in deployment: explicit domain labels for training samples, knowledge of the domain count, or a pool of unlabeled target-domain data available during training. When source domain identity is unknown, only one source is available, or target data cannot be collected before the model is fixed, these methods are simply unusable — not merely degraded. This blocks the realistic scenario where a single labeled dataset is all one has, and the test distribution is only discovered at inference time.

Strong candidate · 4/5 runs10 papers report this57% from 2025+

RL Theory & Markov Decision Processes

Scope to testOpen

Empirical Robustness and Scalability Evaluation of Provably Efficient RL Algorithms

A wide range of theoretical reinforcement learning algorithms offer provable sample-complexity guarantees, yet their empirical validation is almost exclusively restricted to synthetic tabular MDPs with fewer than 30 states. Because these methods have not been evaluated across larger, standardized, or higher-dimensional environments, it remains unknown whether their theoretical efficiency properties hold in practice. Without systematic cross-domain testing, the community lacks evidence on whether provably efficient exploration and representation mechanisms scale beyond toy sandbox problems.

Possible candidate · 3/5 runs10 papers report this30% from 2025+

Offline & Model-Based RL

Scope to testOpen

Cross-Modality and Sim-to-Real Robustness Evaluation of Offline and Model-Based Reinforcement Learning

Current offline and model-based reinforcement learning algorithms are almost exclusively benchmarked on state-observation continuous-control locomotion tasks in simulation (e.g., standard MuJoCo suites). Because papers routinely restrict evaluations to low-dimensional proprioceptive states, the performance, sample efficiency, and stability of these algorithms remains unknown when applied to visual observations, discrete action spaces, manipulation tasks, or physical platforms. Consequently, practitioners cannot determine whether published algorithmic gains reflect genuine advancements in decision-making or overfitting to low-dimensional simulation dynamics.

Possible candidate · 3/5 runs10 papers report this50% from 2025+

Causal Inference & Discovery

Scope to testPartly addressed

Benchmarking Causal Effect Estimation Across Multi-Valued and Continuous Treatment Regimes

The vast majority of causal inference and discovery algorithms are formulated and evaluated strictly under binary treatment assumptions ($T \in \{0, 1\}$). In real-world domains such as medicine and econometrics, interventions are routinely dose-dependent, multi-categorical, or continuous, rendering binary-only methods inapplicable without ad-hoc binarization. Because existing methods have not been systematically evaluated across non-binary treatment spaces, practitioners lack empirical evidence on how severely estimation and discovery performance degrade when moving beyond the binary regime.

Possible candidate · 3/5 runs10 papers report this60% from 2025+

Federated Learning

Effect to explainOpen

Client-Scalable Federated Learning Under Severe Data Partitioning and Heterogeneity

Federated learning algorithms consistently exhibit severe accuracy collapses when the client population scales from small toy cohorts (5–10 clients) to realistic deployments (30–100+ clients). As the client count increases for a fixed dataset, each client receives fewer training samples while cross-client data heterogeneity and gradient conflict intensify, causing standard aggregation operators to produce destructive model updates. This scaling wall structurally limits federated systems from deploying in environments where data is naturally fragmented across large numbers of low-volume edge devices. Current literature masks this limitation by evaluating primarily on small, well-provisioned client counts where local sample sizes remain artificially high.

Strong candidate · 5/5 runs9 papers report this67% from 2025+

Language Model Interpretability & In-Context Learning

Scope to testOpen

Cross-Linguistic Robustness of Interpretability and Cognitive Alignment in Language Models

Mechanistic interpretability claims and brain-to-LM alignment results currently rest almost entirely on English-only stimuli, English-trained models, and English-speaking participant data. Because cross-linguistic evaluation has not been conducted across these methods, it remains unknown whether identified circuits, induction heads, and representational alignments reflect general linguistic processing or are artifacts of English syntax and data abundance. Consequently, researchers cannot rely on existing interpretability mechanisms when analyzing multilingual models or studying cognitive language processing across typologically diverse languages.

Strong candidate · 5/5 runs9 papers report this33% from 2025+

Semantic & Instance Segmentation

Barrier to removeOpen

Overcoming the Mask Proposal Recall Ceiling in Open-Vocabulary Segmentation

Current open-vocabulary and weakly supervised segmentation pipelines decouple class-agnostic mask proposal generation from segment classification, freezing an upstream proposal model like SAM, FreeSOLO, or Mask2Former. Because the downstream classifier operates solely on candidate masks, segmentation accuracy on novel classes and sub-object regions is strictly bounded by the proposal generator's recall, creating a ceiling that downstream improvements cannot fix. Failures to isolate novel or fine-grained instances at the proposal stage permanently propagate through the pipeline without any mechanism for recovery.

Strong candidate · 4/5 runs9 papers report this22% from 2025+

Protein & Biomolecular Sequence Modeling

Scope to testOpen

A Systematic Robustness Study of Genomic Foundation Models Beyond the Human Reference Genome

Current genomic sequence models are predominantly pretrained and evaluated strictly on the single human reference genome (GRCh38/hg38). Consequently, it remains unknown whether representations trained on a single reference sequence maintain their predictive capabilities across distinct chromosomes, across individual genetic variation in population-scale cohorts, or across non-human species genomes. Without cross-setting robustness benchmarks, practitioners cannot determine whether model performance transfers to non-reference genomic contexts or degrades under distribution shifts.

Possible candidate · 3/5 runs9 papers report this67% from 2025+

Hand-Object Interaction & Affordance

Effect to explainOpen

Robust 3D Hand-Object Pose and Interaction Estimation Under Upstream 2D Segmentation and Localization Noise

Current hand-object interaction and 3D hand pose estimation models rely heavily on upstream 2D bounding box crops, hand segmentation masks, or pre-extracted interaction regions. In real-world egocentric scenes, upstream 2D segmenters and detectors degrade severely due to object occlusions, motion blur, and cluttered backgrounds, causing downstream 3D estimation to fail or hallucinate. Because models are trained and benchmarked assuming near-ideal 2D crops or masks, existing pipelines cannot operate reliably in autonomous, in-the-wild video streams.

Possible candidate · 3/4 runs9 papers report this56% from 2025+

RLHF & Preference Optimization

Barrier to removePartly addressed

Quantifying Evaluation Circularity and Judge Bias in Preference Optimization

Current preference optimization and RLHF methods evaluate policy improvements almost exclusively through proprietary LLM judges (primarily GPT-4/GPT-4o) and automated reward models that often share architectures or training data with the policies being tested. This reliance introduces severe evaluation circularity and metric mismatch, where reported win-rate gains frequently reflect alignment with a specific judge's stylistic preferences rather than genuine policy improvements. Consequently, researchers cannot determine whether novel preference optimization algorithms generalize or merely exploit the biases of automated proxies.

Possible candidate · 3/5 runs9 papers report this89% from 2025+

3D Shape Modeling & Surface Reconstruction

Barrier to removeOpen

Reducing the Computational and Fidelity Cost of Distilling 2D Diffusion Priors into Explicit 3D Representations

Every pipeline that extracts 3D from pretrained 2D diffusion models currently pays a massive per-object tax: hundreds of rendered views, hundreds of diffusion evaluations, per-scene hyperparameter hand-tuning, and per-identity training — and still inherits the 2D prior's view-inconsistency and bias toward common categories. Nobody can train a single shared 3D generative model end-to-end on a shape dataset at acceptable cost, so 3D diffusion remains per-instance rather than dataset-scale. This blocks both generalization guarantees and practical deployment.

Possible candidate · 3/5 runs9 papers report this40% from 2025+

Recommender Systems

Scope to testOpen

Cross-Domain and Interaction-Density Robustness of Sequential Recommender Systems

Current sequential and language-model-based recommender algorithms are overwhelmingly developed and evaluated on 5-core filtered subsets of Amazon product reviews (e.g., Books, Electronics, Beauty). Consequently, it is unknown whether reported performance advantages hold across non-e-commerce domains (such as news browsing, streaming video, or social feeds), under varying interaction densities, or at larger catalog scales. Researchers cannot determine whether recent architectural improvements reflect general algorithmic progress or overfitting to the statistical idiosyncrasies of filtered Amazon review datasets.

Possible candidate · 2/5 runs9 papers report this67% from 2025+

Human Pose & Motion Estimation

Barrier to removePartly addressed

Robust 3D Pose Lifting Under Empirical Upstream 2D Detector Errors

Monocular 3D human pose estimators depend fundamentally on upstream 2D keypoint detectors, suffering severe accuracy degradation (e.g., MPJPE increases of 40% to over 100%) when moving from ground-truth 2D annotations to actual detector outputs. Current lifting architectures are predominantly trained on clean annotations or evaluated against simplistic Gaussian perturbation, leaving them brittle to real detector artifacts such as keypoint swaps, missing joints from occlusion, and redundant bounding boxes. Consequently, isolated benchmark gains on clean 3D datasets fail to translate into robust end-to-end motion capture in unconstrained real-world environments.

Possible candidate · 2/5 runs9 papers report this67% from 2025+

Multimodal Fusion & Multi-Task Learning

Barrier to removePartly addressed

Systematic Robustness Benchmarking of Incomplete Multimodal Learning Under Non-Random and Structured Missingness

Existing multimodal learning methods designed for missing data are predominantly evaluated under Missing at Random (MAR) or uniform independent drop assumptions, often requiring fully complete training data. In practical applications like clinical diagnostics and robotics, missingness is frequently structured, correlated across sensors, or Missing Not at Random (MNAR) due to cost or clinical decisions. Because existing literature only tests idealized missingness or uses simple mean imputation fallbacks, practitioners cannot determine whether current alignment, distillation, or generative imputation methods generalize to realistic incomplete data regimes.

Possible candidate · 2/5 runs9 papers report this67% from 2025+

Multimodal Fusion & Multi-Task Learning

Scope to testOpen

Empirical Robustness and Scalability Benchmarking for Multi-View Multimodal Fusion

Current multi-view and multimodal fusion methods are evaluated almost exclusively on small benchmarks (typically under 10,000 samples) with very few modalities or views (2 to 6 views). Because scaling behavior across large sample counts ($N \gg 10^4$) and high view counts ($V > 6$) remains untested, it is unknown whether reported algorithmic advantages persist, collapse, or become computationally intractable in larger settings. Practitioners cannot reliably select fusion algorithms for complex multi-sensor or high-view applications because no empirical baseline exists across these broader operating conditions.

Strong candidate · 5/5 runs8 papers report this62% from 2025+

Offline & Model-Based RL

Scope to testOpen

Cross-Domain Evaluation and Robustness of Offline Model-Based RL Across State-Action Representations

Offline and model-based reinforcement learning algorithms are almost universally developed, theoretically analyzed, and benchmarked exclusively on either discrete (tabular/gridworld) or continuous (continuous control) domains. Because individual methods are rarely evaluated across alternative state-action representations or stochasticity regimes, it is currently unknown whether standard conservatism mechanisms, planning rollouts, and uncertainty bounds remain robust outside their design scope. As a result, practitioners facing mixed, hybrid, or converted state-action spaces cannot anticipate failure modes or know whether existing offline RL algorithms transfer reliably.

Strong candidate · 5/5 runs8 papers report this50% from 2025+

Digital Humans, Avatars & Virtual Try-On

Barrier to removePartly addressed

Joint Pose Refinement and Error-Tolerant Conditioning for Downstream Human Avatar Synthesis

Current avatar creation and virtual try-on pipelines treat SMPL/SMPL-X estimations, camera calibrations, and segmentation masks as fixed, ground-truth preconditions. When off-the-shelf estimators inevitably produce misaligned or out-of-distribution body parameters on in-the-wild images, these errors cascade into severe geometric distortion, stitching artifacts, and generation failure. Because existing downstream pipelines do not jointly optimize human body parameters or account for parameter uncertainty, high-fidelity human generation remains restricted to controlled studio captures with pristine initializations.

Strong candidate · 5/5 runs8 papers report this25% from 2025+

Text-to-Image & Video Generation/Editing

Barrier to removePartly addressed

Dense Motion Supervision Without Optical-Flow Ground Truth

Nearly every pipeline that needs dense correspondence—video generation control, camera pose estimation, motion attribution—currently anchors its supervision or its inference to optical flow, either ground-truth annotations or outputs of a pretrained model (RAFT, AllTracker). This couples training data to static-camera, high-quality, curated footage, caps dataset size (22k–630k pairs), injects flow errors directly into downstream quality (cycle-consistency errors >1000 px), and bakes in simplifying assumptions like first-frame conditioning or pinhole intrinsics. Every consumer of flow inherits the same ceiling set by flow estimators and the data they were trained on.

Strong candidate · 5/5 runs8 papers report this100% from 2025+

Submodular Optimization & Approximation Algorithms

UnclassifiedOpen

Robustness and Generalization of Graph Submodular Optimization to Weighted, Directed, and Attributed Networks

Current submodular optimization algorithms on graphs are almost exclusively formulated and evaluated on static, unweighted, and simple undirected graphs. Practitioners working with networks containing edge capacities, asymmetric directed relations, or node attributes have no evidence regarding whether these algorithms retain their approximation quality or degrade catastrophically. Without systematic testing across these graph variants, it remains unknown which algorithms generalize gracefully and which require fundamentally new theoretical formulations.

Strong candidate · 4/5 runs8 papers report this88% from 2025+

RL Theory & Markov Decision Processes

Barrier to removeOpen

Sample-Efficient Reinforcement Learning Under Markovian Trajectory Sampling Without a Generative Model

Theoretical analysis in these MDP frameworks relies on a generative model that provides independent transition samples for arbitrary state-action queries. In real-world physical systems and streaming applications, arbitrary state resetting is impossible and data must be gathered along continuous online or episodic trajectories. As a result, existing sample complexity and convergence guarantees fail to hold under realistic Markovian data generation, leaving theoretical bounds disconnected from deployable online settings.

Strong candidate · 4/5 runs8 papers report this50% from 2025+

Inference Efficiency & Long Context

Effect to explainOpen

Systematic Cross-Scale Robustness and Evaluation of Efficient Inference and Long-Context Methods

Existing inference efficiency and long-context techniques are predominantly evaluated only on small models (1B–13B) and narrow synthetic tasks or single-pipeline setups. When applied to larger architectures (e.g., 70B) or realistic extended contexts, several methods exhibit severe accuracy regressions, unanticipated KV-cache memory overheads, or net throughput slowdowns. Because no unified evaluation exists across model scales and standardized inference runtimes, practitioners cannot predict whether reported efficiency gains will hold or catastrophically fail in real-world deployments.

Possible candidate · 3/5 runs8 papers report this75% from 2025+

Digital Humans, Avatars & Virtual Try-On

Barrier to removePartly addressed

Reducing Downstream Dependency of Garment/Avatar Generation on Upstream 2D and 3D Human-Parsing Predictions

Every pipeline in this space chains off-the-shelf estimators — pose keypoint detectors, DensePose, SMPL/SMPL-X fitting, segmentation/parsing models, SAM — and treats their outputs as ground truth. When these upstream predictions are noisy, fail on exaggerated poses, unusual garments, or garment-only imagery, the downstream model inherits or amplifies the errors (garment leaking, misinitialized garment geometry, failed reposing). No one measures or controls how error propagates through this preprocessing chain, so each new pipeline is brittle in its own undocumented way.

Possible candidate · 3/5 runs8 papers report this60% from 2025+

Low-Resource Languages & Dialects

Effect to explainPartly addressed

A Benchmark and Labeling Framework for Community-Contingent Hate Speech Annotations, Replacing Single Gold Labels with Distributions of Plausible Ones

Every current hate speech dataset collapses a genuinely plural judgment — what counts as hateful depends on annotator identity, community norms, reclamation practices, and exposure to implicit forms — into one label, then trains and evaluates models against it. This means models cannot represent the disagreement they were built from, audits of pretraining data and model outputs in low-resource settings have no trustworthy ground truth to measure against, and moderation systems calibrated on one community's thresholds misfire (or stay silent) on another's counter-speech, quotes, and reclaimed slurs. Structurally, nobody today can state, for a given utterance, how hateful-label assignments vary across the populations that matter — so "biased model" and "ambiguous data" are permanently confounded.

Possible candidate · 3/5 runs8 papers report this25% from 2025+

Retrieval-Augmented Generation

Effect to explainPartly addressed

Cross-Task Calibration and Failure Modes in RAG Hallucination Detectors

Current hallucination detection and mitigation mechanisms in retrieval-augmented generation fail to transfer across task formats: span-level detectors achieve near-zero recall on subtle conflicts, agreement modules collapse into high-abstention trivial classifiers, and score distributions require manual per-task calibration. Consequently, deploying RAG systems in heterogeneous pipelines either results in pervasive false alarms that discard valid outputs or unflagged hallucinations in grounded summarization and multi-hop QA. Without understanding the operational boundaries and calibration dynamics across varying answer structures, hallucination filters cannot reliably act as automated triggers or guardrails.

Possible candidate · 3/5 runs8 papers report this88% from 2025+

Watermarking & Deepfake Detection

Barrier to removePartly addressed

Black-Box Verification and Trigger Generation for Model Watermarking and Detection

Current watermarking and deepfake detection methods require full white-box access to target model parameters and gradients during trigger construction or verification. When proprietary models are deployed solely behind inference APIs or distributed as encrypted binaries, these techniques cannot be applied at all. This leaves model owners and auditors unable to verify intellectual property theft, track provenance, or detect misuse across commercial black-box deployments.

Possible candidate · 2/5 runs8 papers report this38% from 2025+

Deep Learning Theory & Optimization Dynamics

Scope to testPartly addressed

Extending Two-Layer Network Theory to Deep and Structured Architectures

A substantial body of rigorous results—on training dynamics, predictor statistics, generalization, and under-fitting—is proven only for one- or two-layer fully-connected networks, often with fixed or special activations. Practitioners use deep residual, convolutional, and attention models, so none of these guarantees can be checked against the systems that matter. Until the theory carries over, each paper's conclusions remain unvalidated in practice, and the specific mechanisms they identify may be artifacts of the shallow, unstructured setting.

Weak candidate · 1/4 runs8 papers report this40% from 2025+

Model Safety, Alignment & Jailbreaks

Barrier to removeOpen

Safety Monitoring and Alignment Interventions Under Black-Box and API-Only Constraints

Current state-of-the-art alignment monitoring, jailbreak detection, and activation-steering interventions explicitly require full white-box access to hidden state activations, KV caches, attention logits, or gradient flows. In practice, third-party safety auditors, downstream system builders, and end users interact with models exclusively via black-box query APIs or restricted endpoints. Because existing methods are structurally coupled to internal state inspection and modification, they cannot be deployed or evaluated on proprietary commercial models or in privacy-preserving environments.

Strong candidate · 5/5 runs7 papers report this100% from 2025+

Language Model Interpretability & In-Context Learning

Barrier to removeOpen

Benchmarking Whether LLM-Based Methods Survive a Change of Model: A Transfer Study Across Access Levels and Model Families

Methods built on one LLM — whether they depend on its logits, its instruction-following behavior, its pretraining data, or its annotation quality — are currently reported on one or two models and simply assumed to generalize. Because the assumptions are implicit, a follow-up researcher cannot tell whether a failure on a new model is a flaw of the method or an unexamined dependency on the original model's scale, data, or API surface. This blocks cumulative progress: every new model release forces methods to be silently re-validated or abandoned, and conclusions attributed to "LLMs" are really conclusions about a single backbone.

Strong candidate · 5/5 runs7 papers report this50% from 2025+

Protein & Biomolecular Sequence Modeling

Scope to testPartly addressed

Benchmarking Surrogate-Trained Sequence Optimization Against Empirical Fitness Landscapes

Current machine learning models for protein and biomolecular sequence design are evaluated almost exclusively against learned surrogate oracles (such as CNN fitness predictors) rather than verified experimental outcomes. Because black-box optimizers can easily exploit surrogate artifacts and out-of-distribution prediction errors, reported in-silico fitness gains often fail to reflect actual biological activity. Without systematic evaluation across empirically ground-truthed landscapes, it is unknown which sequence design methods reliably find functional sequences versus which merely overfit to computational evaluator flaws.

Strong candidate · 4/5 runs7 papers report this86% from 2025+

Protein & Biomolecular Sequence Modeling

Effect to explainPartly addressed

Benchmarking and Mitigating Model Sensitivity to Predicted Structure Quality in Protein Sequence Modeling

Current structure-conditioned protein sequence and representation models routinely filter out test proteins with predicted confidence scores below pLDDT 70 or assume near-experimental coordinate fidelity. Consequently, these models cannot reliably be applied to disordered proteins, orphan sequences, or hard-to-predict folds where high-confidence static predictions do not exist. Downstream pipelines experience unquantified error propagation when forced to consume predicted structures from AlphaFold2 or ESMFold, yet no systematic evaluation exists mapping how performance degrades across the full spectrum of structural prediction quality.

Possible candidate · 3/5 runs7 papers report this57% from 2025+

Language Model Evaluation & Benchmarking

Barrier to removeOpen

Methods and Benchmarks for LLM Judging, Ranking, and Evaluation That Work Without Token-Level Probabilities

A large fraction of scoring, ranking, and error-detection techniques for LLMs depend on logprobs, logits, or perplexities, which commercial APIs increasingly do not expose. This structurally blocks researchers and practitioners using closed models from applying state-of-the-art evaluation and ranking methods, forcing them onto weaker text-only alternatives or into hosting expensive open models. There is currently no systematic account of how much quality is lost when moving from probability-based to text-only judgments, nor of which probability signals can be reliably elicited from a model's own generated text.

Possible candidate · 2/5 runs7 papers report this100% from 2025+

Meta-Learning

Barrier to removeOpen

Empirical Robustness and Precondition Boundaries of Meta-Learning Under Task Distribution Shifts

Meta-learning methods fundamentally rely on the precondition that meta-training task distributions accurately reflect downstream target tasks, requiring structured multi-task splits and aligned data. When target tasks deviate from the meta-training distribution, cross-task transfer frequently fails, causing adaptation performance to degrade below standard single-task training. Because these preconditions are rarely satisfied in uncurated real-world deployments where diverse task collections are unavailable, meta-learning remains largely unusable outside narrow synthetic benchmarks.

Possible candidate · 2/5 runs7 papers report this57% from 2025+

RL Theory & Markov Decision Processes

Barrier to removeOpen

Reinforcement Learning Algorithms Without Generative Model and Reset Oracles

A wide class of theoretical and model-based RL methods rely on access to a generative model or simulator capable of querying arbitrary state-action pairs and performing instant resets. In physical systems, streaming data settings, and non-resettable environments, arbitrary state querying is impossible, restricting these algorithms to synthetic simulations. Consequently, the theoretical guarantees and policy optimization mechanisms derived under generative oracle assumptions fail completely when deployed in realistic single-trajectory or trajectory-constrained online environments.

Possible candidate · 2/5 runs7 papers report this100% from 2025+

Language Agents & Tool Use

Effect to explainOpen

Benchmarking and Diagnosing Agent Framework Failures Across Open-Weight LLM Backbones

Current language agent and tool-use frameworks are overwhelmingly benchmarked on proprietary frontier models (GPT-4 class), masking severe brittleness on open-weight and smaller models. When executed on 7B–8B backbones, these frameworks experience catastrophic failure rates, violating task constraints in over 37% of cases and frequently generating invalid API queries or malformed function calls. This restricts practical agent deployment exclusively to cloud APIs, blocking autonomous tool use in privacy-sensitive, air-gapped, or resource-constrained local environments.

Possible candidate · 2/5 runs7 papers report this100% from 2025+

3D Gaussian Splatting & Novel View Synthesis

Scope to testOpen

Cross-Domain Robustness and Generalization Benchmarking for 3D Gaussian Splatting

Novel view synthesis and 3D Gaussian Splatting methods are currently evaluated almost exclusively on room-scale indoor datasets such as Replica, ScanNet++, and Gibson. Because no single pipeline has been systematically evaluated across indoor, outdoor unbounded, panoramic, and object-centric domains, the robustness and failure modes of these representations under domain shift remain entirely unknown. Practitioners and researchers cannot determine whether observed performance gains are artifacts of bounded indoor geometries or if current densification and optimization heuristics translate to unconstrained environments.

Possible candidate · 2/5 runs7 papers report this71% from 2025+

Hand-Object Interaction & Affordance

Scope to testPartly addressed

Benchmarking Hand-Object Interaction Models Across Diverse Hand Morphologies and Embodiments

Existing hand representation and hand-object interaction models are predominantly evaluated on small demographic cohorts or fixed kinematic setups (e.g., standard anthropomorphic models like MANO or fixed robotic hands like ShadowHand). Because existing pipelines implicitly bind kinematics, palm-to-finger ratios, or subject identities into their representations, their performance on unseen human hand proportions and non-standard robotic end-effectors remains untested. As a result, it is unknown whether current interaction and affordance estimators degrade gracefully or fail catastrophically when applied outside their narrow training morphologies.

Possible candidate · 2/5 runs7 papers report this29% from 2025+

Vision-Language Understanding & Captioning

Scope to testPartly addressed

Benchmarking the Cross-Model and Cross-Failure Generalization of Vision-Language Hallucination Mitigation

Current vision-language hallucination mitigations are evaluated under fragmented and narrow experimental conditions—often restricted to nominal entity metrics like CHAIR, small image sets, and single model checkpoints. Because existing techniques target specific decoding-time attention dynamics without assessing transfers across diverse architectures, hyperparameter shifts, or non-nominal hallucination types (such as relation errors or missing pretraining facts), their actual robustness is unverified. Consequently, researchers cannot determine whether observed gains reflect genuine mitigation of multimodal hallucinations or hyperparameter overfitting to specific evaluation setups.

Possible candidate · 2/5 runs7 papers report this86% from 2025+

RL Theory & Markov Decision Processes

Effect to explainOpen

Regret Bounds with Problem-Dependent Constants Removed: Minimax Characterization of the Gap Between Instance-Independent and Instance-Dependent Guarantees

Every one of these bounds is technically instance-independent, but each carries a "constant" — κ, exponential variation budgets, hitting-time factors, (1−γ)^−6, K^n, binomial coefficients — that is itself an instance-dependent quantity scaling exponentially or polynomially-huge on exactly the instances practitioners care about. The result is that published algorithms are provably efficient only on benign instances, and no one knows which of these constants are artifacts of analysis versus intrinsic hardness. Without a theory of which dependence is real, practitioners cannot tell whether a tight-constraint CMDP or a large-state bandit is fundamentally hard or merely badly analyzed.

Possible candidate · 1/2 runs7 papers report this25% from 2025+

Camera Pose, SLAM & Structure-from-Motion

Barrier to removeOpen

Robust Pose and 3D Structure Estimation Under Imperfect and Open-World Instance Masks

Current object-centric camera pose estimation, SLAM, and structure-from-motion pipelines treat 2D or 3D instance segmentation as a frozen upstream precondition (predominantly off-the-shelf Mask R-CNN or Mask3D). When masks contain boundary leakage, missed detections, or category misclassifications, geometric estimation degrades sharply because segmentation noise directly injects outlier correspondences that standard robust estimators cannot fully filter out. Consequently, existing methods cannot operate reliably outside closed object benchmarks with clean instance masks.

Strong candidate · 5/5 runs6 papers report this33% from 2025+

Diffusion Models

Barrier to removeOpen

Black-Box and Few-Step Control of Diffusion Models Without Intermediate State Access

Existing diffusion manipulation, guidance, and acceleration methods almost universally require white-box model access, intermediate noise predictions, or dozens of trajectory steps. Consequently, these techniques fail when interacting with commercial API endpoints that only return final image outputs, or with modern few-step distilled models (e.g., 1-to-4-step generators) where temporal redundancy across timesteps does not exist. Practitioners cannot apply state-of-the-art guidance, watermarking, or editing pipelines to closed-source foundation models or ultra-fast distilled backends. Resolving how to steer or audit generation under sample-only access is required to bridge the gap between white-box research methods and deployed diffusion infrastructure.

Strong candidate · 4/5 runs6 papers report this50% from 2025+

Model Safety, Alignment & Jailbreaks

Scope to testOpen

Evaluating the Cross-Modal Robustness of Text-Centric LLM Safety and Jailbreak Defenses

Safety filters, alignment mechanisms, and jailbreak detectors are predominantly developed and validated on text-only language models. When these models are extended to multimodal settings—such as processing visual text, image-text queries, or non-textual inputs—it remains unknown whether text-derived guardrails maintain their defensive efficacy. Deploying multimodal systems without evaluating these defenses risks catastrophic safety failures through visual or cross-modal bypasses that text-based benchmarks never expose.

Strong candidate · 4/5 runs6 papers report this67% from 2025+

Question Answering & Reading Comprehension

Effect to explainOpen

Principled Allocation of Low-Rank Adaptation Capacity When Full Fine-Tuning Is Unaffordable

LoRA is the default substitute for full fine-tuning, yet across tasks it produces large, sometimes catastrophic, performance drops relative to full updates — and practitioners currently have no way to predict or prevent them. The failures stem from arbitrary capacity choices: uniform ranks across layers and clients, tiny adapter budgets, and no guidance on when the low-rank constraint itself (rather than the task) is responsible. Until rank and layer allocation are treated as first-class decisions, LoRA deployments remain gambles that either waste compute or silently degrade quality.

Strong candidate · 4/5 runs6 papers report this33% from 2025+

Scientific ML: PDEs & Neural Operators

Scope to testOpen

Systematic Generalisation and Robustness Evaluation of Neural Operators in 3D PDE Systems

Current neural operator architectures are evaluated almost exclusively on 2D planar domains, leaving their predictive accuracy and robustness in 3D systems completely unmeasured. Without multi-dimensional benchmarks and controlled cross-dimensional studies, it is unknown whether empirical performance, discretization invariance, and error characteristics observed in 2D transfer to three-dimensional physical systems. Consequently, researchers and practitioners cannot determine whether 2D results provide reliable evidence for full 3D simulation tasks.

Possible candidate · 3/5 runs6 papers report this50% from 2025+

Model Safety, Alignment & Jailbreaks

Scope to testOpen

Evaluating the Generalisation of Small-Model Safety and Jailbreak Techniques to 70B+ and Frontier Models

Current safety, alignment, and jailbreak methodologies are overwhelmingly developed and empirically evaluated on small models under 14B parameters. Safety defenses and attack surfaces developed on smaller models are assumed to transfer to production scale, but whether their efficacy degrades, persists, or exhibits non-linear scaling transitions at 70B+ parameters remains unverified. As a result, the literature cannot determine whether safety interventions tested on 7B–13B models provide meaningful security guarantees for deployed frontier systems.

Possible candidate · 3/5 runs6 papers report this100% from 2025+

Video Tracking & Action Recognition

Effect to explainOpen

Diagnosing and Overcoming the Short-Horizon Saturation Barrier in Dense Video Models

Across multiple video tracking, segmentation, and action recognition architectures, empirical performance saturates or degrades once temporal context exceeds roughly 4 to 8 frames, regardless of model capacity. As a result, dense video models remain structurally constrained to sub-second temporal horizons despite sequence backbones theoretically capable of ingesting longer sequences. Current systems cannot reliably leverage extended frame histories to resolve occlusions, long-term state changes, or temporally dispersed actions, capping effective temporal receptive fields across tracking and recognition benchmarks.

Possible candidate · 3/5 runs6 papers report this83% from 2025+

Protein & Biomolecular Sequence Modeling

Barrier to removePartly addressed

A Benchmark of Experimentally Validated Functional Motifs for Training and Evaluating Motif-Conditioned Generative Models

Motif-conditioned generators in both protein structure and molecular design are trained and scored against motifs that are artifacts of the pipeline — random residue crops, rigid bond-breaking heuristics, or proxy-evaluated structures — rather than motifs whose function has been experimentally verified. Because these synthetic motifs may not match the distribution of true functional motifs, reported success rates measure the ability to satisfy a surrogate objective, not the ability to produce functional designs. Until a curated set of validated motifs with defined success criteria exists, no model comparison (structure-based vs. sequence-based, heuristic vs. expert extraction) can be trusted, and designs cannot be credibly advanced toward wet-lab validation.

Possible candidate · 3/5 runs6 papers report this67% from 2025+

RLHF & Preference Optimization

Barrier to removeOpen

Preference Optimization Under Unlabeled and Heterogeneous Real-World Annotator Noise

Current theoretical and algorithmic advances in noise-robust and personalized reward modeling fundamentally rely on preconditions absent in standard preference datasets: either assuming synthetic noise models generated by golden reward models or requiring persistent annotator identifiers and graphs. In practice, public preference datasets are largely anonymous, aggregated across heterogeneous annotator pools, and exhibit non-monotone human error patterns that synthetic noise models fail to capture. As a result, noise-mitigation and debiasing techniques developed under idealized assumptions remain unvalidated and often brittle when deployed on actual crowdsourced preference data.

Possible candidate · 3/5 runs6 papers report this83% from 2025+

Continual Learning & Catastrophic Forgetting

UnclassifiedOpen

Sub-Quadratic Covariance Sketching for Exemplar-Free Continual Learning

Exemplar-free continual learning methods increasingly rely on storing per-class feature statistics (prototypes and full covariance matrices) to generate pseudo-replay features or parameterize Mahalanobis classifiers without retaining raw inputs. However, storing a dense $D \times D$ covariance matrix for each of $C$ classes incurs $\mathcal{O}(C \cdot D^2)$ storage, which quickly exceeds device memory when scaling to large label spaces ($C \ge 1000$) or modern high-dimensional embeddings ($D \ge 1024$). Practitioners currently resort to naive diagonal approximations that discard inter-feature correlations and cause severe classification degradation under continual distribution shifts. As a result, covariance-based exemplar-free learning is structurally blocked from scaling to realistic large-vocabulary benchmarks on memory-constrained hardware.

Possible candidate · 3/5 runs6 papers report this50% from 2025+

Information Extraction & Entity Recognition

Effect to explainOpen

Semantic Typing and Grounding of Pronominal and Nominal Mentions Across Divergent Coreference Schemas

Current entity recognition and extraction pipelines exhibit sharp performance drops when evaluated on schemas that include pronominal and nominal mentions, dropping to as low as 42% recall on datasets like ACE 2005. Because most taggers assume proper named entities, models fail to type clusters composed entirely of pronouns or misclassify ambiguous pronouns (such as mistaking collective 'we' as person rather than organization) in the absence of cross-sentence antecedent context. Consequently, roughly 10% or more of coreference mentions remain completely ungrounded or unassigned, preventing reliable document-level entity extraction.

Possible candidate · 3/5 runs6 papers report this50% from 2025+

Dialogue & Conversational Systems

Scope to testOpen

Benchmarking Dialogue Management and Planning Beyond Predefined Action Spaces and Static Plans

Existing dialogue management and conversational planning frameworks have predominantly been evaluated within restricted operational scopes, such as predefined sets of dialogue acts, single-turn updates, static initial plans, and two-party interactions. Consequently, it remains unknown how robustly current dialogue planning architectures generalise when moving to open-ended token generation, mid-dialogue plan revision, and multi-party settings without pre-annotated action schemas. Without systematic empirical evaluation across these untested conditions, dialogue system designers cannot determine whether structured planning methods degrade or transfer to unconstrained, dynamic conversations.

Possible candidate · 2/5 runs6 papers report this67% from 2025+

Visual Reasoning & Embodied Navigation

Barrier to removePartly addressed

Visual Reasoning Without Dense Scene-Graph and Rationale Preconditions

Current visual reasoning methods depend on restrictive preconditions, including ground-truth answer hints during rationale generation, curated contrastive image pairs, and dense GQA-style scene-graph annotations. Because these annotations do not exist in most real-world datasets or uncurated domains, existing techniques cannot be transferred to novel environments or extended to multi-image tasks. Furthermore, evaluations have remained confined to narrow, single-image benchmarks, leaving model robustness in unannotated settings unknown. Consequently, visual reasoning systems cannot be deployed outside the handful of bespoke benchmarks engineered to support them.

Possible candidate · 2/5 runs6 papers report this67% from 2025+

Bandits & Online Learning

Barrier to removeOpen

Instance-Dependent Robustness Guarantees Beyond Worst-Case Regret Bounds for Constrained Online Learning

Today, online learning guarantees are only available under restrictive structural conditions—honest agents dominating malicious ones, known adversary classes, Slater's condition, smoothness assumptions—and when these fail, the bounds become vacuous or degrade catastrophically. Practitioners cannot distinguish between "the bound is loose" and "the problem is genuinely hard," because no instance-dependent analysis exists that quantifies what is actually achievable on a given problem instance. This blocks deployment in settings with few honest agents, near-zero constraint slacks, or moderately hard function classes, precisely where robust methods matter most.

Possible candidate · 2/4 runs6 papers report this67% from 2025+

Knowledge & Dataset Distillation

Barrier to removeOpen

Cross-Modal 3D Distillation Without Paired and Synchronized LiDAR Streams

Existing cross-modal distillation frameworks for 3D perception strictly require synchronized, spatially calibrated LiDAR point clouds and pre-trained LiDAR teachers alongside camera or radar feeds during training. This strict precondition limits distillation to expensive research vehicles with high-end sensor rigs, preventing its application to the vast majority of production fleet datasets where LiDAR is absent, uncalibrated, or asynchronously logged. As a result, low-cost sensor suites cannot leverage pre-trained LiDAR representations unless rigid multimodal collection setups are maintained during training.

Possible candidate · 2/5 runs6 papers report this0% from 2025+

Machine Unlearning

Effect to explainPartly addressed

Benchmarking the Adversarial Robustness and Reversibility of LLM Unlearning

Existing LLM unlearning methods are predominantly evaluated on standard, benign queries, giving a false sense of compliance with privacy and copyright demands. Empirical evidence shows that "unlearned" knowledge remains extractable via adversarial jailbreaks, latent-space elicitation, and few-shot relearning. Without systematic evaluation across these extraction vectors, practitioners have no way to verify whether a model has actually eliminated sensitive data or merely applied a superficial suppression mask.

Possible candidate · 2/5 runs6 papers report this83% from 2025+

Molecular Modeling & Drug Design

Effect to explainPartly addressed

Lossless Generalization Across Molecular Scaffold Diversity: Closing the Accuracy Gap on Unseen Compounds

Molecular generative and predictive models currently trade off reconstruction fidelity, chemical validity, and generalization: models that reconstruct accurately lose validity on novel chemistry, and models that stay valid degrade sharply on scaffolds, functional groups, and complex molecules absent or rare in training. This means any model's utility on genuinely new chemistry — the cases that matter most in drug discovery — is unpredictable and unmeasured per scaffold. Structurally, the field cannot yet certify that a model generalizes to uncorrelated compounds rather than interpolating within its training distribution.

Possible candidate · 1/2 runs6 papers report this100% from 2025+

Deep Learning Theory & Optimization Dynamics

Scope to testPartly addressed

Cross-Architecture Empirical Validation of Deep Learning Theory Predictors

Theoretical analysis and optimization dynamic guarantees in deep learning are overwhelmingly developed and proved on shallow fully connected ReLU models or two-layer networks. As a result, theoretical claims regarding convergence rates, generalization bounds, and optimization statistics are rarely verified on non-feedforward or deeper architectures such as CNNs, ResNets, and Transformers. Without systematic empirical evaluation across these structural settings, it remains unknown whether these theoretical quantities correlate with actual training dynamics or are merely artifacts of idealized shallow feedforward assumptions.

Weak candidate · 1/4 runs6 papers report this17% from 2025+

Language Model Interpretability & In-Context Learning

UnclassifiedOpen

A benchmark of naturally occurring hallucinations with model-generated ground truth, replacing synthetic entity-swap surrogates for detection evaluation

Today, hallucination detection methods are trained and evaluated against synthetic artifacts — entity-swapped summaries, outputs from a different model than the one under test — rather than hallucinations that actually arise during generation. This means reported detection numbers do not measure whether a method works on the failure mode that matters, and the modest predictive results (e.g., R² ≈ 0.27 on hallucination scores) may partly reflect this distribution mismatch. Detection validated only on 7B–13B open-weight models with synthetic data cannot be trusted as evidence about frontier-model behavior. Without a natural-hallucination benchmark, every detection paper inherits the same unquantified gap between "detects injected errors" and "detects real hallucinations."

Weak candidate · 1/3 runs6 papers report this33% from 2025+

Audio, Speech & Music Generation

Barrier to removeOpen

Black-Box Optimization and Guidance for Audio Generation Without Target ASR White-Box Access

Current audio and speech generation frameworks that optimize outputs against target ASR systems require white-box access to model internals to compute gradients or rely on explicit phoneme alignment pipelines. This precondition blocks the deployment, alignment, and evaluation of generative models against commercial, proprietary, or closed-source ASR APIs. Additionally, the requirement for differentiable phoneme alignments restricts applicability to languages that lack dedicated phonetic toolkits. As a result, generative audio models cannot be systematically guided or audited against real-world speech recognizers whose parameters and architectures are hidden.

Strong candidate · 5/5 runs5 papers report this20% from 2025+

Offline & Model-Based RL

Effect to explainOpen

Stabilizing Value Learning and Policy Extraction in Sparse-Reward Offline Reinforcement Learning

Current offline reinforcement learning algorithms suffer from catastrophic training instability, late-training policy collapse, and extreme variance on sparse-reward, long-horizon tasks such as AntMaze. To achieve reported benchmark numbers, practitioners routinely resort to unprincipled workarounds, including online pre-training interactions, domain-specific hyperparameter schedules, and checkpoint selection via test-environment rollouts. Without stable training dynamics under sparse feedback, offline RL cannot be deployed in high-stakes or real-world settings where offline checkpoint selection and online environment querying are strictly prohibited.

Strong candidate · 4/5 runs5 papers report this60% from 2025+

Language Model Interpretability & In-Context Learning

Scope to testPartly addressed

Benchmarking Interpretability and In-Context Learning Mechanisms Beyond Toy Synthetic PCFGs

Mechanistic interpretability and in-context learning findings in formal linguistic settings currently rely on tiny, unambiguous PCFGs with vocabularies of fewer than 100 words and corpora of only ~20K sentences evaluated on toy models like nanoGPT. It remains completely unknown whether the specific internal circuits, parsing behaviors, and induction mechanisms documented in these toy studies survive when evaluated on richer formal grammars, ambiguous grammars, larger vocabularies, or natural syntax. Without systematic robustness evaluations across these broader settings, interpretability claims cannot be reliably extrapolated beyond the micro-benchmarks on which they were discovered.

Possible candidate · 3/5 runs5 papers report this60% from 2025+

Semantic & Instance Segmentation

Barrier to removePartly addressed

Benchmarking and Bridging Idealized Preconditions in Promptable Segmentation

Current promptable segmentation models systematically rely on idealized preconditions during evaluation and training, including ground-truth-derived prompts (points or tight boxes extracted directly from GT masks), dense instance annotations, and perfect spatial feature alignment. In real-world interactive deployment, prompts originate from noisy human clicks, coarse bounding boxes, or imperfect auxiliary detectors, leading to uncharacterized performance degradation. Furthermore, models optimized exclusively on point prompts fail to maintain performance across alternative prompt modalities or without high-resolution decoders. Consequently, reported benchmark gains reflect idealized prompting setups rather than actual operational capability.

Possible candidate · 3/5 runs5 papers report this60% from 2025+

Language Model Interpretability & In-Context Learning

Scope to testOpen

Testing the Transferability of Toy-Model Theoretical Guarantees to Deep Transformer Architectures

Theoretical analyses of in-context learning and mechanistic interpretability rely almost entirely on heavily idealized architectures, such as single-layer or two-layer transformers with linear attention, orthogonal embeddings, and no residual connections or normalization. Because these proofs make structural assumptions that diverge substantially from production models, it is currently unknown whether their theoretical guarantees hold in real large language models or function purely as loose analogies. Without empirical validation across the gap between toy settings and deep models, downstream interpretability methods risk relying on mechanistic claims that fail in standard architectures.

Possible candidate · 3/5 runs5 papers report this60% from 2025+

Person Re-Identification & Face Recognition

Effect to explainOpen

Cross-Distribution Robustness and Granularity Evaluation in Cloth-Changing Person Re-Identification

Existing cloth-changing person re-identification (CC-ReID) and gait recognition models are trained and tested on datasets with restricted clothing variation profiles, such as upper-body-only changes or invariant footwear. Models trained on these biased distributions experience severe performance drops (such as single-digit Rank-1 accuracy) when evaluated on full clothing or pants-only changes. Without a standardized cross-setting evaluation that separates upper, lower, full, and footwear changes, reported benchmark metrics obscure localized overfitting and fail to measure true clothing-invariant representations.

Possible candidate · 3/5 runs5 papers report this40% from 2025+

Language Model Evaluation & Benchmarking

Barrier to removeOpen

Label-Free and Transferable Calibration for Large Language Model Outputs

Existing calibration and uncertainty quantification techniques for large language models (LLMs) strictly require task-specific, ground-truth-labeled calibration sets with binarizable correctness annotations. In open-ended, domain-specific, or novel tasks where ground-truth labels do not exist or cannot be easily thresholded, these methods cannot be instantiated. Consequently, practitioners are structurally blocked from obtaining calibrated confidence scores, prediction intervals, or coverage guarantees for LLMs without first performing expensive per-task and per-model annotations.

Possible candidate · 2/5 runs5 papers report this100% from 2025+

Domain Adaptation & Generalization

Scope to testPartly addressed

Cross-Architecture Robustness and Evaluation of Domain Adaptation in Object Detection

Current domain adaptation and domain generalization methods for object detection are almost exclusively evaluated on two-stage anchor-based architectures, primarily Faster R-CNN with ResNet backbones. Because these adaptation mechanisms often couple directly to region proposal networks, RoI pooling representations, or batch normalization layers, it is completely unknown whether reported domain adaptation gains transfer to one-stage detectors (e.g., FCOS, YOLO) or query-based transformers (e.g., DETR variants). Consequently, practitioners deploying modern transformer or anchor-free architectures have no empirical guidance on whether established domain adaptation methods function outside Faster R-CNN.

Possible candidate · 2/5 runs5 papers report this60% from 2025+

Equivariance & Symmetry in Neural Networks

UnclassifiedOpen

Auditing the Gap Between Theoretical Equivariance Guarantees and Empirical Performance in Finite, Symmetry-Broken Architectures

Current theoretical guarantees for equivariant neural networks rely on assumptions that practical pipelines routinely violate: infinite network width, exact group representations, strict absence of symmetry breaking, and fully equivariant backbones. Practitioners regularly deploy approximate, learned, or symmetry-broken equivariant models alongside non-equivariant backbones (e.g., ResNet, Swin), where theoretical equivariance guarantees formally collapse. Because no unified evaluation has tested these isolated scope limitations across common benchmarks, it remains unknown whether reported empirical gains stem from geometric inductive biases or trivial capacity increases, and why learned equivariance fails to extrapolate under distribution shift.

Possible candidate · 2/4 runs5 papers report this40% from 2025+

Knowledge & Dataset Distillation

Barrier to removeOpen

Knowledge Distillation Under Strict Black-Box Teacher Preconditions

Current high-performance distillation techniques rely on white-box access to intermediate representations, full logit vectors, or direct control over the teacher's training dynamics. When the teacher is a proprietary API or closed-source system that only returns discrete text or top-1 predictions, these methods cannot run. Downstream practitioners are structurally blocked from transferring capabilities from frontier closed models into compact architectures using state-of-the-art distillation losses.

Strong candidate · 5/5 runs4 papers report this75% from 2025+

Language Model Evaluation & Benchmarking

Barrier to removeOpen

Probability-Free LLM Evaluation and Benchmarking for Black-Box Models

Many standard evaluation, uncertainty estimation, and benchmarking techniques rely on token-level log-probabilities or soft output distributions. Proprietary and closed-access models (e.g., commercial chat interfaces and restrictive API endpoints) often only expose discrete text outputs, completely excluding them from probability-dependent benchmark comparisons. Relying on Monte Carlo sampling frequencies as a fallback is computationally prohibitive and prone to distortion from proprietary sampling temperatures or top-$p$ truncations. Consequently, existing evaluation suites cannot rigorously compare open-weight and closed-access models under a uniform probability-dependent protocol.

Strong candidate · 5/5 runs4 papers report this25% from 2025+

Depth & Stereo Estimation

Effect to explainPartly addressed

Systematic Cross-Domain Real-World Evaluation of Synthetic-Trained Depth and Stereo Models

Current depth and stereo estimation models are trained almost exclusively on synthetic datasets (such as TartanAir, VKITTI2, and OmniHouse) and are typically evaluated either on synthetic test splits or on narrow qualitative real-world samples. Because quantitative ground-truth evaluation across diverse physical environments is missing, practitioners cannot predict failure modes or transfer performance outside synthetic driving and indoor scenes. Without standardized multi-domain real-world evaluation, competing architectures cannot be fairly compared on real generalization capability.

Strong candidate · 4/5 runs4 papers report this100% from 2025+

Watermarking & Deepfake Detection

Barrier to removeOpen

Black-Box and Model-Agnostic Deepfake and Watermark Verification for Closed-Source Generative Models

Current watermark verification and deepfake detection methods require white-box access to target model parameters, intermediate layer activations, or gradient information. Because leading generative models and vision-language systems are served exclusively behind closed commercial APIs, internal inspection methods cannot be executed by downstream verifiers or auditors. Consequently, defenders cannot detect deepfakes or verify watermarks when the generating model is proprietary, unknown, or inaccessible. Removing the precondition of internal weight and activation access is necessary for auditing real-world deployed models.

Strong candidate · 4/5 runs4 papers report this100% from 2025+

Aerial, Satellite & BEV Perception

Barrier to removeOpen

Robust Cross-View Perception and Localization Without Paired Ground-Level Street Imagery

Current cross-view visual geolocalization, 3D scene reconstruction, and BEV perception models strictly depend on dense, spatially calibrated ground-level panoramas paired with satellite imagery or pre-existing 3D city meshes. In rural, sparsely populated, or underdeveloped regions, ground-level street-view coverage is either completely absent or uncalibrated. This reliance on dense ground-level imagery structurally restricts cross-view perception algorithms to major metropolitan areas with dedicated mapping fleets.

Strong candidate · 4/5 runs4 papers report this25% from 2025+

Neural Text Generation & Decoding

Barrier to removePartly addressed

Zero-Shot Machine-Generated Text Detection Under Strict API and Log-Probability Constraints

Current zero-shot machine-generated text detectors impose heavy access preconditions on the suspected source model, requiring either exact token log-probabilities, specialized completion API endpoints exposing top-K logprobs, or multiple live regeneration queries per candidate text. These requirements fail when evaluating outputs from proprietary models that only provide text-level chat interfaces without log probabilities, when source API access is restricted or cost-prohibitive, or when the generating model identity is unknown. Under these realistic operational constraints, falling back to surrogate scoring models results in steep performance drops (cross-model AUROC falling to 0.6–0.9), leaving auditors without viable tools to verify proprietary or uncooperative model outputs.

Strong candidate · 4/5 runs4 papers report this25% from 2025+

Efficient Attention & State-Space Models

Scope to testOpen

Evaluating Toy-Task Inductive Biases and Mechanistic Claims Across Standard Pre-training and Deep Architectures

Architectural mechanisms and efficiency claims developed for state-space models and efficient attention variants are overwhelmingly validated only on shallow (1–6 layer) networks trained on small synthetic or algorithmic tasks. Because these mechanisms have not been evaluated across standard natural language pre-training corpora or deeper architectures, the community cannot determine whether the reported properties persist, degrade, or dissolve at scale. Consequently, theoretical insights and inductive biases identified on synthetic benchmarks remain unverified hypotheses for real-world language modeling.

Strong candidate · 1/1 runs4 papers report this75% from 2025+

Machine Unlearning

Effect to explainPartly addressed

Empirical Robustness and Sensitivity of Machine Unlearning Under Realistic Data and Checkpoint Constraints

Existing machine unlearning algorithms for language models are predominantly evaluated under idealized conditions: complete access to forget sets, clean entity anchors, standard prose, and checkpoints saved immediately before target data exposure. In practice, unlearning requests frequently present partial forget data, domain variations like code, noisy or alias-heavy entity mentions, and checkpoints separated from target exposure by billions or trillions of tokens. Practitioners currently cannot predict whether an unlearning method that succeeds on curated benchmarks like TOFU will retain any efficacy when deployed under these real-world data and provenance constraints.

Strong candidate · 1/1 runs4 papers report this50% from 2025+

Federated Learning

Barrier to removePartly addressed

Eliminating Server-Side Auxiliary Data Dependencies in Federated Model Alignment and Calibration

Many advanced federated learning (FL) algorithms for feature alignment, knowledge distillation, and client temperature calibration require an auxiliary (public or proxy) dataset located at the central server. In privacy-critical cross-silo or cross-device deployments (such as medical imaging or sensitive mobile data), representative public proxy data simply does not exist or cannot be stored on the server due to compliance constraints. Consequently, these federated methods cannot be deployed in true zero-server-data environments, forcing practitioners to fall back to basic parameter-averaging baselines that suffer severe degradation under heterogeneous data distributions.

Possible candidate · 3/5 runs4 papers report this0% from 2025+

Adversarial Robustness & Backdoor Attacks

Barrier to removePartly addressed

Backdoor and Poisoning Attacks Under Constrained Training-Time Access

Current advanced backdoor injection and model manipulation techniques strictly require full white-box access to the target architecture, parameters, and unconstrained control over the training optimization process. In realistic supply chain or outsourced training scenarios, an adversary or security auditor rarely has end-to-end control over the victim's training loss, optimization hyperparameters, and internal architecture. Because existing techniques break down when these white-box preconditions are removed, security assessments dramatically overestimate attacker capabilities or fail to evaluate practical threat surfaces.

Possible candidate · 3/5 runs4 papers report this50% from 2025+

Watermarking & Deepfake Detection

Barrier to removePartly addressed

Black-Box and Architecture-Agnostic Watermarking and Defense for Diffusion Models

Current proactive watermarking, attribution, and anti-customization defenses structurally assume white-box access to target latent diffusion models, requiring direct inspection of latent spaces, VAE fine-tuning, or gradient backpropagation through internal U-Net attention layers. In practical deployment scenarios, defenders and copyright holders must protect content against proprietary black-box APIs, pixel-space diffusion models, or non-VAE architectures where internal activations and weights are inaccessible. Because existing methods are tethered to specific LDM preconditions, there is currently no verified mechanism to provide provenance or anti-fine-tuning protection across arbitrary or black-box diffusion pipelines.

Possible candidate · 3/5 runs4 papers report this100% from 2025+

Protein & Biomolecular Sequence Modeling

Barrier to removePartly addressed

Quantifying and Mitigating Surrogate Oracle Exploitation in Biomolecular Sequence Optimization

Generative and reinforcement learning methods for protein design optimize against learned fitness predictors or computational proxies (such as ESMFold or AutoDock Vina) rather than ground-truth biophysical assays. As optimization progresses, generated sequences exploit the blind spots and low-label estimation errors of these surrogates, producing candidates with inflated reward scores that fail in true physical or experimental evaluations. Because contemporary literature routinely evaluates generated designs using the same surrogates that guided optimization, reported performance gains frequently reflect proxy over-optimization rather than genuine biophysical efficacy.

Possible candidate · 3/5 runs4 papers report this50% from 2025+

Watermarking & Deepfake Detection

Scope to testPartly addressed

Cross-Family Generalization and Robustness of Text Watermarking Attacks and Defenses

Existing evaluations of text watermark attacks, detection, and removal methods have almost exclusively tested narrow subsets of logit-based schemes (e.g., KGW and SIR) evaluated on fixed base models like OPT-1.3B. Because no single study has evaluated these attack and defense methods across fundamentally different watermark families—including sampling-based, dynamic-hash, content-adaptive, and error-correction-augmented schemes—practitioners cannot determine whether reported evasion and detection results reflect general properties or family-specific artifacts. Consequently, deployment decisions for provenance tracking in production language models rely on unverified theoretical compatibility claims rather than measured empirical robustness.

Possible candidate · 3/5 runs4 papers report this75% from 2025+

RLHF & Preference Optimization

Scope to testOpen

Empirical Robustness of Preference Optimization Across Synthetic, Multi-Annotator, and Real-World Feedback Distributions

Current alignment algorithms (e.g., DPO, PPO, KTO) are almost exclusively evaluated on synthetic user proxies, converted scalar scores, or curated single-annotator datasets like UltraFeedback. Because individual studies restrict evaluation to narrow, semi-synthetic testbeds (such as session-similarity graph samplings or single-dimension datasets like HelpSteer2), it remains unknown whether standard preference optimization algorithms generalize reliably when exposed to genuine multi-annotator disagreement and raw binary human signals. Without a cross-setting evaluation spanning these distinct regimes, practitioners cannot determine whether reported algorithmic gains reflect true alignment robustness or artifacts of synthetic data curation.

Possible candidate · 3/5 runs4 papers report this75% from 2025+

Hand-Object Interaction & Affordance

Barrier to removeOpen

Fine-Grained 3D Hand Affordance Grounding and Generation for Small-Part Articulated Interactions

Current affordance learning pipelines rely on self-training loops, video extraction, or upstream generative HOI models that fail when contact targets are small functional components such as bottle caps, zippers, or rotary knobs. Because generated hand poses easily drift to adjacent object surfaces, models exhibit catastrophic failure on precision affordances (e.g., 0% accuracy on pull or sub-50% on twist), propagating noisy pseudo-labels across unobserved geometries. Downstream manipulation systems are consequently blocked from executing precision tasks like uncapping, pulling sliders, or fine bimanual manipulation.

Possible candidate · 3/5 runs4 papers report this75% from 2025+

Video-Language Understanding & Temporal Grounding

Effect to explainOpen

Benchmarking and Mitigating Spatio-Temporal Grounding Degradation in Continuous Video-LLMs

Current video-language models evaluate spatial grounding primarily on sparse frames (e.g., 0.5 FPS), leaving performance in continuous video unmeasured. When models attempt spatio-temporal grounding, performance degrades significantly under multi-point trajectory supervision (e.g., mIoU drops beyond 3 tracking points) and dynamic multi-resolution token sampling raises spatio-temporal error relative to plain spatial pooling. Because full spatio-temporal box annotation is prohibitively expensive, visual grounding in video remains unresolved with sub-10% IoU baselines.

Possible candidate · 3/5 runs4 papers report this50% from 2025+

Visual Reasoning & Embodied Navigation

Scope to testOpen

Systematic Generalisation and Robustness Evaluation of Visual Reasoning Methods Across Distinct Perceptual Regimes

Current visual reasoning techniques are predominantly evaluated on static images containing localized, explicitly present visual evidence where language reasoning dominates. Consequently, it is unknown whether their performance improvements transfer to vision-dominant perceptual tasks, temporal video sequences, global visual synthesis problems, or queries involving implicit visual evidence. Practitioners cannot determine whether a method’s gains represent generalizable visual reasoning or narrow optimizations for localized static VQA. A systematic robustness evaluation across these untested dimensions is necessary to chart where existing methods hold and where they fail.

Possible candidate · 3/5 runs4 papers report this100% from 2025+

Out-of-Distribution & Anomaly Detection

Barrier to removePartly addressed

Representation Adaptation and End-to-End Learning for Anomaly Detection Beyond ImageNet Backbones

Modern visual anomaly and out-of-distribution (OOD) detection methods rely almost entirely on frozen feature representations extracted from ImageNet-pretrained backbones. In specialized domains such as medical imaging, industrial inspection, or scientific imaging, ImageNet features fail to encode domain-specific normal variation, capping downstream anomaly scoring accuracy. Because current pipelines do not adapt or learn backbone representations end-to-end from in-distribution data, anomaly detection is structurally unviable wherever high-quality pre-trained foundation models do not already exist.

Possible candidate · 3/5 runs4 papers report this50% from 2025+

Vision-Language Understanding & Captioning

Effect to explainOpen

Robust Vision-Language Pre-training Under Cascading Pseudo-Label Noise

Modern fine-grained vision-language and region-captioning models rely predominantly on automated annotation pipelines combining specialist detectors, segmenters, OCR, and large multimodal models without human verification. These multi-stage machine annotations suffer from compounded errors, such as mislocalized bounding boxes and hallucinated visual attributes, which directly propagate into downstream vision-language representations and degrade generation quality. Current learning-with-noisy-labels methods assume human inter-annotator disagreement or synthetic classification noise, failing to model the structured, cascading failure modes inherent to vision-language pseudo-labeling pipelines.

Possible candidate · 3/5 runs4 papers report this50% from 2025+

3D Gaussian Splatting & Novel View Synthesis

Effect to explainPartly addressed

Dynamic 3D Gaussian Splatting Under Non-Smooth and Discontinuous Motion Dynamics

Current dynamic novel view synthesis and 4D Gaussian Splatting frameworks assume smooth, continuous spatio-temporal deformations, linear velocities, and fixed temporal windows. When scenes exhibit abrupt impacts, rapid accelerations, chaotic object trajectories, or zooming and variable frame rates, these continuous deformation priors fail and yield severe geometric tearing, blurring, or total tracking collapse. This fundamentally blocks dynamic radiance fields from capturing real-world interactions such as sports action, mechanical impacts, and sudden dynamic occlusions.

Possible candidate · 3/5 runs4 papers report this75% from 2025+

Knowledge & Dataset Distillation

Barrier to removeOpen

Distillation Under Zero-Data and Unlabeled-Stream Preconditions

Current distillation frameworks depend on strict data preconditions, including access to original downstream training sets, in-domain unlabeled datasets, out-of-distribution reference data, or sample-level identity annotations. When proprietary, privacy, or security constraints prevent the release of both the underlying training set and external proxy data, these methods cannot execute at all. Furthermore, when teacher outputs arrive as an uncurated stream without identity labels, existing feature-bank and prototype-matching methods break down. Resolving these preconditions allows model compression and dataset synthesis in strictly zero-data or privacy-restricted deployments.

Possible candidate · 3/5 runs4 papers report this25% from 2025+

Multi-Agent & Goal-Conditioned RL

Scope to testOpen

Evaluating the Robustness and Limits of Tree-Search RL Under Stochastic Dynamics and Function Approximation

Existing theoretical guarantees and empirical evaluations for tree-search and tree-MDP RL algorithms are strictly restricted to deterministic, tabular, or binary-tree settings. Real-world decision-making domains involve transition stochasticity and large state spaces requiring neural function approximation, where deterministic branching assumptions are untested. Without systematic evaluation across stochastic transitions and continuous state spaces, it is unknown whether deterministic tree-guided methods degrade gracefully or fail catastrophically.

Possible candidate · 2/5 runs4 papers report this25% from 2025+

Scientific ML: PDEs & Neural Operators

Barrier to removePartly addressed

Semi-Supervised Neural Operator Learning for Parameter-Scarce Trajectories

Existing neural operators and trajectory surrogate models for continuum mechanics require dense, ground-truth physical parameter labels for every training trajectory. In experimental and real-world physical systems, recording state trajectories is feasible, but measuring underlying continuum or constitutive parameters is often costly, destructive, or impossible. Because existing methods lack semi-supervised formulations, they cannot leverage unlabeled experimental trajectories and remain restricted to fully labeled synthetic simulations.

Possible candidate · 2/5 runs4 papers report this75% from 2025+

Classical Supervised Learning: Trees, Ranking & Aggregation

Barrier to removeOpen

Empirical Robustness and Scaling of Jacobian- and Pairwise-Based Aggregation Beyond BERT-Base

Several classical supervised learning, ranking, and aggregation methods rely on per-example parameter Jacobians and all-pairs batch formulations, restricting their empirical verification to the linear readout layers of English BERT-base and ResNet-50. As a result, it is unknown whether the spectral variance assumptions and downstream performance of these algorithms hold on larger transformer architectures, decoder-only models, or multilingual representations. Practitioners currently have no evidence indicating whether these techniques fail gracefully or break down entirely when scaled beyond small, English-specific encoder backbones.

Possible candidate · 2/5 runs4 papers report this50% from 2025+

Uncertainty Quantification & Conformal Prediction

Scope to testOpen

Evaluating Uncertainty Quantification and Conformal Prediction Methods Beyond Binary Outcomes

Existing uncertainty quantification and conformal prediction formulations in this family have been developed and tested exclusively on binary outcome spaces ($y \in \{0, 1\}$). As a result, it is unknown whether their theoretical properties, coverage guarantees, and set efficiencies persist or degrade when applied to multiclass and continuous regression targets. Practitioners in non-binary domains are currently blocked from deploying these methods with verified reliability.

Possible candidate · 2/5 runs4 papers report this50% from 2025+

Machine Unlearning

Barrier to removePartly addressed

Machine Unlearning Under Degraded Operational Preconditions

Existing machine unlearning methods rely on strict operational preconditions: access to clean retain datasets ($D_r$), full white-box parameter access, historical pre-training checkpoints, or original pre-unlearned calibration weights. In production environments, intermediate training checkpoints are routinely deleted to save storage, retain data is often inaccessible due to data governance policies, and deployment interfaces may restrict full weight access. Without methods that function in the absence of these preconditions, deployed models cannot legally or practically comply with data deletion requests.

Possible candidate · 2/5 runs4 papers report this75% from 2025+

Digital Humans, Avatars & Virtual Try-On

Effect to explainPartly addressed

Benchmarking and Improving 3D Human Body Model Robustness Across Underrepresented Demographics and Non-Standard Morphologies

Existing 3D statistical body models and avatar pipelines are predominantly trained on narrow demographic subsets (such as adult European CAESAR scans) and evaluated solely on standard normative body templates. Consequently, downstream avatars and virtual try-on methods systematically produce distorted geometry, unnatural surface artifacts, and inaccurate anthropometric measurements on children, adolescents, diverse global populations, individuals with disabilities, and non-standard body proportions. Because existing benchmarks do not measure generalization across diverse body shapes, failure modes on these populations remain unquantified and unaddressed.

Possible candidate · 2/5 runs4 papers report this50% from 2025+

Domain Adaptation & Generalization

Barrier to removePartly addressed

Systematic Evaluation of Data Augmentation Techniques Across Disparate Distribution Shift Types

Data augmentation methods developed for out-of-distribution robustness are overwhelmingly hand-engineered to simulate surface-level texture and appearance variations (e.g., color perturbations and synthetic corruptions). When these methods are deployed under other distribution shift types—such as subpopulation shifts, geometric transformations, or domain shifts in specialized fields like medical imaging—they frequently fail or degrade accuracy below standard Empirical Risk Minimization (ERM). Because augmentation techniques have not been systematically evaluated across these differing shift categories, practitioners have no reliable way to anticipate when a given augmentation strategy will harm downstream generalization.

Possible candidate · 2/5 runs4 papers report this0% from 2025+

Robot Learning & Embodied Control

Barrier to removePartly addressed

Closed-Loop Trajectory Generation for Robot Manipulation Without External Motion Planners

Current keyframe-based robot manipulation policies predict static target end-effector poses and delegate trajectory generation to external sampling-based motion planners such as BiRRT. This decouples visual perception from execution dynamics, preventing real-time reactive corrections when environments change or objects slip mid-motion. Furthermore, relying on external planners introduces high computational latency and frequent planning failures in cluttered scenes where feasible kinematic paths are narrow.

Possible candidate · 2/5 runs4 papers report this100% from 2025+

Inference Efficiency & Long Context

Barrier to removePartly addressed

Calibration-Free and Gradient-Free Predictors for Efficient Inference

Current efficiency mechanisms for large language models and long-context inference—such as dynamic quantization, layer skipping, and speculative routing—rely on offline calibration phases requiring gradient access and curated calibration datasets. Consequently, these acceleration techniques cannot be deployed in privacy-sensitive environments where data cannot be stored, nor on black-box model APIs where backpropagation is unavailable. Furthermore, because predictors are model-specific, serving heterogeneous or frequently updated models incurs repeated offline profiling overhead that impedes zero-shot deployment.

Possible candidate · 2/5 runs4 papers report this50% from 2025+

Differential Privacy

Effect to explainOpen

Robustness and Sensitivity of Public-Data-Assisted Differential Privacy Under Distribution Shift and Proxy Misspecification

Differentially private algorithms increasingly rely on auxiliary public data or pre-set structural parameters—such as covariance proxies, compact feature bounds, and distributional similarity—to mitigate noise. However, in realistic deployments, public datasets rarely match private target distributions exactly, and assuming tight compact bounds or positive-definite covariance proxies can fail. When these auxiliary assumptions degrade, algorithms either suffer severe utility loss or must spend substantial privacy budget to estimate these quantities from private data. Without a systematic evaluation of how public-assisted DP methods perform across varying degrees of proxy error, bound truncation, and distribution shift, practitioners cannot determine when auxiliary public data provides a net benefit over purely private baselines.

Possible candidate · 2/5 runs4 papers report this75% from 2025+

Machine Unlearning

Effect to explainOpen

Scalability and Robustness of Machine Unlearning Under Varying Forget-Set Regimes

Existing machine unlearning techniques are only validated in narrow operating windows, typically limited to small forget budgets ($\le 10\%$). As forget sets scale beyond a few hundred samples or exceed 10–20% of the training distribution, retention accuracy drops significantly (e.g., a 14–15% gap relative to retraining from scratch on CIFAR-10), while tiny splits (e.g., 1%) yield negligible unlearning. Furthermore, unlearned representations remain fragile and easily recoverable when exposed to small fractions of the forgotten data during relearning. Consequently, current unlearning algorithms cannot be reliably deployed for large-scale data deletion compliance without catastrophic utility loss.

Possible candidate · 2/5 runs4 papers report this100% from 2025+

Sentiment & Emotion Analysis

Barrier to removePartly addressed

Benchmarking Emotion Analysis and Decoding Across Model Access Regimes and Representation Constraints

Current emotion analysis and decoding methods operate under mutually incompatible preconditions, requiring either white-box access to hidden layer activations or complete reliance on black-box model outputs. At the same time, existing pipelines restrict emotional representations to hand-assigned discrete coordinate grids (such as Russell's valence-arousal cells or Mikels' discrete categories) or depend on static dictionaries that assign neutral values to missing vocabulary. Because no single study has evaluated these methods across these different access regimes and representation constraints, practitioners cannot determine how much performance degrades when moving from white-box embeddings to black-box APIs, or from continuous spaces to fixed discrete taxonomies. Consequently, emotion-aware systems deployed via commercial APIs or out-of-vocabulary domains operate without quantified reliability bounds.

Possible candidate · 1/2 runs4 papers report this25% from 2025+

Algorithmic Fairness

Barrier to removePartly addressed

Benchmarking Algorithmic Fairness Methods Under Violations of Sensitive Attribute Preconditions

Existing algorithmic fairness methods are almost universally evaluated under rigid preconditions: fully observed or missing-at-random discrete group labels, bounded minority group representations ($\rho$-fraction), and absence of feature-dependent annotation noise. In real-world deployments, sensitive attributes are often unannotated, continuous, severely underrepresented, or corrupted by feature-dependent noise. Because fairness methods have not been systematically evaluated across these non-idealized regimes, practitioners cannot determine which, if any, algorithmic debiasing approaches retain their fairness guarantees or utility tradeoffs when theoretical preconditions fail.

Possible candidate · 1/2 runs4 papers report this75% from 2025+

Machine Unlearning

Barrier to removePartly addressed

Machine Unlearning Without External Discriminators or Teacher Reasoning Scaffolds

Current machine unlearning methods for complex language and reasoning tasks depend on external scaffolding—such as proprietary reasoning teacher LLMs, off-the-shelf discriminators (e.g., Detoxify, Scrubadub), pre-unlearned baseline checkpoints, or explicit chain-of-thought traces. In domains where pre-trained detectors do not exist or where target knowledge is elicited implicitly without structured traces, these unlearning pipelines fail to execute. As a result, model developers cannot sanitize models against novel domain-specific liabilities, proprietary data leaks, or implicit reasoning paths without commissioning expensive external supervision models for each new target.

Weak candidate · 1/3 runs4 papers report this75% from 2025+

Hand-Object Interaction & Affordance

Barrier to removeOpen

Hand-Object Affordance Estimation Beyond Clean Pre-Scanned Scenes and Predefined 3D Grids

Current methods for 3D hand-object interaction and affordance estimation depend strictly on clean geometric preconditions: high-fidelity 3D scans, pre-detected object instances, static overhead views, or tightly predefined 3D bounding grids. In practical manipulation and augmented reality scenarios, environments are unsegmented, viewpoints are dynamic, and point clouds are noisy and incomplete. Relying on manually tuned grid volumes and pre-extracted meshes prevents existing models from operating directly on raw, uncurated sensor streams.

Strong candidate · 5/5 runs3 papers report this33% from 2025+

Word Embeddings & Lexical Semantics

Barrier to removeOpen

Robustness Benchmarking of Lexical Semantics Methods Under Language Model Scarcity and Distribution Shift

Current lexical semantics and semantic change methods assume the availability of high-quality, off-the-shelf pretrained language models trained on massive, well-matched text distributions. When applied to historical texts, low-resource languages, or niche domains where such models do not exist or perform poorly, practitioners have no empirical evidence on how severely these methods degrade. This dependence leaves lexical semantics largely untested across ancient corpora, morphologically non-standard texts, and non-mainstream model architectures.

Strong candidate · 5/5 runs3 papers report this0% from 2025+

Watermarking & Deepfake Detection

Scope to testOpen

Benchmarking Watermarking and Deepfake Detection Across Non-Latent and Flow-Matching Generative Architectures

Current watermarking and detection methods in generative media are almost exclusively developed and evaluated on standard latent diffusion models relying on Gaussian noise sampling. Consequently, it is unknown whether these techniques transfer, degrade, or fail entirely when applied to pixel-space diffusion, non-iterative architectures, or modern rectified-flow and flow-matching models such as SD3 and FLUX. As production generative pipelines shift away from standard latent diffusion, safety and provenance mechanisms risk operating under untested assumptions. A systematic evaluation across these architectural families is necessary to establish the empirical boundary of existing detection and watermarking schemes.

Strong candidate · 5/5 runs3 papers report this67% from 2025+

Aerial, Satellite & BEV Perception

Effect to explainOpen

Benchmarking and Mitigating UAV Perception Failure Under Rapid Motion Blur, Scale Shifts, and Occlusion

Current perception models deployed on unmanned aerial vehicles achieve low absolute accuracy on standard downstream tasks (e.g., recall under 7.5% in scene understanding, under 40% in action recognition, and under 30% m_vIoU in tracking). These failures are driven by real-world flight conditions—rapid camera motion blur, extreme resolution and altitude shifts, severe occlusion, and small target sizes—that standard benchmarks do not systematically couple. As a consequence, aerial vision pipelines remain brittle and cannot be reliably trusted for autonomous navigation, search-and-rescue, or monitoring in complex flight regimes.

Strong candidate · 4/5 runs3 papers report this33% from 2025+

Information Extraction & Entity Recognition

Barrier to removeOpen

Quantifying and Mitigating Upstream Parser Brittleness in Syntactic Information Extraction

Syntactic-semantic information extraction pipelines depend strictly on upstream dependency parsers and part-of-speech taggers, leaving downstream relation and entity extraction vulnerable to cascading errors. When applied to domain-specific jargon or subjected to minor parser version updates, upstream accuracy drops and alters the extracted syntactic structures. This leaves practitioners unable to guarantee extraction stability across domains or library releases without manual parser fine-tuning.

Strong candidate · 4/5 runs3 papers report this33% from 2025+

Computational Social Science & Content Moderation

Barrier to removeOpen

Benchmarking Systematic Biases and Reproducibility in LLM Pragmatic Discourse Annotation

Computational social science and online moderation research increasingly substitute human coding with LLM annotation for subjective pragmatic phenomena such as conversational roles, implicit toxicity, and discourse dynamics. However, existing studies rely heavily on proprietary models like GPT-4, which introduce non-determinism, reproducibility barriers, and systematic biases toward surface stylistic features like politeness. Because smaller open-weight models frequently fail or skip messages on complex pragmatic schemas, researchers currently lack a validated, reproducible protocol or understanding of how LLM annotation errors propagate across subjective discourse settings.

Strong candidate · 4/5 runs3 papers report this100% from 2025+

Vision-Language Understanding & Captioning

Barrier to removeOpen

Cross-Cultural Knowledge Grounding for Vision-Language Understanding

Current knowledge-grounded vision-language benchmarks rely almost exclusively on English Wikipedia as their ground-truth reference base. Because Wikipedia has documented geographic and cultural footprint skews, models evaluated on imagery from low-footprint cultures suffer from missing reference facts rather than failures in visual-semantic reasoning. Consequently, benchmarks penalize models on underrepresented cultures and researchers cannot diagnose whether errors stem from visual perception, language generation, or knowledge retrieval omissions.

Strong candidate · 4/5 runs3 papers report this100% from 2025+

RLHF & Preference Optimization

Barrier to removeOpen

Eliminating Matched-Pair Preconditions and Scale Sensitivity in Contrastive Preference Guidance

Current contrastive guidance techniques for preference optimization structurally depend on having access to a matched tuned/untuned small proxy model pair whose latent reward aligns with the target preference. When such pairs are unavailable or poorly matched, guidance quality collapses and caps output quality at the proxy's low baseline capability (~10% win rate), while weak-to-strong self-steering without external reward signals fails to converge. Furthermore, the guidance scale parameter $\gamma$ exhibits high variance in noisy regions and fails to transfer across differing target architectures without periodic, expensive reward re-estimation.

Strong candidate · 4/5 runs3 papers report this33% from 2025+

Neural Text Generation & Decoding

Barrier to removeOpen

Benchmarking and Extending Machine-Generated Text Detection Under Black-Box, Mixed-Author, and Multi-Source Settings

Current machine-generated text detection frameworks rely on preconditions that fail in real-world deployments: access to token-level log-probabilities from the source model, abundant paired human reference corpora, and strictly binary (100% human vs. 100% machine) document inputs. Because these methods have not been evaluated across black-box commercial APIs, partially edited or collaborative human-AI text, and multi-model source attribution, their true utility is unknown. Practitioners cannot reliably determine whether a piece of text was generated, co-authored, or which specific model produced it when log-probabilities are inaccessible.

Strong candidate · 4/5 runs3 papers report this67% from 2025+

Person Re-Identification & Face Recognition

Effect to explainOpen

Robust Cross-Modality and Cross-Platform Retrieval for Person and Vehicle Re-Identification

Current cross-modality re-identification methods report relative improvements on benchmark datasets, yet their absolute retrieval performance drops severely in strongly mismatched sensor and platform scenarios (e.g., 9–16 mAP on RGBNT201 and MSVR310). When spectral shifts (RGB to Near-Infrared or Thermal) coincide with cross-platform viewpoint or domain shifts, identity discrimination fails at a fundamental level. This severe degradation blocks the practical deployment of automated re-identification systems across heterogeneous, multi-sensor surveillance networks operating in varying lighting and platform conditions.

Strong candidate · 4/5 runs3 papers report this100% from 2025+

Machine Unlearning

Barrier to removeOpen

Machine Unlearning and Verification Under Black-Box Model Access Constraints

Current machine unlearning algorithms and formal auditing frameworks overwhelmingly require white-box access to model parameters, gradients, or pre- and post-unlearning checkpoint weights. This structurally prevents unlearning operations and compliance verification on closed-source, API-hosted, or third-party proprietary models where internal weights cannot be inspected or directly modified. Without methods and evaluation protocols that operate under black-box or query-only access, unlearning remains unusable for hosted foundation models and unverifiable by external auditors.

Strong candidate · 1/1 runs3 papers report this100% from 2025+

Causal Inference & Discovery

UnclassifiedOpen

Robustness and Adaptation of Causal Frameworks to Arbitrary DAGs and Non-Linear SCMs

Current causal frameworks often rely on restrictive scope assumptions: fully known causal graphs restricted to simple canonical topologies (causal, anticausal, FIIF), strict linearity of structural equations, and sensitive feasibility thresholds. In real-world applications, true graphs are rarely known or confined to canonical 3-node topologies, and underlying mechanisms routinely exhibit non-linear dynamics. Consequently, practitioners cannot deploy these methods without risking severe model failure or infeasibility due to unverified structural and functional assumptions.

Strong candidate · 1/1 runs3 papers report this33% from 2025+

Audio, Speech & Music Generation

Barrier to removeOpen

Conditioning Audio and Speech Generation on Unstructured Holistic Text Descriptions

Current audio, speech, and music generation models cannot generate coherent multi-event audio directly from free-form, unformatted natural language descriptions. Instead, they depend on rigid preconditions: segment-by-segment timestamped scripts, multi-field prompt partitions, or explicit slot separations (e.g., distinguishing speech transcripts from ambient sounds). This forces users to manually author structured scripts or introduces a brittle dependency on upstream multimodal LLMs to parse and slot-fill prompts before generation can occur.

Possible candidate · 3/5 runs3 papers report this100% from 2025+

GANs, Style Transfer & Image Translation

Barrier to removeOpen

Evaluating and Extending Latent Image Manipulation Beyond Domain-Specific StyleGAN2 Backbones

Current GAN-based semantic editing and adaptation methods strictly require pre-trained, domain-specific StyleGAN2 checkpoints trained on curated, single-object datasets like faces and cars. For complex domains lacking pre-trained models (such as full-body poses or multi-object urban scenes), these manipulation pipelines cannot be deployed. Furthermore, it remains unknown whether existing latent editing techniques generalize to non-StyleGAN generator architectures or scale beyond few-shot target data regimes.

Possible candidate · 3/5 runs3 papers report this0% from 2025+

Aerial, Satellite & BEV Perception

Barrier to removePartly addressed

Robustness and Degradation Benchmarking of Map-Conditioned Perception Under Imperfect OpenStreetMap Priors

Contemporary BEV perception, aerial localization, and landmark-guided navigation models increasingly condition their visual representations on OpenStreetMap (OSM) vector geometry and metadata. In practice, OSM data exhibits severe geographic disparities, missing road classes, topological errors, and incomplete landmark tags. Because existing methods assume clean and complete vector maps as an operational precondition, system performance degrades unpredictably in under-mapped or rural areas, preventing deployment outside densely curated metropolitan regions.

Possible candidate · 3/5 runs3 papers report this100% from 2025+

Deep Learning Theory & Optimization Dynamics

Effect to explainOpen

Empirical Benchmarking of Generalization Failure in Alternative Optimization Dynamics

Alternative optimization methods, such as exact Gauss-Newton (GN) and non-backpropagation dynamics (e.g., NMNC), exhibit severe generalization gaps and early loss saturation when scaled to deep networks and mini-batch settings. Standard regularizers (dropout, weight decay, data augmentation, pseudoinverse regularization) and initialization heuristics developed on shallow models consistently fail to close these train-test gaps. Without a controlled comparative evaluation across these distinct settings, it remains unknown whether these generalization failures share common optimization dynamics or require fundamentally different stabilization interventions.

Possible candidate · 3/5 runs3 papers report this67% from 2025+

Medical Imaging

Barrier to removeOpen

Open-Vocabulary Diagnostic Knowledge Retrieval for Medical Imaging Without Predefined Ontologies

Current diagnostic and retrieval models in medical imaging depend on fixed disease categories and rigid hierarchical ontologies. When patients present with rare, novel, or out-of-vocabulary pathologies—or when clinical environments lack standardized label taxonomies—these systems fail to retrieve relevant medical knowledge or characterize visual findings. Consequently, automated diagnostic support is structurally blocked from assisting on atypical and long-tail conditions where clinical reference support is most needed.

Possible candidate · 2/5 runs3 papers report this0% from 2025+

Domain Adaptation & Generalization

Barrier to removeOpen

Cold-Start Streaming Domain Adaptation and Discovery Without Labeled Base Sessions or Offline Pre-Passes

Current continual domain adaptation and streaming open-world discovery methods structurally depend on an offline initialization phase—either requiring a fully labeled source base dataset or an offline full-dataset inference pass to compute initial prototypes and representations. In truly dynamic environments where data arrives purely as an online stream and domain shifts occur from the first observation, collecting labeled base data or executing multi-pass offline initialization is impossible. Consequently, existing frameworks cannot be deployed in pure cold-start streaming regimes.

Possible candidate · 2/5 runs3 papers report this100% from 2025+

Continual Learning & Catastrophic Forgetting

Barrier to removePartly addressed

Task-Agnostic Inference for Modular and Masked Continual Learning Architectures

Modular continual learning methods prevent catastrophic forgetting by allocating task-specific masks, sub-networks, or adaptive parameters during training. However, these methods structurally depend on having ground-truth task identifiers provided at test time to select the corresponding parameters. In practical deployments, incoming inputs arrive without task metadata, rendering task-dependent parameter masking unusable in class-incremental or task-agnostic settings. Without a reliable mechanism to resolve task identity or route inputs dynamically at test time, parameter-isolation architectures remain restricted to artificial task-incremental benchmarks.

Possible candidate · 2/5 runs3 papers report this0% from 2025+

Semantic & Instance Segmentation

Barrier to removePartly addressed

Benchmarking and Developing Segmentation Without Dense Initial Annotation Preconditions

Current semantic and instance segmentation methods rely on strict supervision preconditions, such as dense pixel-wise ground truth or accurate initial bounding-box-derived masks. When applied to domains where only weak annotations (such as image-level tags or sparse points) are available from scratch, these methods cannot be initialized or trained. Consequently, segmentation models cannot be deployed in annotation-scarce domains without first undertaking costly manual pixel-level or box-level labeling campaigns.

Possible candidate · 2/5 runs3 papers report this0% from 2025+

Quantization & Low-Precision Training

Barrier to removeOpen

Post-Training Realization of Training-Time Quantization Objectives on Fixed Pre-Trained Models

High-performance low-precision quantization techniques depend heavily on active training cycles, relying on mechanisms such as iterative knowledge distillation, multi-stage noise optimization, and auxiliary loss terms (e.g., kurtosis minimization). Because these methods require backpropagation through the model and access to full training infrastructure, they cannot be applied to fixed, pre-trained checkpoints, closed-source models, or settings where original training pipelines and datasets are inaccessible. Practitioners constrained to frozen models are therefore blocked from utilizing these advanced quantization objectives and must fall back to basic post-training heuristics. Bridging this gap enables applying sophisticated distribution-shaping and distillation objectives directly to fixed models.

Possible candidate · 2/5 runs3 papers report this33% from 2025+

Sentiment & Emotion Analysis

UnclassifiedOpen

Cross-Lingual Robustness and Generalizability of Empathetic Dialogue Systems

Current empathetic dialogue and emotion analysis models are evaluated almost exclusively on English-only datasets such as EMPATHETICDIALOGUES. Because existing methods have not been benchmarked across other languages, non-English domains, or diverse cultural contexts, it is unknown whether reported performance gains transfer beyond standard English benchmarks. Without multi-setting evaluations, researchers cannot determine whether existing empathy modeling techniques reflect general emotional reasoning capabilities or are overfitted to English linguistic conventions.

Possible candidate · 2/5 runs3 papers report this33% from 2025+

Language Agents & Tool Use

Barrier to removePartly addressed

Benchmark Verification for Language Agents Without Ground-Truth Test Suites or Reference Repositories

Current benchmarks for code-generating and paper-reproducing language agents rely strictly on pre-existing public GitHub repositories with functional, high-coverage unit tests as ground truth. This filters out the vast majority of scientific literature and closed-source software where code is missing, incomplete, or untested. As a result, agent capabilities cannot be benchmarked on genuine zero-code reproduction tasks or messy real-world codebases without human experts manually engineering executable harnesses.

Possible candidate · 2/5 runs3 papers report this67% from 2025+

Out-of-Distribution & Anomaly Detection

Scope to testPartly addressed

Cross-Modality Robustness and Evaluation of Anomaly Detection Methods Across Tabular and Unstructured Data

Anomaly and out-of-distribution (OOD) detection algorithms are developed in modality-specific silos, tested either exclusively on tabular benchmarks or on high-dimensional unstructured data relying on pre-trained vision/language foundation models. Because no unified empirical evaluation assesses these algorithms across both tabular and unstructured modalities, practitioners cannot determine whether tabular anomaly detectors fail on raw high-dimensional inputs or whether embedding-dependent detectors degrade when pre-trained feature extractors are absent. As a consequence, safety-critical domains with heterogeneous or non-standard data types lack validated guidance on algorithm selection.

Possible candidate · 2/5 runs3 papers report this33% from 2025+

Multi-Agent & Goal-Conditioned RL

Effect to explainPartly addressed

Policy Generalization Across Large Agent-Count Shifts in Multi-Agent Reinforcement Learning

Multi-agent and goal-conditioned reinforcement learning policies trained on small team sizes fail catastrophically when evaluated on larger agent counts, collapsing to near-zero success rates on unseen large-scale tasks. Existing architectures and skill representations fail to capture distribution shifts under scale, forcing practitioners to retrain separate policies from scratch for every team size. This inability to transfer learned coordination across team scales prevents the deployment of multi-agent policies to dynamic environments where entity counts vary at runtime.

Possible candidate · 2/5 runs3 papers report this100% from 2025+

Semantic & Instance Segmentation

Barrier to removePartly addressed

Hierarchical Layout and Vector Segmentation Under Incomplete Multi-Level Supervision

Current instance and semantic segmentation methods for structured documents and vector graphics rely on full, multi-tier hierarchical annotations (simultaneously requiring word, line, and paragraph labels) and fail when vector elements are deeply nested, tangled, or overlapping. Because multi-level hierarchical annotations are significantly more expensive than standard single-level bounding boxes, deploying segmentation pipelines to diverse real-world documents remains cost-prohibitive. Furthermore, the inability to parse overlapping and nested primitives prevents reliable digitization of technical figures, CAD drawings, and complex vector assets.

Possible candidate · 2/5 runs3 papers report this33% from 2025+

Knowledge & Dataset Distillation

Barrier to removePartly addressed

Knowledge Distillation and Rollout Transfer from Black-Box Teachers Without Logit Access

High-performing knowledge distillation and policy transfer techniques rely on white-box teacher access to evaluate conditional entropy, top-$k$ predictive distributions, and log-probabilities during rollouts. This requirement structurally prevents distillation from state-of-the-art proprietary or API-only models that output only discrete decisions or text tokens without internal logits. Furthermore, existing workarounds often rely on fragile parametric assumptions, such as Gaussianity of teacher logits under data augmentation, which fail on complex, multi-modal output distributions.

Possible candidate · 2/5 runs3 papers report this67% from 2025+

Question Answering & Reading Comprehension

Scope to testPartly addressed

Multilingual Robustness of Table Question Answering Models

Current table question answering benchmarks and systems predominantly evaluate questions, tables, and associated text in a single language—typically English or Simplified Chinese. As a result, it is unknown how robust existing table QA architectures are when queries, table contents, or schema descriptions span multiple languages or differ linguistically from one another. Without cross-lingual and multilingual evaluation, practitioners cannot determine whether standard table parsing, semantic parsing, and retrieval-augmented methods generalize to non-English structured data or fail due to language mismatch.

Possible candidate · 2/5 runs3 papers report this67% from 2025+

Offline & Model-Based RL

Scope to testOpen

Empirical Robustness of Model-Based RL Across Simulator Differentiability, Parallelism, and Task Horizons

Current model-based RL and planning algorithms are developed under disjoint simulation assumptions, variously requiring fully differentiable GPU physics engines, massive parallel batch environments, or short-horizon dense rewards. Because these methods are rarely tested outside their native simulator regimes, practitioners cannot determine whether algorithms relying on analytic dynamics or extreme parallel throughput transfer to standard non-differentiable environments or long-horizon agentic tasks. This leaves the empirical boundaries and failure modes of existing model-based approaches unknown across differing simulator capabilities.

Possible candidate · 2/5 runs3 papers report this67% from 2025+

Watermarking & Deepfake Detection

Scope to testPartly addressed

Cross-Paradigm Robustness Benchmarking for Deepfake Detectors

Current deepfake detectors are developed and validated within narrow manipulation regimes, such as full-frame image synthesis, boundary-blended face swaps, or identity-conditioned video pairs. Because detection models rely on artifacts specific to their target generation family, their performance characteristics across alternative forgery paradigms—such as expression reenactment, localized facial attribute edits, and full-image generation—remain untested. Consequently, practitioners cannot determine whether existing detectors offer any protection outside their specific training domain or if they fail completely when deployed against unmodeled manipulation types.

Possible candidate · 2/5 runs3 papers report this33% from 2025+

Word Embeddings & Lexical Semantics

Barrier to removePartly addressed

Cross-Lingual and Natural-Distribution Robustness Benchmarking for Paraphrase-Invariant Semantic Models

Existing paraphrase-dependent representations and watermark defense techniques rely strictly on multi-million-sentence English datasets like ParaBank2 and have only been verified against synthetic paraphrasers such as T5. Consequently, it remains unknown whether these methods retain semantic invariance under diverse, natural human paraphrases or across non-English languages lacking massive parallel corpora. Practitioners cannot deploy these semantic models in multilingual or low-resource settings because their behavior outside synthetic English text distributions is entirely uncharacterized.

Possible candidate · 2/5 runs3 papers report this0% from 2025+

Low-Resource Languages & Dialects

Barrier to removeOpen

Metric Inflation and Reliability Limits in IPA-Space Generation Evaluation

Multilingual systems increasingly evaluate generations in International Phonetic Alphabet (IPA) space to bridge script disparities, yet standard character-level metrics like chrF inflate scores because IPA relies heavily on shared Latin base characters. Furthermore, automatic grapheme-to-phoneme tools (such as Epitran) introduce systematic transcription errors and inventory mismatches that are masked when evaluating against synthetic or machine-translated benchmarks. Researchers currently cannot distinguish whether high IPA-level chrF gains reflect genuine semantic translation quality or superficial phonetic transcription artifacts.

Possible candidate · 2/5 runs3 papers report this100% from 2025+

Graph Neural Networks & Node Classification

Barrier to removePartly addressed

Architecture-Agnostic Black-Box Explanation for Graph Neural Networks

Many current GNN explainers require differentiable access to intermediate node embeddings, meaning they cannot operate under true black-box conditions such as inference-only APIs or proprietary model deployments. Because these methods are tied to specific encoder representations, their explanations fail to generalize across diverse GNN architectures and degrade when underlying embeddings are noisy or poorly calibrated. Practitioners with black-box access to graph classification services are currently blocked from generating post-hoc explanations without access to internal model parameters and activations.

Possible candidate · 2/5 runs3 papers report this0% from 2025+

Person Re-Identification & Face Recognition

Barrier to removeOpen

Cross-Encoder and Target-Free Robustness Benchmarking for Facial Representations

Current facial representation and transfer methods rely on training-time preconditions that frequently fail in real-world deployment. Specifically, techniques depend on collecting unlabeled target-domain demographic data—which is often legally or ethically prohibited—or evaluate strictly against known, seen feature extractors within face-only datasets. When the deployed encoder is changed or target data cannot be sampled in advance, systems experience substantial performance drops that remain unquantified across modern architectures. Without a unified benchmark testing methods under zero-target-data and unseen-encoder constraints, practitioners cannot determine which representation strategies legitimately transfer.

Possible candidate · 2/5 runs3 papers report this33% from 2025+

Language Model Evaluation & Benchmarking

Barrier to removePartly addressed

Multi-Model Evaluation and Ensembling Without Per-Step White-Box Logit Access

Current multi-model evaluation and collaborative generation frameworks require synchronous, step-by-step execution across multiple 32B–70B open-source LLMs with direct access to intermediate token logits. This white-box precondition structurally excludes closed-source, API-only models (which do not expose unconstrained per-step logits) from participating in these evaluation pipelines. It also prevents deployment on resource-constrained hardware unable to fit multiple large model footprints into memory simultaneously. Consequently, these benchmarking and decoding techniques cannot be applied to leading proprietary models or decentralized, asynchronous workflows.

Possible candidate · 2/5 runs3 papers report this100% from 2025+

Equivariance & Symmetry in Neural Networks

Effect to explainPartly addressed

Benchmarking Equivariance Degradation Across Discretization, Domain Shift, and Hybrid Pipelines

Exact equivariant architectures suffer numerical degradation when implemented on discretized grids, while learned approximate equivariance fails to extrapolate under test-time distribution shifts. Furthermore, inserting non-equivariant intermediate layers (such as high-dimensional feature GCNs) into geometric backbones breaks end-to-end SE(3) equivariance, risking pipeline instability. Because no systematic study evaluates these distinct methods and failure modes across shared out-of-distribution and discretized settings, practitioners have no principled way to know whether exact algebraic constraints, learned symmetries, or hybrid pipelines will retain their geometric guarantees in deployment.

Possible candidate · 1/2 runs3 papers report this67% from 2025+

Algorithmic Fairness

UnclassifiedOpen

Robustness and Sensitivity of Fairness Interventions Under Misspecified Merit Distributions and Latent Parameters

Many algorithmic fairness frameworks rely on theoretical assumptions that cannot be verified in deployment, specifically requiring known ground-truth merit distributions, unobservable transfer constants, or rigid pre-specified group partitions. In real-world applications, practitioners have access only to observed proxy labels and imperfect covariates, leaving the true merit distributions and transfer constants unknown and unbounded. Consequently, when fairness interventions are calibrated against estimated or misspecified latent parameters, theoretical parity guarantees can break down unpredictably, risking both utility loss and unintended disparity.

Possible candidate · 1/2 runs3 papers report this33% from 2025+

Molecular Modeling & Drug Design

Scope to testPartly addressed

Benchmarking Cross-System and Cross-Ensemble Generalization in Molecular Generative and Transition Models

Current deep learning models for molecular transitions, ensemble generation, and structure-based drug design are routinely trained and evaluated on isolated systems, fixed thermodynamic conditions, or homogeneous configuration spaces. When deployed on unseen molecular systems, new temperature/pressure regimes, novel pairs of meta-stable states, or shifted configuration spaces (such as moving from isolated ligands to pocket environments), these models require costly de novo retraining. Because prior studies only evaluate models within their narrow training distribution, the extent to which existing architectures can transfer across these varying physical and structural conditions remains unknown.

Weak candidate · 1/3 runs3 papers report this33% from 2025+

Hand-Object Interaction & Affordance

Barrier to removePartly addressed

Dynamic Closed-Loop Execution and Task-Conditioned Hand-Object Manipulation

Current hand-object interaction and grasp synthesis frameworks evaluate grasp success almost exclusively as static contact configurations generated from full, pre-captured object point clouds. In dynamic deployment, these methods suffer severe rotational tracking drift and mid-trajectory drops because visual observations during motion are partial, occluded, and streaming rather than static and complete. Consequently, high static grasp success rates fail to translate into functional multi-step manipulation or object placement in the wild.

Weak candidate · 1/4 runs3 papers report this33% from 2025+

3D Gaussian Splatting & Novel View Synthesis

Barrier to removeOpen

Robust Wide-Baseline 3D Gaussian Splatting Under Unreliable Foundation-Model Geometry Priors

Modern feed-forward 3D Gaussian Splatting pipelines rely on pairwise foundation models (such as MASt3R or VGGT) to supply initial coordinate pointmaps and relative camera poses. Because these foundation predictors degrade significantly as angular view differences increase and inject persistent geometric noise, downstream view synthesis fails when views are not captured in dense, small-step sequences. Treating these imperfect upstream predictions as rigid pseudo-ground truth structurally blocks feed-forward novel view synthesis on sparse, wide-baseline image collections.

Weak candidate · 1/3 runs3 papers report this100% from 2025+

3D Shape Modeling & Surface Reconstruction

UnclassifiedOpen

Cross-Category and Vocabulary Generalization Benchmark for Text-to-3D Shape Models

Current text-driven 3D shape generation and modeling methods are predominantly trained and benchmarked strictly on Text2Shape, which is limited to chairs and tables (~11.5k shapes and a ~3.6k-word vocabulary). Consequently, it is unknown whether existing architectural designs, text encodings, and shape priors maintain generative fidelity when exposed to out-of-domain object categories or open-vocabulary text prompts. Without evaluating existing methods across broader category and vocabulary distributions, the field cannot determine if current performance gains reflect generalizable shape modeling or dataset-specific memorization.

Weak candidate · 1/4 runs3 papers report this0% from 2025+

Image Restoration & Super-Resolution

Effect to explainOpen

Cross-Domain Robustness Audit of Image Restoration Models Under Physical Sensor and Optical Degradations

Image restoration and super-resolution models are primarily developed and evaluated under synthetic degradation models, including bicubic downsampling, Wald-protocol Gaussian blur, and simplified sensor calibrations. Consequently, it remains unknown how well these methods transfer to actual physical captures, such as high-magnification (8×/16×) microscopy and real-world non-uniform sensor acquisitions. Practitioners deploying models to physical instruments cannot anticipate whether synthetic-trained models will retain their reported performance or fail catastrophically. Without a systematic cross-setting robustness evaluation, the domain-gap penalty between synthetic simulations and real optical sensors remains unquantified.

Weak candidate · 1/3 runs3 papers report this67% from 2025+