Open Problems

Continual Learning & Catastrophic Forgetting

Sub-Quadratic Covariance Sketching for Exemplar-Free Continual Learning

UnclassifiedOpen
Possible candidate · 3/5 runs6 papers report this50% from 2025+

Generated automatically from the limitations stated in 6 papers (CVPR, ICML, NeurIPS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Exemplar-free continual learning methods increasingly rely on storing per-class feature statistics (prototypes and full covariance matrices) to generate pseudo-replay features or parameterize Mahalanobis classifiers without retaining raw inputs. However, storing a dense $D \times D$ covariance matrix for each of $C$ classes incurs $\mathcal{O}(C \cdot D^2)$ storage, which quickly exceeds device memory when scaling to large label spaces ($C \ge 1000$) or modern high-dimensional embeddings ($D \ge 1024$). Practitioners currently resort to naive diagonal approximations that discard inter-feature correlations and cause severe classification degradation under continual distribution shifts. As a result, covariance-based exemplar-free learning is structurally blocked from scaling to realistic large-vocabulary benchmarks on memory-constrained hardware.

Why it matters

Enables exemplar-free and statistical-replay continual learning to scale to tens of thousands of classes and high-dimensional foundation model backbones on memory-constrained edge hardware.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Low-Rank Plus Shared-Basis Covariance Decomposition: Model class covariances as a rank-$k$ class-specific factor ($k \ll D$) over a shared global subspace plus diagonal residual noise; measure total parameter storage (in megabytes) against incremental accuracy on Split ImageNet-1k.

  2. 2

    Streaming Frequent Directions for Class Statistics: Adapt deterministic streaming sketching to update compact low-rank covariance sketches online as each class stream arrives; measure relative spectral norm error ($\|\Sigma_c - \hat{\Sigma}_c\|_2$) and task-retention performance relative to full-covariance baselines.

  3. 3

    Kronecker / Block-Structured Covariance Approximations: Exploit layer-wise or channel-group structures in pretrained Vision Transformer embeddings to store class covariances as compact Kronecker products; measure classification accuracy and peak memory across ViT and ConvNeXt feature backbones.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

If frozen global covariance matrices or fixed feature extractors from large foundation models reduce intra-class covariance variations to negligible noise, class-specific covariance storage would become entirely redundant.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Continual Learning & Catastrophic Forgetting

Scope to testPartly addressed

A Standardized Transfer-Validation Benchmark for Methods Validated Only on Small-Scale Image Classification

Nearly every training-time method in the modern literature—continual learning, distillation, unlearning, robustness, federated aggregation—is validated exclusively on CIFAR/ImageNet-subset image classification with ResNet-18/50-class backbones, and each paper explicitly defers scale, architecture, and modality transfer to future work. Because every paper defers this simultaneously, the field has no evidence base at all for whether any of these methods survive the shift to transformers, higher resolutions, dense tasks, or non-vision data; practitioners choosing between methods are choosing based on measurements made in a regime none of them deploy in. The barrier is not one method's limitation but a collective, unpriced evaluation debt: nobody can tell which small-scale conclusions are regime-robust and which are artifacts of ResNet-18 on 32×32 images.

Possible candidate · 3/5 runs152 papers report this44% from 2025+

Continual Learning & Catastrophic Forgetting

Scope to testOpen

Characterizing and Reducing Pre-Trained ViT Dependence in Transfer-Based Methods

A wide range of methods built on frozen ImageNet-pretrained ViT features—whether for adapting to new tasks, constraining representations, or detecting anomalies—inherit an unexamined dependency: their guarantees hold only when the backbone's feature space is already good. Today nobody knows whether these methods degrade gracefully with weaker pretraining, generalize to domains far from ImageNet, or transfer to non-transformer architectures, because the dependence is never ablated. The consequence is that reported gains may be properties of the backbone, not of the methods themselves, and the methods are silently inapplicable wherever a strong ImageNet ViT does not exist.

Possible candidate · 3/5 runs20 papers report this78% from 2025+

Continual Learning & Catastrophic Forgetting

Barrier to removePartly addressed

Task-Agnostic Inference for Modular and Masked Continual Learning Architectures

Modular continual learning methods prevent catastrophic forgetting by allocating task-specific masks, sub-networks, or adaptive parameters during training. However, these methods structurally depend on having ground-truth task identifiers provided at test time to select the corresponding parameters. In practical deployments, incoming inputs arrive without task metadata, rendering task-dependent parameter masking unusable in class-incremental or task-agnostic settings. Without a reliable mechanism to resolve task identity or route inputs dynamically at test time, parameter-isolation architectures remain restricted to artificial task-incremental benchmarks.

Possible candidate · 2/5 runs3 papers report this0% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.