Open Problems

Deep Learning Theory & Optimization Dynamics

Extending Two-Layer Network Theory to Deep and Structured Architectures

Scope to testPartly addressed
Weak candidate · 1/4 runs8 papers report this40% from 2025+

Generated automatically from the limitations stated in 8 papers (NeurIPS, COLT, ICML), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

A substantial body of rigorous results—on training dynamics, predictor statistics, generalization, and under-fitting—is proven only for one- or two-layer fully-connected networks, often with fixed or special activations. Practitioners use deep residual, convolutional, and attention models, so none of these guarantees can be checked against the systems that matter. Until the theory carries over, each paper's conclusions remain unvalidated in practice, and the specific mechanisms they identify may be artifacts of the shallow, unstructured setting.

Why it matters

The existing shallow-network results become usable predictions about the deep architectures actually deployed, and the point-of-failure analysis itself becomes a research tool.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Layer-by-layer induction. Take an existing two-layer result whose proof factors through the feature kernel and attempt to propagate it through depth by showing the required statistics stabilize under repeated composition; measure how tightly the per-layer bounds compose (do they blow up exponentially, linearly, or hold?). A student can start within a month on depth-3–10 ReLU/softplus networks on synthetic data.

  2. 2

    Fix-then-release the second layer. Several statements block on a fixed top layer; study the same models with the second layer trained, measuring whether the quantities the theory tracks (kernel spectra, predictor statistics) are stable under training, as a prerequisite for any deep guarantee.

  3. 3

    Structured architectures as weight-sharing constraints. Treat convolution/attention as linear constraints on an otherwise fully-connected weight space and test whether the existing proofs survive under those constraints, measuring where the argument first breaks.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

The induction genuinely fails: the quantities tracked in two layers are simply not stable under composition, so each additional depth layer requires a new analysis and no unified result exists.

Sub-problems

  • Extending Optimization Dynamics and Representation Guarantees of Two-Layer ReLU Networks to Multivariate Inputs

    Theoretical analyses of optimization dynamics and representation in shallow ReLU networks currently depend on the total ordering and simple geometry of one-dimensional inputs and fixed internal weights. As a consequence, existing proofs cannot predict whether training dynamics, underfitting phenomena, or convergence rates persist when inputs span multiple dimensions with non-trivial covariance structures. Extending these analyses beyond univariate settings is necessary to determine if current theoretical explanations capture general deep learning behavior or are merely artifacts of 1D scalar geometry.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Deep Learning Theory & Optimization Dynamics

Barrier to removePartly addressed

Bridging the Activation Divide in Neural Network Optimization Dynamics Theory

Current theoretical analyses of neural network training dynamics are split across incompatible mathematical preconditions: one major body of literature strictly requires high-order smoothness ($C^2$, $C^\infty$, bounded derivatives, or odd symmetry), which explicitly excludes standard piecewise-linear activations like ReLU, while another body relies strictly on piecewise linearity and positive homogeneity, failing to transfer to smooth activations like GELU, SiLU, or tanh. As a result, foundational theoretical predictions—such as edge-of-stability behavior, width-independent convergence rates, and representation learning bounds—cannot be generalized across activation families. Without analytical tools that either handle subgradient boundary crossings or bound the divergence between smooth approximations and exact piecewise-linear trajectories during training, optimization theory remains fragmented into mutually exclusive activation regimes.

Possible candidate · 2/5 runs13 papers report this38% from 2025+

Deep Learning Theory & Optimization Dynamics

Scope to testPartly addressed

Cross-Architecture Empirical Validation of Deep Learning Theory Predictors

Theoretical analysis and optimization dynamic guarantees in deep learning are overwhelmingly developed and proved on shallow fully connected ReLU models or two-layer networks. As a result, theoretical claims regarding convergence rates, generalization bounds, and optimization statistics are rarely verified on non-feedforward or deeper architectures such as CNNs, ResNets, and Transformers. Without systematic empirical evaluation across these structural settings, it remains unknown whether these theoretical quantities correlate with actual training dynamics or are merely artifacts of idealized shallow feedforward assumptions.

Weak candidate · 1/4 runs6 papers report this17% from 2025+

Deep Learning Theory & Optimization Dynamics

Effect to explainOpen

Empirical Benchmarking of Generalization Failure in Alternative Optimization Dynamics

Alternative optimization methods, such as exact Gauss-Newton (GN) and non-backpropagation dynamics (e.g., NMNC), exhibit severe generalization gaps and early loss saturation when scaled to deep networks and mini-batch settings. Standard regularizers (dropout, weight decay, data augmentation, pseudoinverse regularization) and initialization heuristics developed on shallow models consistently fail to close these train-test gaps. Without a controlled comparative evaluation across these distinct settings, it remains unknown whether these generalization failures share common optimization dynamics or require fundamentally different stabilization interventions.

Possible candidate · 3/5 runs3 papers report this67% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.