Multimodal Fusion & Multi-Task Learning
Empirical Robustness and Scalability Benchmarking for Multi-View Multimodal Fusion
Generated automatically from the limitations stated in 8 papers (CVPR, NeurIPS, ICLR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current multi-view and multimodal fusion methods are evaluated almost exclusively on small benchmarks (typically under 10,000 samples) with very few modalities or views (2 to 6 views). Because scaling behavior across large sample counts ($N \gg 10^4$) and high view counts ($V > 6$) remains untested, it is unknown whether reported algorithmic advantages persist, collapse, or become computationally intractable in larger settings. Practitioners cannot reliably select fusion algorithms for complex multi-sensor or high-view applications because no empirical baseline exists across these broader operating conditions.
Why it matters
Provides the first empirical ground truth on how multi-view fusion algorithms perform and scale when pushed beyond small benchmark limits. Enables practitioners to select robust fusion architectures for large-scale, high-view deployments based on demonstrated empirical behavior.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Unified Scaling Testbed: Construct a standardized evaluation pipeline containing top-performing multi-view fusion paradigms (e.g., cross-view alignment, tensor decomposition, and deep shared-space encoders) and evaluate them on datasets scaled from $V=2$ to $V=30$ views and $N=10^3$ to $N=500,000$ samples. Measure task accuracy (classification/clustering), memory footprints, and wall-clock training time against theoretical complexity bounds.
- 2
View-Robustness and Missing-View Stress Testing: Systematically vary the number of available views and view-to-sample ratios on high-dimensional benchmarks to evaluate model stability. Measure the performance delta and ranking stability of methods as view count scales upward.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If standard baseline implementations encounter out-of-memory errors that prevent running full multi-view baselines without bespoke algorithmic re-engineering, completing the benchmark matrix may prove intractable. Additionally, if the scaling behavior across all methods simply mirrors trivial theoretical parameter complexity without unexpected performance shifts, the empirical contribution would be marginal.
Evidence
Each paper's own statement of the limitation, verbatim.
- Imbalanced View Contribution Evaluation and Refinement for Deep Incomplete Multi-View ClusteringCVPR 2026
Evaluation is on small benchmark datasets (600 to 10158 samples, 2-6 views); scalability to large or high-view-count data is not demonstrated
- Confident Block Diagonal Structure-Aware Invariable Graph Completion for Incomplete Multi-view ClusteringICLR 2026
Empirically evaluated only on relatively small benchmark subsets ($n \le 1,474$ samples).
- Beyond Independence: Learning Correlated Views for Variational Incomplete Multi-View ClusteringICML 2026
Evaluation is limited to small/medium benchmarks (600–70,000 samples, 2–6 views); scaling behavior to many views is only indirectly suggested by the V(V−1) deviation bound and quadratic parameter growth in view count
- SparseMVC: Probing Cross-view Sparsity Variations for Multi-view ClusteringNeurIPS 2025
Evaluation covers small-to-medium datasets (up to 10,800 samples, at most 4 views in most cases), leaving very high view counts or large-scale data untested
- Incomplete Multi-view Clustering via Hierarchical Semantic Alignment and Cooperative CompletionNeurIPS 2025
Evaluated only on relatively small datasets (up to 10,000 samples); authors acknowledge need for scalability improvements in large-scale real-world applications
- Differentiable Information Bottleneck for Deterministic Multi-view ClusteringCVPR 2024
Evaluated only on small-to-mid-size benchmark datasets with a fixed number of views; no large-scale or high-view-count experiments
- MetaViewer: Towards a Unified Multi-View RepresentationCVPR 2023
Gains are on small benchmarks (2,000-10,158 samples, 2-6 views); scalability to large datasets or many views is not demonstrated
- Cross-view Topology Based Consistent and Complementary Information for Deep Multi-view ClusteringICCV 2023
Gains are shown only on five small benchmark datasets (≤10,800 samples, ≤100 clusters); behavior on high-dimensional or many-view real-world data is untested.
Nearest existing work
- Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal PerceptionNeurIPS 2023
- Efficient Low-rank Multimodal Fusion With Modality-Specific FactorsACL 2018
- Provable Dynamic Fusion for Low-Quality Multimodal DataICML 2023
- Data-Efficient Multimodal Fusion on a Single GPUCVPR 2024
- Beyond Magnitude: Scale-Invariant Evidential Fusion for Multi-View ClassificationICML 2026
- What Makes Multi-Modal Learning Better than Single (Provably)NeurIPS 2021
- CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation ModelsICML 2025
- Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated PerturbationICCV 2025
- MultiMoDN—Multimodal, Multi-Task, Interpretable Modular NetworksNeurIPS 2023
- Multimodal Nested Learning for Decoupled and Coordinated OptimizationICML 2026
- SiMO: Single-Modality-Operable Multimodal Collaborative PerceptionICLR 2026
- Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in BiomedicineICLR 2025
- CoRiM: Conflict-driven Risk Minimization for Dynamic Multimodal FusionCVPR 2026
- MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-TuningICML 2026
- Predictive Dynamic FusionICML 2024