Causal Inference & Discovery
Benchmarking Causal Effect Estimation Across Multi-Valued and Continuous Treatment Regimes
Generated automatically from the limitations stated in 10 papers (ICML, ICLR, AISTATS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
The vast majority of causal inference and discovery algorithms are formulated and evaluated strictly under binary treatment assumptions ($T \in \{0, 1\}$). In real-world domains such as medicine and econometrics, interventions are routinely dose-dependent, multi-categorical, or continuous, rendering binary-only methods inapplicable without ad-hoc binarization. Because existing methods have not been systematically evaluated across non-binary treatment spaces, practitioners lack empirical evidence on how severely estimation and discovery performance degrade when moving beyond the binary regime.
Why it matters
Provides the first standardized empirical characterization of estimator performance across non-binary treatment types, enabling practitioners to select appropriate causal estimators for dose-response and multi-arm intervention settings.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Systematic Empirical Benchmark: Implement standard causal effect estimators (e.g., propensity-based, meta-learners, doubly robust methods) across synthetic and semi-synthetic benchmarks featuring graded multi-valued and continuous treatment variables, measuring mean squared error on individual and average dose-response curves.
- 2
Discretization vs. Direct Continuous Modeling: Conduct a comparative study evaluating naive discretization schemes (binning continuous treatments into binary/multi-class bins) against continuous treatment estimators (e.g., generalized propensity score methods, kernel estimators) across varying sample sizes and confounding strengths, measuring estimation bias and variance.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Standard continuous treatment estimators (e.g., generalized propensity scores or continuous causal forests) may already be well-established in the econometrics literature, leaving the machine learning benchmark without sufficient novel empirical surprise if standard tools trivially transfer.
Evidence
Each paper's own statement of the limitation, verbatim.
- SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival AnalysisICLR 2026
Restricted to static binary treatments and fixed baseline covariates; excludes time-varying treatments, instrumental variables, longitudinal covariates, and multi-valued/continuous treatments
- Spiked-CFR: Causal Representation Learning from LLMs via Wasserstein Projection PursuitICML 2026
Restricted to binary treatments; multi-valued, continuous, and time-varying treatments are not handled
- $\epsilon$-Identifiability of Causal QuantitiesAISTATS 2026
Only binary treatment and outcome variables are treated; multi-valued extensions are deferred to a footnote reference, and only two graph-dependent conditions (Theorems 14–15) exploit causal structure
- Data Fusion for Partial Identification of Causal EffectsNeurIPS 2025
Only single experimental and single observational datasets, binary treatments, and static (non-dynamic) regimes are supported; continuous treatments and multiple studies are excluded
- Counterfactual Contrastive Learning with Normalizing Flows for Robust Treatment Effect EstimationICML 2025
Evaluated only on binary treatments with static, time-independent outcomes; the authors note extension to time-dependent outcome variables is future work
- CATE Estimation With Potential Outcome Imputation From Local RegressionUAI 2025
Restricted to binary treatment (T ∈ {0,1}); continuous or multi-valued treatments are not covered
- Tight Partial Identification of Causal Effects with Marginal Distribution of Unmeasured ConfoundersICML 2024
Restricted to discrete random variables for the treatment, outcome, and confounders.
- Empirical Analysis of Model Selection for Heterogeneous Causal Effect EstimationICLR 2024
Only binary treatments W ∈ {0,1} are considered; continuous or multi-valued treatments are out of scope
Show all 10 papers
- Towards Balanced Representation Learning for Credit Policy EvaluationAISTATS 2023
Restricted to the potential outcomes framework with binary treatment; multi-valued or continuous treatments are excluded
- Optimal Treatment Regimes for Proximal Causal LearningNeurIPS 2023
Assumes a binary treatment space, limiting applicability to multi-valued or continuous intervention settings.
Nearest existing work
- Partial identification of dose responses with hidden confoundersUAI 2023
- Answering Complex Causal Queries With the Maximum Causal Set EffectNeurIPS 2021
- Scalable Sensitivity and Uncertainty Analyses for Causal-Effect Estimates of Continuous-Valued InterventionsNeurIPS 2022
- Classifying Treatment Responders Under Causal Effect MonotonicityICML 2019
- Disentangling Causal Effects from Sets of Interventions in the Presence of Unobserved ConfoundersNeurIPS 2022
- Hierarchical Bias-Driven Stratification for Interpretable Causal Effect EstimationAISTATS 2025
- Learning Joint Interventional Effects from Single-Variable Interventions in Additive ModelsICML 2025
- Continuous Treatment Effects with Surrogate OutcomesICML 2024
- Probabilities of Causation for Continuous and Vector VariablesUAI 2024
- Sharp Bounds for Generalized Causal Sensitivity AnalysisNeurIPS 2023
- Operationalizing Complex Causes: A Pragmatic View of MediationICML 2021
- Two-Stage Shadow Inclusion Estimation: An IV Approach for Causal Inference under Latent Confounding and Collider BiasICML 2024
- End-to-End Balancing for Causal Continuous Treatment-Effect EstimationICML 2022
- Double Debiased Machine Learning for Mediation Analysis with Continuous TreatmentsAISTATS 2025
- Generalization Bounds for Estimating Causal Effects of Continuous TreatmentsNeurIPS 2022