Open Problems

Causal Inference & Discovery

Benchmarking Causal Effect Estimation Across Multi-Valued and Continuous Treatment Regimes

Scope to testPartly addressed
Possible candidate · 3/5 runs10 papers report this60% from 2025+

Generated automatically from the limitations stated in 10 papers (ICML, ICLR, AISTATS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

The vast majority of causal inference and discovery algorithms are formulated and evaluated strictly under binary treatment assumptions ($T \in \{0, 1\}$). In real-world domains such as medicine and econometrics, interventions are routinely dose-dependent, multi-categorical, or continuous, rendering binary-only methods inapplicable without ad-hoc binarization. Because existing methods have not been systematically evaluated across non-binary treatment spaces, practitioners lack empirical evidence on how severely estimation and discovery performance degrade when moving beyond the binary regime.

Why it matters

Provides the first standardized empirical characterization of estimator performance across non-binary treatment types, enabling practitioners to select appropriate causal estimators for dose-response and multi-arm intervention settings.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Systematic Empirical Benchmark: Implement standard causal effect estimators (e.g., propensity-based, meta-learners, doubly robust methods) across synthetic and semi-synthetic benchmarks featuring graded multi-valued and continuous treatment variables, measuring mean squared error on individual and average dose-response curves.

  2. 2

    Discretization vs. Direct Continuous Modeling: Conduct a comparative study evaluating naive discretization schemes (binning continuous treatments into binary/multi-class bins) against continuous treatment estimators (e.g., generalized propensity score methods, kernel estimators) across varying sample sizes and confounding strengths, measuring estimation bias and variance.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Standard continuous treatment estimators (e.g., generalized propensity scores or continuous causal forests) may already be well-established in the econometrics literature, leaving the machine learning benchmark without sufficient novel empirical surprise if standard tools trivially transfer.

Evidence

Each paper's own statement of the limitation, verbatim.

Show all 10 papers

Nearest existing work

Related open problems

Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.