<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>doubly robust | Chen Xing</title>
    <link>https://chenxing.space/tag/doubly-robust/</link>
      <atom:link href="https://chenxing.space/tag/doubly-robust/index.xml" rel="self" type="application/rss+xml" />
    <description>doubly robust</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Tue, 17 Mar 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>doubly robust</title>
      <link>https://chenxing.space/tag/doubly-robust/</link>
    </image>
    
    <item>
      <title>Notes on Doubly Robust Censoring Unbiased Transformation</title>
      <link>https://chenxing.space/blog/notes-on-doubly-robust-censoring-unbiased-transformation/</link>
      <pubDate>Tue, 17 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-doubly-robust-censoring-unbiased-transformation/</guid>
      <description>&lt;p&gt;Predicting outcomes with right-censored survival data forces a choice: do we model the outcome distribution, or the censoring mechanism? Classical transformations require us to commit to one. Rubin and van der Laan (2007) tell us we don&amp;rsquo;t have to. Their &lt;mark&gt;&lt;strong&gt;doubly robust censoring unbiased transformation&lt;/strong&gt;&lt;/mark&gt; fuses both approaches, remaining valid as long as &lt;em&gt;at least one&lt;/em&gt; of the two nuisance models is correctly specified. This post walks through the setup, the classical transformations, and how the doubly robust version combines them — drawing the analogy to AIPW along the way.&lt;/p&gt;
&lt;h2 id=&#34;1-setup&#34;&gt;1. Setup&lt;/h2&gt;
&lt;p&gt;We observe an i.i.d. sample  $\{O_i\}_{i=1}^n$  where each observation is&lt;/p&gt;
 $$O = \bigl(W,\; \Delta = \mathbf{1}(Y \le C),\; \tilde{Y} = Y \wedge C\bigr)$$ 
&lt;ul&gt;
&lt;li&gt;$W$: covariates&lt;/li&gt;
&lt;li&gt;$Y$: true (possibly unobserved) survival time&lt;/li&gt;
&lt;li&gt;$C$: random censoring time&lt;/li&gt;
&lt;li&gt;$\tilde{Y} = \min(Y, C)$ is what we actually record&lt;/li&gt;
&lt;li&gt;$\Delta = 1$ if the event is observed ($Y \le C$), and $\Delta = 0$ if the outcome is right-censored ($Y &amp;gt; C$)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Our goal&lt;/strong&gt; is to estimate the regression function&lt;/p&gt;
&lt;p&gt;$$m(w) = \mathbb{E}[Y \mid W = w]$$&lt;/p&gt;
&lt;p&gt;Two nuisance functions appear throughout:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$\bar{F}(\cdot \mid W)$: conditional survival function of the response $Y$ given $W$&lt;/li&gt;
&lt;li&gt;$\bar{G}(\cdot \mid W)$: conditional survival function of the censoring time $C$ given $W$&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We maintain the standard assumption $Y \indep C \mid W$.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;2-the-challenge-unidentifiability&#34;&gt;2. The Challenge: Unidentifiability&lt;/h2&gt;
&lt;p&gt;When the censoring time $C$ corresponds to a fixed study endpoint, the true response $Y$ may exceed it. Beyond that endpoint, &lt;strong&gt;nothing can be learned&lt;/strong&gt; about the tail of the survival distribution — making $m(W) = \mathbb{E}[Y \mid W]$ unidentifiable in general.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; truncate the response at a known horizon $\tau$,&lt;/p&gt;
&lt;p&gt;$$Y \longmapsto Y \wedge \tau = \min(Y, \tau)$$&lt;/p&gt;
&lt;p&gt;and estimate $w \mapsto \mathbb{E}[Y \wedge \tau \mid W = w]$ instead.&lt;/p&gt;
&lt;h3 id=&#34;the-surrogate-response-strategy&#34;&gt;The Surrogate-Response Strategy&lt;/h3&gt;
&lt;p&gt;The general approach to prediction with right-censored data is:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Replace&lt;/strong&gt; the possibly unavailable responses  $\{Y_i\}_{i=1}^n$  with surrogate values  $\{Y^*(O_i)\}_{i=1}^n$  using an imputation map $Y^*(\cdot)$ built from observed data.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Plug&lt;/strong&gt; the imputed dataset  $\{W_i, Y^*(O_i)\}_{i=1}^n$  into any standard regression algorithm.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The imputation map $Y^*(\cdot)$ is called a &lt;mark&gt;&lt;strong&gt;censoring unbiased transformation&lt;/strong&gt;&lt;/mark&gt; (Fan and Gijbels 1996) if it satisfies&lt;/p&gt;
&lt;p&gt;$$\mathbb{E}[Y^*(O) \mid W] = \mathbb{E}[Y \mid W] = m(W)$$&lt;/p&gt;
&lt;p&gt;That is, the surrogate is an unbiased proxy for the true (latent) response, conditional on covariates.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;3-two-classical-transformations&#34;&gt;3. Two Classical Transformations&lt;/h2&gt;
&lt;h3 id=&#34;a-the-buckleyjames-transformation-depends-on-barf&#34;&gt;A. The Buckley–James Transformation (depends on $\bar{F}$)&lt;/h3&gt;
&lt;p&gt;The Buckley–James transformation imputes a censored observation with its conditional mean given that it exceeds the censoring time:&lt;/p&gt;
&lt;p&gt;$$Y^*(O) = \Delta Y + (1-\Delta) Q_{\bar{F}}(W, C)$$&lt;/p&gt;
&lt;p&gt;where $\Delta = 1$ if $Y$ is observed and $\Delta = 0$ if right-censored, and&lt;/p&gt;
 $$Q_{\bar{F}}(w, y) = \mathbb{E}[Y \mid W = w,\; Y &gt; y] = \frac{1}{\bar{F}(y \mid W=w)} \int_y^{+\infty} u \; dF(u \mid W=w)$$ 
&lt;p&gt;Intuitively: if we observe the event, we keep $Y$; if censored at $C$, we impute with the expected remaining survival time above $C$.&lt;/p&gt;
&lt;p&gt;This transformation requires correctly estimating $\bar{F}(\cdot \mid W)$, i.e., the conditional distribution of the survival time.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id=&#34;b-the-ipcw-transformation-depends-on-barg&#34;&gt;B. The IPCW Transformation (depends on $\bar{G}$)&lt;/h3&gt;
&lt;p&gt;Inverse probability of censoring weighting (IPCW) up-weights the observed events to compensate for the censored ones:&lt;/p&gt;
&lt;p&gt;$$Y^*(O) = \frac{Y \Delta}{\bar{G}(Y \mid W)}$$&lt;/p&gt;
&lt;p&gt;This requires correctly estimating $\bar{G}(\cdot \mid W)$, the conditional survival function of the censoring time. It is the survival-analysis analogue of IPW in the causal inference literature.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;4-the-doubly-robust-censoring-unbiased-transformation&#34;&gt;4. The Doubly Robust Censoring Unbiased Transformation&lt;/h2&gt;
&lt;p&gt;The two classical transformations each stake everything on one nuisance model. The doubly robust approach combines them:&lt;/p&gt;

$$
Y^*(O) = \underbrace{\frac{Y\Delta}{\bar{G}(Y \mid W)}}_{\text{1st term}} + \underbrace{\frac{Q_{\bar{F}}(W,C)\,(1-\Delta)}{\bar{G}(C \mid W)}}_{\text{2nd term}} - \underbrace{\int_{-\infty}^{\tilde{Y}} \frac{Q_{\bar{F}}(W,c)}{\bar{G}^2(c \mid W)}\, dG(c \mid W)}_{\text{3rd term (correction)}}
$$

&lt;p&gt;where $\tilde{Y} = Y \wedge C = \min(Y,C)$ and $Q_{\bar{F}}(w, c) = \mathbb{E}[Y \mid W=w,; Y &amp;gt; c]$.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$\mathbb{E}[Y^*(O) \mid W] = \mathbb{E}[Y \mid W]$ whenever either $\bar{F}(\cdot \mid W)$ or $\bar{G}(\cdot \mid W)$ is correctly specified.

  &lt;/div&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id=&#34;5-intuition-an-aipw-in-disguise&#34;&gt;5. Intuition: An AIPW in Disguise&lt;/h2&gt;
&lt;p&gt;The three-term structure has a clean interpretation. Recall that $Q_{\bar{F}}(W,C) = \mathbb{E}[Y \mid W, Y &amp;gt; C]$ is the &lt;strong&gt;outcome regression&lt;/strong&gt; for censored individuals.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Term&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1st:&lt;/strong&gt; $Y\Delta \,/\, \bar{G}(Y\!\mid\! W)$&lt;/td&gt;
&lt;td&gt;For &lt;em&gt;observed&lt;/em&gt; outcomes — apply IPCW, the analogue of inverse propensity score weighting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2nd:&lt;/strong&gt; $Q_{\bar{F}}(W,C)(1-\Delta) \,/\, \bar{G}(C\!\mid\! W)$&lt;/td&gt;
&lt;td&gt;For &lt;em&gt;censored&lt;/em&gt; outcomes — impute with the outcome regression $Q_{\bar{F}}$, then apply an IPCW-style weight&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3rd:&lt;/strong&gt; $-\int Q_{\bar{F}} / \bar{G}^2 \; dG$&lt;/td&gt;
&lt;td&gt;Bias &lt;strong&gt;correction&lt;/strong&gt; term&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The first two terms together look exactly like &lt;strong&gt;IPW + imputation&lt;/strong&gt; — the two ingredients of AIPW in the standard (binary treatment) setting. The third term is the survival-analysis counterpart of the augmentation correction in AIPW: it removes the bias that accumulates when both nuisance models are slightly off.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;takeaway&#34;&gt;Takeaway&lt;/h2&gt;
&lt;p&gt;The doubly robust censoring unbiased transformation is a drop-in replacement for the classical Buckley–James or IPCW transformations. Once $Y^*(O_i)$ is computed for each observation, any off-the-shelf regression algorithm can be applied to the pairs $\{W_i, Y^*(O_i)\}$.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Rubin, Daniel and Mark J. van der Laan (2007), &amp;ldquo;A Doubly Robust Censoring Unbiased Transformation,&amp;rdquo; &lt;em&gt;The International Journal of Biostatistics&lt;/em&gt;, 3 (1).&lt;/p&gt;
&lt;p&gt;Fan, Jianqing and Irène Gijbels (1996), &lt;em&gt;Local Polynomial Modelling and Its Applications&lt;/em&gt;, Chapman &amp;amp; Hall.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Calibrate Your Nuisances: A Simple Fix for Doubly Robust Inference</title>
      <link>https://chenxing.space/blog/calibrate-your-nuisances-a-simple-fix-for-doubly-robust-inference/</link>
      <pubDate>Fri, 09 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/calibrate-your-nuisances-a-simple-fix-for-doubly-robust-inference/</guid>
      <description>&lt;h2 id=&#34;doubly-robust-inference-via-calibration&#34;&gt;Doubly Robust Inference via Calibration&lt;/h2&gt;
&lt;h3 id=&#34;tldr&#34;&gt;TL;DR&lt;/h3&gt;
&lt;p&gt;Van der Laan, Luedtke, and Carone introduce &lt;strong&gt;&amp;ldquo;calibrated debiased machine learning&amp;rdquo; (calibrated DML)&lt;/strong&gt;, a method that achieves doubly robust asymptotic normality for causal inference estimators by simply adding an isotonic regression calibration step to standard DML pipelines. The key innovation is that valid inference requires only one of two nuisance functions (outcome regression or propensity score) to be estimated well—the other can converge arbitrarily slowly or even inconsistently. This bridges a long-standing gap where consistency was doubly robust but inference was not, and the method can be implemented by adding just a few lines of code to existing workflows.&lt;/p&gt;
&lt;h3 id=&#34;what-is-this-paper-about&#34;&gt;What is this paper about?&lt;/h3&gt;
&lt;p&gt;Doubly robust estimators like AIPW are popular for estimating average treatment effects because they remain consistent if either the outcome regression or propensity score is correctly specified. However, there&amp;rsquo;s a crucial asymmetry: &lt;mark&gt;while consistency requires only one nuisance to be correct, valid inference (asymptotic normality, correct confidence intervals) typically requires both nuisances to converge at sufficiently fast rates $(\approx n^{-1/4})$. When one nuisance is estimated poorly—whether due to model misspecification or slow convergence—standard confidence intervals can have incorrect coverage.&lt;/mark&gt; Prior solutions like the DR-TMLE framework required computationally intensive iterative procedures and case-by-case derivations for each new parameter.&lt;/p&gt;
&lt;h4 id=&#34;motivation&#34;&gt;Motivation&lt;/h4&gt;
&lt;p&gt;The authors ask: can we achieve doubly robust inference using a simple, general-purpose procedure that works with any machine learning estimator?&lt;/p&gt;
&lt;h3 id=&#34;what-do-the-authors-do&#34;&gt;What do the authors do?&lt;/h3&gt;
&lt;p&gt;The authors develop calibrated DML, a two-step procedure that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;step 1: takes cross-fitted nuisance estimators from any machine learning algorithm&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;step 2: calibrates them using isotonic regression before constructing the debiased estimator&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The calibration step ensures that nuisance estimates satisfy certain &lt;strong&gt;empirical orthogonality conditions&lt;/strong&gt; that linearize the bias term in the doubly robust expansion. Specifically, they calibrate the outcome regression using squared error loss and the Riesz representer (inverse propensity weights for ATE) using a tailored Riesz loss.&lt;/p&gt;
&lt;p&gt;&lt;mark&gt;The paper proves that this calibration is sufficient to achieve &amp;ldquo;doubly robust asymptotic linearity&amp;rdquo; (DRAL)—meaning the estimator is asymptotically normal whenever at least one nuisance converges at $n^{-1/4}$ rate, even if the other is inconsistent.&lt;/mark&gt; They also develop bootstrap-assisted confidence intervals that avoid estimating additional nuisance functions. Empirically, they evaluate on simulated data with deliberately misspecified nuisances and on semi-synthetic benchmarks (ACIC, IHDP, Twins, LaLonde), comparing against standard AIPW and the iterative DR-TMLE.&lt;/p&gt;
&lt;h3 id=&#34;why-is-this-important&#34;&gt;Why is this important?&lt;/h3&gt;
&lt;p&gt;Applied researchers using flexible ML methods (random forests, neural networks, gradient boosting) for nuisance estimation face a dilemma: these methods provide consistency but may not converge fast enough to guarantee valid inference, especially in moderate dimensions. &lt;mark&gt;The standard &amp;ldquo;product rate&amp;rdquo; condition requiring both nuisances to converge at $n^{-1/4}$ can fail when even one nuisance is complex or misspecified. Calibrated DML transforms this multiplicative requirement into a requirement on just the better-estimated nuisance.&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;The practical benefits are substantial: better finite-sample coverage (e.g., improving from 32% to 90% coverage in ACIC-2017 simulations), reduced bias, and a method that integrates into existing pipelines with minimal code changes.&lt;/p&gt;
&lt;p&gt;The theoretical contribution is equally significant—it establishes a novel connection between prediction calibration and causal inference validity, showing that calibration of nuisances provides the &amp;ldquo;debiasing&amp;rdquo; needed for doubly robust inference.&lt;/p&gt;
&lt;h3 id=&#34;who-should-care&#34;&gt;Who should care?&lt;/h3&gt;
&lt;p&gt;Applied researchers estimating treatment effects with observational data using ML for nuisance estimation will find immediate practical value—this method provides insurance against nuisance misspecification without computational overhead.&lt;/p&gt;
&lt;p&gt;Econometricians and biostatisticians working on semiparametric inference will appreciate the theoretical framework connecting calibration to doubly robust properties.&lt;/p&gt;
&lt;p&gt;Methodologists developing new estimands can use this as a general recipe: the approach applies to any linear functional of the outcome regression (counterfactual means, ATE, partial covariance, survival outcomes under missingness), not just the ATE.&lt;/p&gt;
&lt;p&gt;Researchers in policy evaluation, epidemiology, marketing, and tech who regularly use AIPW-style estimators should consider adopting calibrated DML as a default given its robustness benefits.&lt;/p&gt;
&lt;h3 id=&#34;do-we-have-code&#34;&gt;Do we have code?&lt;/h3&gt;
&lt;p&gt;Yes, the authors provide both R and Python implementations.&lt;/p&gt;
&lt;p&gt;An R package &lt;code&gt;calibratedDML&lt;/code&gt; and Python code are available on GitHub at &lt;a href=&#34;https://github.com/Larsvanderlaan/calibratedDML&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://github.com/Larsvanderlaan/calibratedDML&lt;/a&gt;. The paper includes complete code listings for calibrating inverse propensity weights and outcome regressions using &lt;code&gt;xgboost&lt;/code&gt; with monotonicity constraints.&lt;/p&gt;
&lt;p&gt;The implementation is straightforward: isotonic regression is performed using gradient-boosted trees with &lt;code&gt;monotone_constraints=1&lt;/code&gt; and a single boosting round, making it computationally efficient and easy to integrate into existing DML workflows.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;In summary, this paper provides a remarkably practical solution to a longstanding theoretical problem. The insight that isotonic calibration—a standard tool from prediction—can unlock doubly robust inference is both elegant and immediately actionable. For anyone running AIPW or related estimators with ML nuisances, calibrating cross-fitted estimates before debiasing is now the obvious default: it costs almost nothing computationally and provides genuine protection against the scenario where one of your nuisance models is less reliable than you hoped.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;van der Laan, Lars, Alex Luedtke, and Marco Carone (2024), “Doubly robust inference via calibration,” arXiv preprint arXiv:2411.02771.&lt;/p&gt;
&lt;p&gt;An R package &lt;code&gt;calibratedDML&lt;/code&gt; and Python code are available on GitHub at &lt;a href=&#34;https://github.com/Larsvanderlaan/calibratedDML&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://github.com/Larsvanderlaan/calibratedDML&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Balancing Weights for Causal Inference</title>
      <link>https://chenxing.space/blog/balancing-weights-for-causal-inference/</link>
      <pubDate>Wed, 22 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/balancing-weights-for-causal-inference/</guid>
      <description>&lt;h2 id=&#34;tldr&#34;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;Cohn et al. (2023) introduces the &lt;strong&gt;balancing approach&lt;/strong&gt; to weighting for causal inference in observational studies. &lt;mark&gt;Unlike traditional methods that model the propensity score directly, balancing weights are estimated by solving an optimization problem that directly targets covariate balance between treatment groups.&lt;/mark&gt; The authors demonstrate that this approach offers protection against model misspecification, connects naturally to bias-variance trade-offs, and can be augmented with outcome modeling for improved performance. Applied to the classic LaLonde job training data, balancing methods achieve better covariate balance than standard propensity score approaches while maintaining reasonable effective sample sizes.&lt;/p&gt;
&lt;h2 id=&#34;what-is-this-paper-about&#34;&gt;What is this paper about?&lt;/h2&gt;
&lt;p&gt;Covariate balance is fundamental to causal inference: randomized experiments achieve it by design, while observational studies must adjust for it. This chapter addresses a key challenge in observational causal inference—how to construct weights that remove confounding by balancing observed covariates between treated and control groups. The traditional modeling approach estimates propensity scores (the probability of treatment given covariates) and inverts them to create weights, but this relies heavily on correct model specification. When the propensity score model is wrong, the resulting weights may fail to balance covariates in the sample, leading to biased treatment effect estimates. &lt;mark&gt;The chapter explores an alternative: directly finding weights that achieve balance in the observed data, rather than first modeling the propensity score.&lt;/mark&gt;&lt;/p&gt;
&lt;h2 id=&#34;what-do-the-authors-do&#34;&gt;What do the authors do?&lt;/h2&gt;
&lt;p&gt;The authors formalize the balancing approach as an optimization problem that jointly minimizes covariate imbalance and weight dispersion (variance). They show how different choices of the &amp;ldquo;model class&amp;rdquo; M—the set of functions of covariates to balance—correspond to different assumptions about the outcome model and lead to different optimization formulations. Using the LaLonde dataset (a constructed observational study where the true treatment effect is known), they compare three designs: balancing main covariate terms only, balancing up to three-way interactions, and balancing an infinite-dimensional reproducing kernel Hilbert space (RKHS). For each design, they evaluate covariate balance using standardized mean differences, examine the effective sample size (a measure of weight dispersion), and explore the bias-variance trade-off by varying regularization parameters. The authors also demonstrate how balancing weights can be augmented with outcome regression to further reduce bias, and they establish asymptotic normality results for inference.&lt;/p&gt;
&lt;h2 id=&#34;why-is-this-important&#34;&gt;Why is this important?&lt;/h2&gt;
&lt;p&gt;This work matters because most observational studies include covariates in their analysis, yet practitioners often don&amp;rsquo;t carefully consider whether their weighting method actually achieves balance on the relevant covariate functions. The balancing approach makes covariate balance a first-order design criterion rather than a post-hoc diagnostic check. It reveals the implicit bias-variance trade-offs in weighting methods and shows that different assumptions about the outcome model (linear, interactive, nonparametric) lead to fundamentally different weighting strategies. &lt;mark&gt;The framework &lt;strong&gt;unifies&lt;/strong&gt; many existing methods (entropy balancing, stable weights, kernel balancing) under one optimization structure and clarifies the connection between the balancing and modeling approaches through Lagrangian duality.&lt;/mark&gt; Importantly, the chapter provides practical guidance on design choices—what to balance, how much dispersion to tolerate, whether to allow negative weights—that applied researchers face but often lack principled ways to resolve.&lt;/p&gt;
&lt;h2 id=&#34;who-should-care&#34;&gt;Who should care?&lt;/h2&gt;
&lt;p&gt;Applied researchers in economics, epidemiology, public policy, education, and medicine who use inverse propensity weighting or other covariate adjustment methods in observational studies. Methodologists working on causal inference, especially those developing new weighting estimators or studying properties of existing ones. Graduate students learning causal inference who need to understand the trade-offs between different adjustment strategies and the implicit assumptions behind common practices. Policy evaluators who must justify their modeling choices and demonstrate that their treatment effect estimates are robust to covariate imbalance. Anyone who has struggled with poor covariate balance after propensity score weighting or wondered how to choose between competing adjustment methods would benefit from this framework.&lt;/p&gt;
&lt;h2 id=&#34;do-we-have-code&#34;&gt;Do we have code?&lt;/h2&gt;
&lt;p&gt;The chapter does not provide standalone replication code or software packages. However, the authors note that many of the specific balancing methods discussed are available in existing R packages: entropy balancing in the &lt;code&gt;ebal&lt;/code&gt; package, covariate balancing propensity scores (CBPS) in the &lt;code&gt;CBPS&lt;/code&gt; package, and stable balancing weights in the &lt;code&gt;sbw&lt;/code&gt; package. The kernel balancing approach can be implemented using standard kernel methods in R or Python. The LaLonde dataset used throughout the examples is publicly available and widely used in the causal inference literature, making it straightforward to reproduce the analyses with these existing tools.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;In summary&lt;/strong&gt;, this chapter reframes propensity score weighting as a balance-optimization problem rather than a pure modeling exercise. By directly targeting the balancing property of inverse propensity weights, the approach offers robustness to propensity score misspecification while making explicit the bias-variance trade-offs inherent in any covariate adjustment. The LaLonde application demonstrates that balancing weights can substantially reduce covariate imbalance compared to standard methods, though at the cost of reduced effective sample size. The framework provides both theoretical insight (connecting balancing to dual regression, establishing asymptotic properties) and practical guidance (how to choose what to balance, when to augment with outcome modeling, whether to allow extrapolation) that fills an important gap in applied causal inference.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Cohn, Eric R., Eli Ben-Michael, Avi Feller, and José R. Zubizarreta (2023), “Balancing Weights for Causal Inference,” in Handbook of Matching and Weighting Adjustments for Causal Inference, Chapman and Hall/CRC. &lt;a href=&#34;https://www.taylorfrancis.com/chapters/edit/10.1201/9781003102670-16/balancing-weights-causal-inference-eric-cohn-eli-ben-michael-avi-feller-jos&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://www.taylorfrancis.com/chapters/edit/10.1201/9781003102670-16/balancing-weights-causal-inference-eric-cohn-eli-ben-michael-avi-feller-jos&lt;/a&gt;é-zubizarreta&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Triply Robust Panel Estimators 🛡: When You Don&#39;t Know Which Assumptions Hold</title>
      <link>https://chenxing.space/blog/triply-robust-panel-estimators-when-you-don-t-know-which-assumptions-hold/</link>
      <pubDate>Mon, 20 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/triply-robust-panel-estimators-when-you-don-t-know-which-assumptions-hold/</guid>
      <description>&lt;h2 id=&#34;tldr&#34;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;Athey, Imbens, Qu, and Viviano introduce the &lt;strong&gt;Triply RObust Panel (TROP) estimator&lt;/strong&gt;, which combines unit weights, time weights, and a flexible low-rank factor model to estimate causal effects in panel data. &lt;mark&gt;The estimator achieves triple robustness: it remains unbiased if any one of three components succeeds (unit balance, time balance, or correct outcome modeling).&lt;/mark&gt; Across 21 empirically calibrated simulation designs, TROP outperforms standard methods like DID, synthetic control, matrix completion, and synthetic difference-in-differences in 20 out of 21 cases, often by substantial margins. The method is particularly valuable when interactive fixed effects violate parallel trends assumptions and when researchers face uncertainty about which modeling assumptions hold.&lt;/p&gt;
&lt;h2 id=&#34;what-is-this-paper-about&#34;&gt;What is this paper about?&lt;/h2&gt;
&lt;p&gt;This paper addresses a fundamental challenge in panel data analysis: choosing among competing estimators when the validity of their underlying assumptions is uncertain. Traditional difference-in-differences (DID) methods assume parallel trends, synthetic control (SC) emphasizes pre-treatment outcome alignment, and matrix completion (MC) relies on low-rank factor structures, but researchers rarely know which assumptions hold in practice. A second problem is that existing methods often treat observations uniformly across time and units, ignoring that recent periods may be more informative for predicting counterfactuals than distant ones, and that some control units may be more comparable to treated units than others. Third, most SC-based methods are designed for a single treated unit and do not naturally extend to complex assignment patterns with multiple treated units and staggered adoption. The authors propose a &lt;strong&gt;unified framework&lt;/strong&gt; that addresses all three issues simultaneously while providing formal robustness guarantees.&lt;/p&gt;
&lt;h2 id=&#34;what-do-the-authors-do&#34;&gt;What do the authors do?&lt;/h2&gt;
&lt;p&gt;The authors develop the TROP estimator, which minimizes a weighted regression objective that incorporates:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;unit-specific weights (to upweight control units similar to treated units)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;time-specific weights (to upweight recent periods)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;a nuclear-norm-penalized low-rank factor model (to capture interactive fixed effects)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;They propose an exponential decay structure for time weights and distance-based weights for units, with all tuning parameters selected via leave-one-out cross-validation.&lt;/p&gt;
&lt;p&gt;To evaluate performance, they conduct extensive semi-synthetic simulations calibrated to seven real datasets (CPS wage data, Penn World Table GDP, Germany reunification, Basque country, California smoking, and Mariel boatlift), generating outcomes from rank-4 factor models with realistic autocorrelation and treatment assignments based on estimated propensity scores from actual policies. They systematically vary features of the data-generating process (removing autocorrelation, interactive effects, or fixed effects) and components of the estimator (shutting down unit weights, time weights, or the regression adjustment) to understand what drives performance differences. &lt;mark&gt;They provide formal theory showing that TROP&amp;rsquo;s bias equals the &lt;i&gt;&lt;u&gt;product&lt;/u&gt;&lt;/i&gt; of (1) unit imbalance, (2) time imbalance, and (3) regression model misspecification, establishing triple robustness.&lt;/mark&gt;&lt;/p&gt;







&lt;figure class=&#34;blog-figure&#34;&gt;
   &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20251020093604943.png&#34; alt=&#34;image-20251020093604943&#34; style=&#34;zoom:50%;&#34; /&gt; 
  &lt;figcaption&gt;
    
      Figure 1: Theorem 5.1 in the paper
    
  &lt;/figcaption&gt;
&lt;/figure&gt;
&lt;blockquote&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20251020094047519.png&#34; alt=&#34;image-20251020094047519&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;/blockquote&gt;
&lt;h2 id=&#34;why-is-this-important&#34;&gt;Why is this important?&lt;/h2&gt;
&lt;p&gt;The paper makes three critical contributions to applied panel data analysis. First, it demonstrates that interactive fixed effects—where unit-specific trends vary over time—are empirically pervasive and cause substantial bias in standard DID estimators, with violations present in over 58% of units across their applications and adding even a single interactive factor reducing RMSE by 10-60%. Second, it shows that no single existing method dominates: DID can have RMSE 900% higher than the best competitor in some settings, SC can be 580% worse, and even modern methods like SDID can be 90% worse, highlighting the risk of mechanically applying any one approach. Third, TROP provides a principled solution by learning from the data which components (unit balance, time balance, or factor modeling) matter most for each application, offering insurance against model misspecification without requiring researchers to know ex ante which assumptions hold. The triple robustness property means that getting any one of three components approximately right suffices for consistency, substantially weakening the conditions needed for valid causal inference.&lt;/p&gt;
&lt;h2 id=&#34;who-should-care&#34;&gt;Who should care?&lt;/h2&gt;
&lt;p&gt;Applied researchers using panel data methods to evaluate policies, interventions, or natural experiments should care deeply about this paper, especially those working with observational data where parallel trends may fail due to heterogeneous trends across units. Economists studying labor markets, public finance, development, or industrial organization who face settings with staggered treatment adoption or multiple treated units will benefit from the method&amp;rsquo;s flexibility. &lt;mark&gt;Methodologists interested in robust causal inference, doubly robust estimation, or combining machine learning with causal identification will find the theoretical framework valuable.&lt;/mark&gt;  Policy analysts and data scientists in government agencies producing impact evaluations need methods that perform well under uncertainty about modeling assumptions. Anyone who has abandoned a difference-in-differences project because pre-trends failed should reconsider whether allowing for interactive fixed effects and flexible weighting could recover valid identification.&lt;/p&gt;
&lt;h2 id=&#34;do-we-have-code&#34;&gt;Do we have code?&lt;/h2&gt;
&lt;p&gt;The paper does not provide replication code, R or Python packages, or links to implementation files. The method requires implementing nuclear-norm-penalized regression combined with weighted least squares and a cross-validation procedure to select three tuning parameters (λ_unit, λ_time, λ_nn), which is computationally involved but feasible for researchers comfortable with optimization. The algorithms are described in detail (Algorithms 1-3 in the paper), and the exponential decay weighting scheme and distance metrics are explicitly specified in equations (3) and surrounding text. Researchers would need to build their own implementation or wait for the authors to release software, though the paper&amp;rsquo;s clarity makes this more tractable than usual.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;In summary&lt;/strong&gt;, this paper introduces a methodologically sophisticated yet practically motivated estimator that combines the strengths of synthetic control (unit weighting), modern DID approaches (time weighting), and matrix completion (low-rank modeling) while avoiding the brittleness of relying on any single component. The extensive simulations demonstrate that TROP consistently outperforms existing methods across diverse empirically relevant settings, with the advantage most pronounced when interactive fixed effects are present—a common feature of real data that violates standard parallel trends assumptions. The triple robustness property provides formal theoretical insurance: if researchers get unit balance approximately right, or time balance approximately right, or the factor model approximately right, the estimator remains consistent. &lt;mark&gt;For applied work, this represents a major advance in practical robustness, moving beyond the &amp;ldquo;which method should I use?&amp;rdquo; question to &amp;ldquo;let the data tell us which features matter most.&amp;quot;&lt;/mark&gt; The lack of available code is the main practical limitation, but the clear exposition of algorithms makes implementation feasible for quantitatively oriented researchers.&lt;/p&gt;
&lt;h2 id=&#34;points-i-enjoyed&#34;&gt;Points I enjoyed&lt;/h2&gt;
&lt;h3 id=&#34;trop-vs-sdid&#34;&gt;TROP vs. SDID&lt;/h3&gt;
&lt;p&gt;Let&amp;rsquo;s compare the TROP vs. SDID:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Triple Robustness (TROP)&lt;/strong&gt;: TROP achieves triple robustness, meaning its asymptotic bias vanishes if either the &lt;strong&gt;unit weights&lt;/strong&gt;, or the &lt;strong&gt;time weights&lt;/strong&gt;, or the &lt;strong&gt;flexible low-rank regression adjustment&lt;/strong&gt; successfully remove the underlying biases. The bias of TROP depends on the product of unit imbalance, time imbalance, and the bias arising from a potentially misspecified regression adjustment.&lt;/p&gt;

  
  
  
  
  
  
  &lt;figure class=&#34;blog-figure&#34;&gt;
     &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/SCR-20251021-kgwh.png&#34; alt=&#34;SCR-20251021-kgwh&#34; style=&#34;zoom:50%;&#34; /&gt; 
    &lt;figcaption&gt;
      
        Figure 2: Key difference between TROP and SDID – adding the low-rank factor component in the outcome model
      
    &lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Double Robustness (SDID)&lt;/strong&gt;: SDID (Arkhangelsky et al. 2021) achieves only double robustness, where bias vanishes if either unit imbalance or time imbalance is negligible. It lacks the third robust channel provided by the flexible regression adjustment.&lt;/p&gt;

  
  
  
  
  
  
  &lt;figure class=&#34;blog-figure&#34;&gt;
     &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/SCR-20251021-kkco.png&#34; alt=&#34;SCR-20251021-kkco&#34; style=&#34;zoom:50%;&#34; /&gt; 
    &lt;figcaption&gt;
      
        Figure 3: In SDID, there is no low-rank component $L_{it}$
      
    &lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;how-to-come-up-with-this-trop-estimator-my-thought&#34;&gt;How to come up with this TROP estimator? My Thought&lt;/h3&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    In my view, the idea behind TROP is to marry the &lt;strong&gt;SDID&lt;/strong&gt; framework with the &lt;strong&gt;matrix completion&lt;/strong&gt; estimator
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Recall that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Matrix completion methods (Athey et al. 2021) focus &lt;strong&gt;solely on modeling the outcomes&lt;/strong&gt; as a function of the latent factors. These methods aim to accurately model the control outcome process using a low-rank factor structure that generalizes the two-way fixed effects (TWFE) model&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;SDID is a &amp;ldquo;local&amp;rdquo; TWFE estimator, &lt;strong&gt;which essentially corresponds to modeling the &lt;u&gt;assignment mechanism&lt;/u&gt;&lt;/strong&gt;, via local weighting (i.e. unit and time weights)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The estimated average &lt;strong&gt;counterfactual&lt;/strong&gt; from the TROP estimator has the following form:&lt;/p&gt;







&lt;figure class=&#34;blog-figure&#34;&gt;
   &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/SCR-20251024-nlpp-2.png&#34; alt=&#34;SCR-20251024-nlpp-2&#34; style=&#34;zoom:50%;&#34; /&gt; 
  &lt;figcaption&gt;
    
      Figure 4: The estimated average counterfactual from the TROP estimator. I think the key intuition here is the &amp;#39;bias correction&amp;#39;
    
  &lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;It contains both &lt;strong&gt;unit/time weights&lt;/strong&gt; (the key idea in SDID) and &lt;strong&gt;outcome model (low-rank factor model)&lt;/strong&gt; in matrix completion methods. The TROP estimator, $\hat{\tau}^{\text{TROP}}$, applies unit and time weights, $\omega_i$ and $\theta_t$ respectively,  to correct the bias from the outcome model $(\mathbf{L})$.&lt;/p&gt;
&lt;p&gt;This form also reminds me the &lt;strong&gt;idea&lt;/strong&gt; in AIPW estimator or the generic debiased framework (Chernozhukov, Newey, and Singh 2022) – &lt;mark&gt;&lt;strong&gt;leveraging the Riesz representer to correct for bias&lt;/strong&gt;&lt;/mark&gt;.&lt;/p&gt;







&lt;figure class=&#34;blog-figure&#34;&gt;
   &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20251024152434011.png&#34; alt=&#34;image-20251024152434011&#34; style=&#34;zoom:100%;&#34; /&gt; 
  &lt;figcaption&gt;
    
      Figure 5: screen shot from my previous post about AutoDML - using Riesz Representer to correct for bias
    
  &lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Check more details from my previous post: &lt;a href=&#34;https://chenxing.space/blog/intuition-for-doubly-robust-estimator/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;&amp;ldquo;Intuition for Doubly Robust Estimator&amp;rdquo;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Athey, Susan, Guido Imbens, Zhaonan Qu, and Davide Viviano (2025), “Triply robust panel estimators.”
&lt;a href=&#34;https://doi.org/10.48550/arXiv.2508.21536&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.48550/arXiv.2508.21536&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Athey, Susan, Mohsen Bayati, Nikolay Doudchenko, Guido Imbens, and Khashayar Khosravi (2021), “Matrix Completion Methods for Causal Panel Data Models,” &lt;i&gt;Journal of the American Statistical Association&lt;/i&gt;, 116 (536), 1716–30.&lt;/p&gt;
&lt;p&gt;Arkhangelsky, Dmitry, Susan Athey, David A. Hirshberg, Guido W. Imbens, and Stefan Wager (2021), “Synthetic Difference-in-Differences,” &lt;i&gt;American Economic Review&lt;/i&gt;, 111 (12), 4088–4118.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, and Rahul Singh (2022), “Automatic Debiased Machine Learning of Causal and Structural Effects,” &lt;i&gt;Econometrica&lt;/i&gt;, 90 (3), 967–1027.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>AIPW vs. Residual-on-Residual regression: Non-Parametric Flexibility or Efficiency?</title>
      <link>https://chenxing.space/blog/aipw-vs-residual-on-residual-regression-non-parametric-flexibility-or-efficiency/</link>
      <pubDate>Wed, 11 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/aipw-vs-residual-on-residual-regression-non-parametric-flexibility-or-efficiency/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Two powerful tools in causal inference are the &lt;strong&gt;Augmented Inverse Propensity Weighting (AIPW)&lt;/strong&gt; estimator and the &lt;strong&gt;Residual-on-Residual regression&lt;/strong&gt; estimator for partially linear models. Drawing from Wager’s notes (2024), this post breaks down how these estimators work, compares their strengths and weaknesses, and offers tips for when to use each.&lt;/p&gt;
&lt;h2 id=&#34;residual-on-residual-regression&#34;&gt;Residual-on-Residual regression&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;semiparamteric partially linear model&lt;/strong&gt; (PLM) assumes the outcome $ Y $ can be written as:&lt;/p&gt;
&lt;p&gt;$$ Y = \theta D + g(X) + \epsilon \tag{1}$$&lt;/p&gt;
&lt;p&gt;Here, $ D $ is a binary treatment, $ \theta $ is the causal parameter of interest, $ g(X) $ is some unknown function of covariates $ X $, and $ \epsilon $ is random noise with $\E[\epsilon \mid D, X] = 0$. The residual-on-residual estimator (or Robinson (1988) estimator) isolates $ \theta $ in two steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Partialing out covariates&lt;/strong&gt;: &amp;ldquo;Partialing out&amp;rdquo; $ X $ from $D$ and $Y$ using nonparametric regression and cross-fitting: $ \tilde{D} = D - \hat{\E}[D \mid X] $ and $ \tilde{Y} = Y - \hat{\E}[Y \mid X] $.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Estimate the effect&lt;/strong&gt;: Use linear regression $ \tilde{Y} \sim \tilde{D} $ to get $ \hat{\theta}$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    However, the model is &lt;strong&gt;not fully general&lt;/strong&gt;, because it  imposes a &lt;strong&gt;parametric specification&lt;/strong&gt; on the key component of interest. It imposes &lt;strong&gt;additivity&lt;/strong&gt; in $g(X)$ and $D$.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;augmented-inverse-propensity-weighting-aipw-estimator&#34;&gt;Augmented Inverse Propensity Weighting (AIPW) Estimator&lt;/h3&gt;
&lt;p&gt;AIPW takes a &lt;strong&gt;fully non-parametric&lt;/strong&gt; approach, aiming to estimate ATE, $ \tau = E[Y(1) - Y(0)] $, where $ Y(1) $ and $ Y(0) $ are potential outcomes under treatment and control.&lt;/p&gt;
 $$ \hat{\tau}_{\text{AIPW}} = \frac{1}{n} \sum_{i=1}^n \left[ \hat{\mu}_1(X_i) - \hat{\mu}_0(X_i) + D_i \frac{Y_i - \hat{\mu}_1(X_i)}{\hat{e}(X_i)} - (1 - D_i) \frac{Y_i - \hat{\mu}_0(X_i)}{1 - \hat{e}(X_i)} \right] $$ 
&lt;p&gt;Here, $ \hat{\mu}_1(X) $ and $ \hat{\mu}_0(X) $ are outcome regression estimators for treated and untreated units, and $ \hat{e}(X) $ is the estimator of propensity score. AIPW combines outcome modeling with inverse propensity weighting, making it &lt;strong&gt;doubly robust&lt;/strong&gt;: it’s consistent if &lt;em&gt;either&lt;/em&gt; the outcome model or propensity score is consistent.&lt;/p&gt;
&lt;h2 id=&#34;the-key-difference-non-parametric-vs-partially-linear&#34;&gt;The Key Difference: Non-Parametric vs. Partially Linear&lt;/h2&gt;
&lt;p&gt;&lt;mark&gt;&lt;strong&gt;AIPW is fully non-parametric&lt;/strong&gt;, imposing no specific parametric form on the treatment effect, while &lt;strong&gt;residual-on-residual regression estimator assumes a partially linear structure&lt;/strong&gt;&lt;/mark&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;AIPW’s flexibility&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;AIPW does not assume a specific parametric form!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AIPW is &lt;strong&gt;efficient&lt;/strong&gt; in the generic &lt;strong&gt;non-parametric&lt;/strong&gt; setting.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250611092857556.png&#34; alt=&#34;image-20250611092857556&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;/blockquote&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Residual-on-residual regression’s structure&lt;/strong&gt;: What partially linear assumption buys us is that residual-on-residual estimators that exploit this constraint can have &lt;strong&gt;smaller variance than AIPW&lt;/strong&gt;. In other words, adding this additional structure makes the residual-on-residual estimator more efficient than AIPW.&lt;/p&gt;
&lt;blockquote&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250611092145259.png&#34; alt=&#34;image-20250611092145259&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;/blockquote&gt;
&lt;div class=&#34;alert alert-warning&#34;&gt;
&lt;div&gt;
  A risk   of using the residual-on-residual estimator is that &lt;strong&gt;constant treatment effect&lt;/strong&gt; model (1) may be misspecified.
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
&lt;h2 id=&#34;why-this-matters&#34;&gt;Why This Matters&lt;/h2&gt;
&lt;p&gt;The choice between AIPW and residual-on-residual regression reflects a deeper trade-off in causal inference: &lt;strong&gt;flexibility versus efficiency&lt;/strong&gt;. AIPW’s non-parametric nature makes it a Swiss Army knife for complex data, while PLM structure is like a precision tool—effective when conditions are right.&lt;/p&gt;
&lt;p&gt;As Wager’s notes highlight:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Both AIPW and residual-on-residual regression are &lt;strong&gt;Neyman-orthogonal&lt;/strong&gt;, making them robust to first‐stage errors.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;However, their assumptions shape their performance. AIPW attains the lowest possible asymptotic variance for ATE under unconfoundedness. The residual-on-residual estimator, by imposing extra structure, can go beyond that bound when its structure is correct but at the cost of vulnerability to misspecification.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;conclusion&#34;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;AIPW’s fully non-parametric approach offers robustness and flexibility, while residual-on-residual regression’s partially linear structure prioritizes efficiency when assumptions hold.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Wager, S. (2024). &lt;em&gt;Causal inference: A statistical learning approach&lt;/em&gt;. &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/causal_inf_book.pdf&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Propensity Score Methods</title>
      <link>https://chenxing.space/blog/notes-on-propensity-score/</link>
      <pubDate>Tue, 03 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-propensity-score/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Here are my notes on propensity scores, mainly from Prof. Ding&amp;rsquo;s textbook (2024).&lt;/p&gt;
&lt;p&gt;The traditional propensity score analysis workflow is shown in the image below, which I will not cover in detail. Instead, I will summarize the key theorems and results from Ding&amp;rsquo;s textbook.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/Image%20from%20Chap3.3_observational_PS,%20page%2016.png&#34; alt=&#34;Image from Chap3.3_observational_PS, page 16&#34; style=&#34;zoom:50%;&#34; /&gt;
  &lt;figcaption&gt;Figure 1: Traditional propensity score analysis workflow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I will also provide some connections with &lt;strong&gt;Riesz Representer (RR)&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Why connect with the Riesz Representer (RR)? The connection provides a powerful generalization of the foundational Rosenbaum-Rubin (1983) result.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Rosenbaum and Rubin showed that &lt;strong&gt;conditioning on the propensity score is sufficient for removing confounding bias&lt;/strong&gt; when estimating causal effects. The Riesz representer extends this principle: &lt;strong&gt;it suffices to regress on the Riesz representer&lt;/strong&gt; to obtain unbiased estimates of the average treatment effect.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;strong&gt;key insight&lt;/strong&gt; is that the Riesz representer, like the propensity score, serves as a sufficient statistic – it captures all the confounding information necessary for unbiased estimation of your target causal parameter.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;setting--notation&#34;&gt;Setting &amp;amp; Notation&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Binary treatment $Z$&lt;/li&gt;
&lt;li&gt;Potential outcomes $\{Y(0), Y(1)\}$  &lt;/li&gt;
&lt;li&gt;Propensity score: $\P(Z = 1 \mid X)$, where $X$ represents covariates&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Two approaches learning causal relationships:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Outcome process (via outcome regression)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Treatment assignment mechanism&lt;/strong&gt; (via propensity score)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following summarizes the key theorems and results related to propensity scores from Prof. Ding&amp;rsquo;s textbook.&lt;/p&gt;
&lt;h2 id=&#34;1-the-propensity-score-as-a-markdimension-reductionmark-tool&#34;&gt;1. The propensity score as a &lt;mark&gt;dimension reduction&lt;/mark&gt; tool&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(pscore as dimension reduction tool)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$
\text { If } Z \indep \{Y(1), Y(0)\} \mid X, \text { then } Z \indep \{Y(1), Y(0)\} \mid e(X) .
$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Covariates $X$ can be &lt;strong&gt;high dimensional&lt;/strong&gt;, but the propensity score, $e(X) \in \R$, is a  &lt;strong&gt;1-dimensional scalar&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;We can view the propensity score as a &lt;strong&gt;dimensional reduction&lt;/strong&gt; tool&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;2-propensity-score-stratification&#34;&gt;2. Propensity score stratification&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;: Discretize the estimated propensity score by its $K$ quantiles:&lt;/p&gt;
 
$$
Z \indep \{Y(1), Y(0)\} \mid \hat{e}^{\prime}(X)=e_k \quad(k=1, \ldots, K) .
$$

&lt;p&gt;Estimate ATE within each subclass and then average by the block size&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Advantage&lt;/strong&gt;: The propensity score stratification estimator &lt;strong&gt;only requires the correct ordering&lt;/strong&gt; of the estimated propensity scores rather than their exact values, which makes it &lt;strong&gt;relatively robust&lt;/strong&gt; compared with other methods&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;3-propensity-score-weighting&#34;&gt;3. Propensity score weighting&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Invese propensity score weighting (IPW))&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
If $Z \indep \{Y(1), Y(0)\} \mid X$ and $ 0 &lt; e(X) &lt; 1$, then
$$E\{Y(1)\}=E\left\{\frac{Z Y}{e(X)}\right\}, \quad E\{Y(0)\}=E\left\{\frac{(1-Z) Y}{1-e(X)}\right\}$$
and 
$$
\begin{aligned}
\tau &amp;=E\{Y(1)-Y(0)\}\\
&amp;=E\left\{\frac{Z Y}{e(X)}-\frac{(1-Z) Y}{1-e(X)}\right\} \\
&amp;=E\left\{HY \right\}
\end{aligned}
$$
where
$H := \left[\frac{Z}{e(X)}-\frac{(1-Z) }{1-e(X)}\right]$ is called the &lt;strong&gt;Horvitz-Thompson transform&lt;/strong&gt;.

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Connection the Riesz Representer (RR)&lt;/p&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(RR in the case of ATE)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
    
    In the case of ATE, the Riesz Representer, $\alpha(Z, X)$, has the same form as above Horvitz-Thompson transform,
    $$
    \alpha(Z, X) = \left[\frac{Z}{e(X)}-\frac{(1-Z) }{1-e(X)}\right]
    $$
    
  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;31-estimation&#34;&gt;3.1 Estimation&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The sample version of IPW is called the &lt;strong&gt;Horvitz–Thompson (HT) estimator&lt;/strong&gt;,&lt;/p&gt;
 $$
\hat{\tau}^{\mathrm{ht}}=\frac{1}{n} \sum_{i=1}^n \frac{Z_i Y_i}{\hat{e}\left(X_i\right)}-\frac{1}{n} \sum_{i=1}^n \frac{\left(1-Z_i\right) Y_i}{1-\hat{e}\left(X_i\right)}
$$ 
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;HT estimator $\hat{\tau}^{\mathrm{ht}}$ has many problems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Problem: lack of invariance&lt;/strong&gt;, i.e. if we replace $Y_i$ by $Y_i + c$, $\hat{\tau}^{\mathrm{ht}}$ changed because it depends on $c$. This is not reasonable.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Solution: normalizing the weights&lt;/strong&gt;&lt;/p&gt;
 $$
\hat{\tau}^{\text {hajek }}=\frac{\sum_{i=1}^n \frac{Z_i Y_i}{\hat{e}\left(X_i\right)}}{\sum_{i=1}^n \frac{Z_i}{\hat{e}\left(X_i\right)}}-\frac{\sum_{i=1}^n \frac{\left(1-Z_i\right) Y_i}{1-\hat{e}\left(X_i\right)}}{\sum_{i=1}^n \frac{1-Z_i}{1-\hat{e}\left(X_i\right)}} .
$$ 
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Hajek estimator is invariant to the location transformation&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;32-strong-overlap-condition&#34;&gt;3.2 Strong overlap condition&lt;/h3&gt;
&lt;p&gt;Many asymptotic analyses require a &lt;em&gt;strong overlap&lt;/em&gt; condition,&lt;/p&gt;
&lt;p&gt; $$
0&lt;\alpha_{\mathrm{L}} \leq e(X) \leq \alpha_{\mathrm{U}}&lt;1
$$ 
In practice,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Crump et al. (2009) suggested $α_L = 0.1$ and $α_U = 0.9$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Kurth et al. (2005) suggested $α_L = 0.05$ and $α_U = 0.95$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;4-balancing-property&#34;&gt;4. Balancing property&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(balancing property)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182217046.png&#34; alt=&#34;image-20250603182217046&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Conditional on $e(X)$, the treatment and the covariates are independent&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Within the same level of the propensity score, the covariate distributions are &lt;strong&gt;balanced&lt;/strong&gt; across the treatment and control groups&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Useful implication&lt;/strong&gt;: we can check whether the propensity score model is specified well enough to ensure the &lt;strong&gt;covariate balance&lt;/strong&gt; in the data&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;41-propensity-score-is-a-balancing-score&#34;&gt;4.1 Propensity score is a balancing score&lt;/h3&gt;







&lt;div class=&#34;math-environment definition&#34; id=&#34;definition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Definition 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182507172.png&#34; alt=&#34;image-20250603182507172&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-4&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 4&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Propensity score is a balancing score)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182648222.png&#34; alt=&#34;image-20250603182648222&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182857253.png&#34; alt=&#34;image-20250603182857253&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This is relevant in &lt;strong&gt;subgroup analysis&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The conditional independence in (11.5) &lt;mark&gt;ensures &lt;strong&gt;unconfoundedness&lt;/strong&gt; holds given the propensity score, within each level of $X_1$&lt;/mark&gt;. Therefore, we can perform the same analysis based on the propensity score, within each level of $X_1$, yielding estimates for two subgroup effects&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;5-doubly-robust-or-aipw&#34;&gt;5. Doubly Robust or AIPW&lt;/h2&gt;
&lt;p&gt;The following Theorem is summarized from Prof. Wager&amp;rsquo;s lecture notes (2024).&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-5&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 5&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(strong double robustness of AIPW estimator)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Define the outcome regression as $$
\mu_{(z)}(x)=\mathbb{E}\left[Y_i(z) \mid X_i=x\right],
$$
Define AIPW estimator as
$$
\begin{aligned}
\hat{\tau}_{A I P W} &amp; =\underbrace{\frac{1}{n} \sum_{i=1}\left(\hat{\mu}_{(1)}\left(X_i\right)-\hat{\mu}_{(0)}\left(X_i\right)\right)}_{\text { outcome regression estimator }} \\
&amp; +\underbrace{\frac{1}{n} \sum_{i=1}^n\left(\frac{Z_i}{\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(1)}\left(X_i\right)\right)-\frac{1-Z_i}{1-\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(0)}\left(X_i\right)\right)\right)}_{\text { applying IPW to the regression residuals }}
\end{aligned}
$$

If we use estimators $\hat{\mu}_{(z)}(x)$ and $\hat{e}(x)$ that are both consistent with root-mean squared error (RMSE) decaying faster than $n^{-\alpha_\mu}$ and $n^{-\alpha_e}$ respectively, and if furthermore $\alpha_\mu+\alpha_e \geq 1 / 2$, then

$$
\begin{aligned}
&amp; \sqrt{n}\left(\hat{\tau}_{A I P W}-\tau\right) \Rightarrow \mathcal{N}\left(0, V_{A I P W}\right) \\
&amp; V_{A I P W}=\operatorname{Var}\left[\tau\left(X_i\right)\right]+\mathbb{E}\left[\frac{\sigma_0^2\left(X_i\right)}{1-e\left(X_i\right)}\right]+\mathbb{E}\left[\frac{\sigma_1^2\left(X_i\right)}{e\left(X_i\right)}\right]
\end{aligned}
$$
where
$$
\sigma_{(z)}^2(x)=\operatorname{Var}\left[Y_i(z) \mid X_i=x\right]
$$


  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Check my previous post: &lt;a href=&#34;https://chenxing.space/blog/intuition-for-doubly-robust-estimator/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Intuition for Doubly Robust Estimator&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AIPW provides a natural starting point for understanding Double Machine Learning&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key insight&lt;/strong&gt; of RR in DML framework: Leverage the Riesz Representer, a &amp;ldquo;generalized version of propensity score&amp;rdquo; to &amp;ldquo;correct the bias&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;6-other-estimands-related-to-ipw&#34;&gt;6. Other Estimands related to IPW&lt;/h2&gt;
&lt;p&gt;More general, Li et al. (2018a) gave a unified discussion of the causal estimands in observational studies.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-6&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 6&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Ding (2024), Section 13.4)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602144337279.png&#34; alt=&#34;image-20250602144337279&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;br&gt;
Summary Table of common estimands: 
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602145602468.png&#34; alt=&#34;image-20250602145602468&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This table provides us a good way to understand and remember IPW estimator for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to remember $\tau^h$? Apply IPW on &lt;mark&gt;&amp;ldquo;pseudo outcome&amp;rdquo; $Yh(X)$ &lt;/mark&gt; then divide by $E(h(X))$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;When the parameter of interest is ATT, then $$E(h(X)) = E(e(X)) = E(E(Z \mid X)) = E(Z) = \P(Z = 1) = e$$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use it to better understand IPW for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;7-propensity-score-in-regression&#34;&gt;7. Propensity Score in Regression&lt;/h2&gt;
&lt;h3 id=&#34;ps-as-a-covariate&#34;&gt;PS as a covariate&lt;/h3&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-7&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 7&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(regression with pscore as a covariate)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under unconfoundedness, the coefficient of $Z$ in the population OLS fit of

$$
Y \sim 1+Z + e(X)
$$

equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;, $$
\tau_{\mathrm{O}}=\frac{E[e(X)\{1-e(X)\} \tau(X)]}{E[e(X)\{1-e(X)\}]}
,$$ which is the &lt;mark&gt;&lt;strong&gt;overlap-weighted average treatment effect&lt;/strong&gt;&lt;/mark&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Based on above Theorem, we also have:&lt;/p&gt;







&lt;div class=&#34;math-environment corollary&#34; id=&#34;corollary-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Corollary 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under unconfoundedness, &lt;br&gt;

1. the coefficient of $Z$ in the population OLS fit of

$$
Y \sim 1+Z + e(X) + X
$$

also equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;, &lt;br&gt;&lt;/br&gt;

2. the coefficient of $Z-e(X)$ in the population OLS fit of

$$
Y \sim [Z - e(X)] \quad \text{or} \quad Y \sim 1 + [Z - e(X)]
$$

also equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;ps-as-a-weight&#34;&gt;PS as a weight&lt;/h3&gt;
&lt;p&gt;There is a convenient way to obtain $\hat{\tau}^{\text{hajek}}$ based on WLS.&lt;/p&gt;







&lt;div class=&#34;math-environment proposition&#34; id=&#34;proposition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Proposition 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(convenient to obtain $\hat{\tau}^{\text{hajek}}$ based on WLS)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250604101708438.png&#34; alt=&#34;image-20250604101708438&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Need to use bootstrap for standard error&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Why does the WLS give a consistent estimator for $\tau$ ?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In RCT with a constant propensity score, we can simply use the coefficient of $Z_i$ in the OLS fit of $Y_i$ on ( $1, Z_i$ ) to estimate $\tau$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In observational studies, we need to deal with the selection bias. The key idea is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;If we weight the treated units by $\frac{1}{e(X_i)}$ and the control units by $\frac{1}{1-e(X_i)}$, then both treated and control groups can represent the whole population&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Thus, &lt;strong&gt;by weighting, we effectively have a pseudo-randomized experiment&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(IPCW)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Inverse Probability of Censoring Weighting (IPCW) follows the same idea — it adjusts for censoring bias by reweighting observations based on their probability of being uncensored.

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Consequently, the difference between the weighted means is consistent for $\tau$. The numerical equivalence of $\hat{\tau}^{\text {hajek }}$ and WLS is not only a fun numerical fact itself but also useful for motivating more complex estimators with covariate adjustment&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Ding, Peng (2024), &lt;i&gt;A First Course in Causal Inference&lt;/i&gt;, CRC Press.&lt;/p&gt;
&lt;p&gt;Wager, S. (2024). Causal inference: A statistical learning approach. &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/causal_inf_book.pdf&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on DML for DiD: A Unified Approach</title>
      <link>https://chenxing.space/blog/notes-on-dml-for-did/</link>
      <pubDate>Mon, 02 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-dml-for-did/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;This blog post explores how Double Machine Learning (DML) extends to conditional Difference-in-Differences (DiD), focusing on doubly robust estimators. The &lt;strong&gt;key insight&lt;/strong&gt; is that conditional DiD can be understood through the lens of cross-sectional ATT estimation.&lt;/p&gt;
&lt;h2 id=&#34;foundation-cross-sectional-att-estimation&#34;&gt;Foundation: Cross-Sectional ATT Estimation&lt;/h2&gt;
&lt;p&gt;To build intuition, we start with the familiar cross-sectional setting. Standard identification requires three assumptions: &lt;strong&gt;SUTVA, unconfoundedness, and overlap&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&#34;step-1-propensity-score-approach-for-att&#34;&gt;Step 1: Propensity Score Approach for ATT&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Unlike ATE, ATT estimation requires only &amp;ldquo;one-sided&amp;rdquo; unconfoundedness and overlap conditions.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Identification Assumptions)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$Z \indep Y(0) \mid X \text{ and } e(X) &lt; 1$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Estimate ATT using &lt;strong&gt;IPW&lt;/strong&gt;&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Ding (2024), Section 13.2)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602144838745.png&#34; alt=&#34;image-20250602144838745&#34; style=&#34;zoom:40%;&#34; /&gt; 

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;More general, Li et al. (2018a) gave a unified discussion of the causal estimands in observational studies.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Ding (2024), Section 13.4)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602144337279.png&#34; alt=&#34;image-20250602144337279&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;br&gt;
Summary Table of common estimands: 
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602145602468.png&#34; alt=&#34;image-20250602145602468&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This table provides us a good way to understand and remember IPW estimator for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to remember $\tau^h$? Apply IPW on &lt;mark&gt;&amp;ldquo;pseudo outcome&amp;rdquo; $Yh(X)$ &lt;/mark&gt; then divide by $E(h(X))$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;When the parameter of interest is ATT, then $$E(h(X)) = E(e(X)) = E(E(Z \mid X)) = E(Z) = \P(Z = 1) = e$$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use it to better understand IPW for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;step-2-doubly-robust-att-estimator&#34;&gt;Step 2: Doubly Robust ATT Estimator&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Combines outcome regression and IPW methods&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;For DR estimator of ATT, check &lt;a href=&#34;https://chenxing.space/blog/intuition-for-doubly-robust-estimator/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;my previous post&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;More generally, we have&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(DR for general estimand, see Ding (2024), page 191)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602151637808.png&#34; alt=&#34;image-20250602151637808&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;extension-to-conditional-did&#34;&gt;Extension to Conditional DiD&lt;/h2&gt;
&lt;h3 id=&#34;identification-assumptions&#34;&gt;Identification Assumptions&lt;/h3&gt;
&lt;p&gt;Conditional DiD relies on two core assumptions: conditional parallel trends and no anticipation, plus an overlap condition.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(CausalML Book, page 457)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602152200078.png&#34; alt=&#34;image-20250602152200078&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;How to understand the &lt;strong&gt;overlap condition&lt;/strong&gt; (16.3.3)? It essentially imposes that there are control observations available for every value of $X$.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;the-key-insight-transformation-to-cross-sectional-problem&#34;&gt;The Key Insight: Transformation to Cross-Sectional Problem&lt;/h3&gt;
&lt;p&gt;By taking the difference,&lt;/p&gt;
&lt;p&gt;$$
\Delta Y = Y_{\text{after}} - Y_{\text{before}}
$$&lt;/p&gt;
&lt;p&gt;we transform panel data into a cross-sectional problem. This allows us to apply the same doubly robust framework used for cross-sectional ATT.&lt;/p&gt;
&lt;h3 id=&#34;the-unified-result&#34;&gt;The Unified Result&lt;/h3&gt;
&lt;p&gt;The Neyman orthogonal score for conditional DiD is &lt;strong&gt;identical&lt;/strong&gt; to the cross-sectional ATT score, where the outcome variable is simply the difference $\Delta Y$.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Neyman orthogonal score for ATT in conditional DiD&lt;/p&gt;







&lt;div class=&#34;math-environment proposition&#34; id=&#34;proposition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Proposition 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(see CausalML Book)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602153412853.png&#34; alt=&#34;image-20250602153412853&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Neyman orthogonal score for ATT in cross-sectional setting&lt;/p&gt;







&lt;div class=&#34;math-environment proposition&#34; id=&#34;proposition-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Proposition 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(see CausalML Book)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602153919836.png&#34; alt=&#34;image-20250602153919836&#34; style=&#34;zoom:40%;&#34; /&gt; &lt;br&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602154015540.png&#34; alt=&#34;image-20250602154015540&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;mark&gt;Comparing to the score for the ATT in cross-sectional setting, we see that DiD score is &lt;strong&gt;identical&lt;/strong&gt; to that for learning the ATT under &lt;strong&gt;unconfoundedness&lt;/strong&gt; where the outcome variable is simply defined as $\Delta Y$ &lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    This elegant connection demonstrates that the doubly robust estimator for conditional DiD is equivalent to the doubly robust ATT estimator applied to the differenced outcome $\Delta Y$.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;p&gt;Chernozhukov, Victor, Christian Hansen, Nathan Kallus, Martin Spindler, and Vasilis Syrgkanis (2024), “Applied causal inference powered by ML and AI.”&lt;/p&gt;
&lt;p&gt;Ding, P. (2024). A First Course in Causal Inference. CRC Press.&lt;/p&gt;
&lt;p&gt;Callaway, Brantly and Pedro H. C. Sant’Anna (2021), “Difference-in-Differences with multiple time periods,” &lt;i&gt;Journal of Econometrics&lt;/i&gt;, Themed Issue: Treatment Effect 1, 225 (2), 200–230.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K Newey, and Rahul Singh (2022), “Debiased machine learning of global and local parameters using regularized Riesz representers,” &lt;i&gt;The Econometrics Journal&lt;/i&gt;, 25 (3), 576–601.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, Victor Quintas-Martinez, and Vasilis Syrgkanis (2024), “Automatic debiased machine learning via riesz regression.”&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Callaway &amp; Sant’Anna (2021) – Staggered Adoption DiD</title>
      <link>https://chenxing.space/blog/notes-on-callaway-sant-anna-2021-staggered-adoption-did/</link>
      <pubDate>Tue, 27 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-callaway-sant-anna-2021-staggered-adoption-did/</guid>
      <description>&lt;h2 id=&#34;0-motivation&#34;&gt;0. Motivation&lt;/h2&gt;
&lt;p&gt;Staggered‐adoption policies break the canonical &lt;strong&gt;two-period / two-group&lt;/strong&gt; DiD model. It has been shown that the traditional two-way fixed-effects (TWFE) regression can assign &lt;strong&gt;negative weights&lt;/strong&gt; to treatment effects, thereby obscuring their dynamic and heterogeneous patterns.&lt;/p&gt;
&lt;p&gt;Callaway &amp;amp; Sant’Anna (2021) propose a &lt;strong&gt;divide-and-conquer&lt;/strong&gt; strategy:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Divide&lt;/strong&gt; the messy staggered panel into many honest $2\times2$ DiDs&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Conquer&lt;/strong&gt; by estimating each &amp;ldquo;little&amp;rdquo; DiD under familiar assumptions, then combine them with user-chosen weights to answer specific questions&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The key takeaway:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Use $\operatorname{ATT}(g, t)$ as a building block so we can transparently see how things are constructed&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Many different aggregation schemes are possible: they deliver different parameters&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Can allow for covariates via regressions adjustments, IPW, and DR.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;1-setup&#34;&gt;1. Setup&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Data structure&lt;/strong&gt;: Panel of units $i$ over time $t = 1,\dots,T$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $D_{it}$ be a binary variable. $D_{it} = 1$ if unit $i$ is treated in period $t$; $D_{it} = 0$ otherwise&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cohorts&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Define $G$ as the time period when a unit &lt;strong&gt;first becomes treated&lt;/strong&gt;. For all units that eventually get treated, $G$ defines which &amp;ldquo;group&amp;rdquo; they belong to&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Define $G_{g}$ as a binary variable. &lt;mark&gt;$G_{g} = 1$ if a unit is first treated in period $g$&lt;/mark&gt; (i.e.  $G_{i,g} = \1\{G_i = g\}$ )&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Define $G=\infty$ as &amp;ldquo;never treated&amp;rdquo; group.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Potential outcomes&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$Y_{it}(g)$: outcome at time $t$ if first treated in period $g$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$Y_{it}(\infty)$: outcome if never treated.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cohort-time ATT&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Assume $\mathrm{iid}$, drop unit index $i$. A parameter of interest that has clear interpretation is the $\operatorname{ATT}(g, t)$: $$
\operatorname{ATT}(g, t)=\mathbb{E}\left[Y_t(g)-Y_t(\infty) \mid G_g=1\right], \text { for } t \geq g .
$$ This defines one &amp;ldquo;clean&amp;rdquo; $2\times2$ DiD for each pair $(g,t)$.&lt;/p&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Why focus on time after treatment starting period, $t \ge g$? Because we need the &lt;mark&gt;no anticipation&lt;/mark&gt; assumption. Before treatment taking place, $t &lt; g$, there is no treatment effect.

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;2-key-assumptions&#34;&gt;2. Key Assumptions&lt;/h2&gt;
&lt;p&gt;Given that we never observe $Y(\infty)$ in post-treatment periods among units that have been treated, we need to make assumptions to identify $\operatorname{ATT}(g, t)$.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(No anticipation)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$
Y_{i t}(g)=Y_{i t}\left(g^{\prime}\right) \quad \forall \ i \text{, } t&lt;\min \left\{g, g^{\prime}\right\}
$$
Treatment cannot affect pre-treatment outcomes.

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;The no anticipation assumption has exactly the same content as in the $2\times2$ case.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Parallel trends based on &amp;#39;never-treated&amp;#39; control)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    


For each $t \in\{2, \ldots, T\}, g \in \mathcal{G}$ such that $\textcolor{red}{t \geq g}$,

$$
\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid G_g=1\right]=\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid G=\infty\right], \tag{1}
$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Note that, (1) is equivalent to the following: for  $t \in\{2, \ldots, T\}, g \in \mathcal{G}, t \geq g$ ,&lt;/p&gt;
&lt;p&gt;$$
\mathbb{E}\left[Y_t(\infty)-Y_{g-1}(\infty) \mid G_g=1\right]=\mathbb{E}\left[Y_t(\infty)-Y_{g-1}(\infty) \mid G=\infty\right], \tag{1&amp;rsquo;}
$$&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Parallel trends based on &amp;#39;Not-Yet-Treated&amp;#39; groups)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
For each $(s, t) \in\{2, \ldots, T\} \times\{2, \ldots, T\}, g \in \mathcal{G}$ such that $\textcolor{red}{t \geq g, s \geq t}$

$$
\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid G_g=1\right]=\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid D_s=0, G_g=0\right], \tag{2}
$$


  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Similarly, (2) is equivalent to (2&amp;rsquo;), that is, changing $Y_{t-1}(\infty)$ to $Y_{g-1(\infty)}$ in (2).&lt;/p&gt;
&lt;h2 id=&#34;3-identification-long-difference-estimands&#34;&gt;3. Identification: Long-Difference Estimands&lt;/h2&gt;
&lt;p&gt;Under no anticipation and one of the parallel-trends assumptions, each $\mathrm{ATT}(g,t)$ equals a simple &lt;mark&gt;&lt;strong&gt;long-difference DID&lt;/strong&gt;&lt;/mark&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Using never-treated&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
 $$
\operatorname{ATT}^{\text {never}}(g, t)=\mathbb{E}\left[Y_{t}-Y_{g-1} \mid G_g=1 \right]-\mathbb{E}\left[Y_{ t}-Y_{g-1} \mid G=\infty\right]
$$ 
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Using not-yet-treated&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
 $$
\operatorname{ATT}^{\text {not-yet}}(g, t)=\mathbb{E}\left[Y_{t}-Y_{g-1} \mid G_g=1 \right]-\mathbb{E}\left[Y_{ t}-Y_{g-1} \mid D_t = 0, G=\infty\right]
$$ 
&lt;figure style=&#34;text-align: center;&#34;&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250527100613626.png&#34; alt=&#34;image-20250527100613626&#34; style=&#34;zoom:50%;&#34; caption=&#34;sdf&#34;/&gt;
&lt;figcaption&gt;Why it’s called &lt;strong&gt;&#34;long difference&#34;&lt;/strong&gt;? Longer Time Span! &lt;/figcaption&gt;
&lt;/figure&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 2&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Q: Why it’s called &lt;strong&gt;&#34;long difference&#34;&lt;/strong&gt;? &lt;br&gt;&lt;/br&gt;

A: Longer Time Span! Check my plot above. The &#34;long difference&#34; refers to the fact that the comparison often spans from a pre-treatment period (i.e. $g-1$) to a later period (post-treatment, $t \ge g$), potentially covering multiple time periods. This contrasts with &#34;short differences,&#34; which might involve comparing outcomes in consecutive periods or shorter time windows.&lt;br&gt;&lt;/br&gt;

For each treated cohort, the method computes the difference in outcomes between the pre-treatment period and a specific post-treatment period, potentially far apart in time. This extended gap emphasizes the &#34;long&#34; aspect, as it captures the cumulative effect of the treatment over time.



  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Moreover, with covariates, one can form &lt;strong&gt;doubly-robust&lt;/strong&gt; estimators that combine generalized propensity scores $p_g(X_i)$ and outcome models $m_{g,t}(X_i)$. At high level, the form of this estimator is identical to AIPW–ATT in cross–section, but we need to replace &amp;ldquo;levels&amp;rdquo; (e.g. $Y$) with &amp;ldquo;changes&amp;rdquo; (e.g. $\Delta Y$).&lt;/p&gt;
&lt;p&gt;For more details, check Theorem 1 in the paper.&lt;/p&gt;
&lt;h2 id=&#34;4-aggregation-of-mathrmattgt&#34;&gt;4. Aggregation of $\mathrm{ATT}(g,t)$&lt;/h2&gt;
&lt;p&gt;Any overall summary $\theta$ is a weighted average of the cell-specific ATTs:&lt;/p&gt;
&lt;p&gt;$$
\theta=\sum_{g=2}^T \sum_{t=g}^T w_{g, t} \operatorname{ATT}(g, t), \quad \sum_{g, t} w_{g, t}=1 .
$$&lt;/p&gt;
&lt;p&gt;Common choices:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Cohort-heterogeneity: Average effect of participating in the treatment that units in group $g$ experienced,&lt;/li&gt;
&lt;/ol&gt;
 $$
\theta_S(g)=\frac{1}{T-g+1} \sum_{t=2}^T 1\{g \leq t\} \mathrm{ATT}(g,t)
$$ 
&lt;ol start=&#34;2&#34;&gt;
&lt;li&gt;
&lt;p&gt;Calendar time heterogeneity&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Event-study / dynamic treatment effects&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;5-limitation--extension&#34;&gt;5. Limitation &amp;amp; Extension&lt;/h2&gt;
&lt;p&gt;Lee &amp;amp; Wooldridge (2023) argue that Callaway &amp;amp; Sant’Anna (2021) method is &lt;strong&gt;less efficient&lt;/strong&gt; but more resilient to functional form of covariates. The following is from their working paper:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;hellip; CS (2021) method &lt;strong&gt;uses only the period just prior to the intervention&lt;/strong&gt; in defining the control group, thereby discarding potentially useful information in earlier time periods. &lt;br&gt;&lt;/br&gt;In fact, Wooldridge (2021) shows that, under the standard “error components” structure on the error, with a homoskedastic time-constant component and homoskedastic and serially uncorrelated idiosyncratic errors, the POLS estimator is both best linear unbiased (BLUE) and asymptotically efficient. These theoretical results imply that the CS (2021) estimators are inefficient under a standard set of assumptions. The simulations in Wooldridge (2021) bear this out, showing the CS approach can be very inefficient. Balanced against the loss in precision is that the CS approach can be less biased when parallel trends are violated.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;To improve the efficiency, instead of using long differences, Lee and Wooldridge (2023) use all suitable control observations in transforming the outcome variable. Specifically, &lt;mark&gt;rather than using the single period just prior to the treatment, $Y_{g-1}$, they use the &lt;strong&gt;pre-treatment average&lt;/strong&gt;, $\frac{1}{g-1}\sum_{s = 1}^{g-1}Y_s$&lt;/mark&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;rolling method&lt;/strong&gt; (Lee and Wooldridge, 2023) transforms panel data with staggered interventions by subtracting each unit&amp;rsquo;s average outcome across &lt;strong&gt;all pre-treatment periods&lt;/strong&gt; from their outcome in the current period of interest. This transformation, combined with &lt;strong&gt;no anticipation&lt;/strong&gt; and &lt;strong&gt;parallel trends&lt;/strong&gt; assumptions, makes the treatment assignment &lt;strong&gt;unconfounded&lt;/strong&gt; for the transformed outcome in each cohort/time cross-section. With unconfoundedness holding, we can then apply &lt;strong&gt;standard treatment effects estimators&lt;/strong&gt;, including &lt;strong&gt;doubly robust methods&lt;/strong&gt; and matching, utilizing &lt;strong&gt;all not-yet-treated units&lt;/strong&gt; as the valid control group for that specific cross-section.&lt;/p&gt;
&lt;h2 id=&#34;6-r-code-example&#34;&gt;6. R code Example&lt;/h2&gt;
&lt;p&gt;The following R code example is provided by Professor &lt;a href=&#34;https://github.com/scunning1975&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Scott Cunningham&lt;/a&gt;.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;readstata13&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;ggplot2&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;did&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# Callaway &amp;amp; Sant&amp;#39;Anna&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;data.frame&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;read.dta13&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;s&#34;&gt;&amp;#39;https://github.com/scunning1975/mixtape/raw/master/castle.dta&amp;#39;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;$&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;effyear&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;[is.na&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;$&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;effyear&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;]&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# untreated units have effective year of 0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Estimating the effect on log(homicide)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;att_gt&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;yname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;l_homicide&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# LHS variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;tname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;year&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# time variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;idname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;sid&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# id variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;gname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;effyear&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# first treatment period variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;data&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# data&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;xformla&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;NULL&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# no covariates&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;c1&#34;&gt;#xformla = ~ l_police, # with covariates&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;est_method&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;dr&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# &amp;#34;dr&amp;#34; is doubly robust. &amp;#34;ipw&amp;#34; is inverse probability weighting. &amp;#34;reg&amp;#34; is regression&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;control_group&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;nevertreated&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# set the comparison group which is either &amp;#34;nevertreated&amp;#34; or &amp;#34;notyettreated&amp;#34; &lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;bstrap&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;TRUE&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# if TRUE compute bootstrapped SE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;biters&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# number of bootstrap iterations&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;print_details&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;FALSE&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# if TRUE, print detailed results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;clustervars&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;sid&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# cluster level&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;panel&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;TRUE&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# whether the data is panel or repeated cross-sectional&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Aggregate ATT&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;agg_effects&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;aggte&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;type&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;group&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;summary&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;agg_effects&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Group-time ATTs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;summary&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Plot group-time ATTs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;ggdid&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Event-study&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;agg_effects_es&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;aggte&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;type&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;dynamic&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;summary&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;agg_effects_es&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Plot event-study coefficients&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;ggdid&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;agg_effects_es&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Callaway, Brantly and Pedro H. C. Sant’Anna (2021), “Difference-in-Differences with multiple time periods,” &lt;i&gt;Journal of Econometrics&lt;/i&gt;, Themed Issue: Treatment Effect 1, 225 (2), 200–230.&lt;/p&gt;
&lt;p&gt;Lee, S. J., &amp;amp; Wooldridge, J. M. (2023). A Simple Transformation Approach to Difference-in-Differences Estimation for Panel Data (SSRN Scholarly Paper No. 4516518). Social Science Research Network. &lt;a href=&#34;https://doi.org/10.2139/ssrn.4516518&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.2139/ssrn.4516518&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Sant’Anna, Pedro H. C. and Jun Zhao (2020), “Doubly robust difference-in-differences estimators,” &lt;i&gt;Journal of Econometrics&lt;/i&gt;, 219 (1), 101–22.&lt;/p&gt;
&lt;p&gt;How does doubly robust DiD estimator works? Check this: &lt;a href=&#34;https://psantanna.com/DiD/05_Covariates.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Lecture 5: How Covariates can make your DiD More Plausible&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://bcallaway11.github.io/did/index.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;did&lt;/a&gt; R package 📦&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Intuition for Doubly Robust Estimator</title>
      <link>https://chenxing.space/blog/intuition-for-doubly-robust-estimator/</link>
      <pubDate>Sun, 18 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/intuition-for-doubly-robust-estimator/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;To estimate ATE, we can either use outcome regression or inverse propensity weighting (IPW). While each approach has merits, combining them offers significant advantage – double robustness. In this post, I summarize the intuition for doubly robust estimator from Professor Ding&amp;rsquo;s textbook (Ding 2024), and connects this framework to debiased machine learning (DML) through Riesz representation theory. By understanding these connections, we can gain some insights into how modern causal inference methods effectively correct for bias in treatment effect estimation.&lt;/p&gt;
&lt;h2 id=&#34;two-characterizations-of-the-ate&#34;&gt;Two characterizations of the ATE&lt;/h2&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Basic setting)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
SUTVA, unconfoundedness and overlap

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Let $Y(0), Y(1)$ be potential outcomes and $D$ be a binary treatment variable. Consider ATE.&lt;/p&gt;
&lt;p&gt;First, we can use the outcome regression,&lt;/p&gt;

$$
\tau = \E\{\mu_1(X) - \mu_2(X) \},
$$

where

$$
\mu_1 = \E\{Y(1) \mid X \} = \E\{Y \mid D = 1, X \}, 
$$

$$
\mu_0 = \E\{Y(0) \mid X \} = \E\{Y \mid D = 0, X \}
$$

&lt;p&gt;Second, we can use the inverse propensity score weighting (IPW) approach,&lt;/p&gt;

$$\tau = \E\left\{\frac{DY}{e(X)} \right\} - \E\left\{\frac{(1-D)Y}{1-e(X)} \right\},$$

&lt;p&gt;where $e(X) = \P(D = 1 \mid X)$ is the propensity score.&lt;/p&gt;
&lt;p&gt;It completely ignores the outcome model. However, if there exist covariates $X$ that are predictive of $Y$, then even when the outcome model is misspecified, including it can reduce the variance compared to using IPW alone.&lt;/p&gt;
&lt;h2 id=&#34;key-insights&#34;&gt;Key Insights&lt;/h2&gt;
&lt;p&gt;This motivates combining the two approaches to:&lt;/p&gt;
&lt;p&gt;&lt;mark&gt;1. Reduce the variance of the IPW estimator&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;&lt;mark&gt;2. Reduce the bias of the outcome regression&lt;/mark&gt;&lt;/p&gt;
&lt;h3 id=&#34;reducing-the-variance&#34;&gt;Reducing the Variance&lt;/h3&gt;

$$
\mu_1=E\{Y(1)\}=E\left\{Y(1)-\mu_1\left(X, \beta_1\right)\right\}+E\left\{\mu_1\left(X, \beta_1\right)\right\} .
$$

&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;View &lt;mark&gt;$Y-\mu_1\left(X, \beta_1\right)$&lt;/mark&gt; as a &amp;ldquo;pseudo potential outcome&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Then apply IPW to it:&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

$$
\begin{aligned}
\mu_1 &amp; = E\left\{\frac{D\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}\right\}+E\left\{\mu_1\left(X, \beta_1\right)\right\} \\
&amp; =E\left\{\frac{D\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}+\mu_1\left(X, \beta_1\right)\right\},
\end{aligned}
$$

&lt;p&gt;Similarly,&lt;/p&gt;

$$
\begin{aligned}
\mu_0 &amp; =E\left\{\frac{(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}}{1-e(X)}\right\}+E\left\{\mu_0\left(X, \beta_0\right)\right\} \\
&amp; =E\left\{\frac{(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}}{1-e(X)}+\mu_0\left(X, \beta_0\right)\right\},
\end{aligned}
$$

&lt;p&gt;Notice that,&lt;/p&gt;

$$
\begin{aligned}
\mu_1 - \mu_0 &amp; = E\left\{\frac{D\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}+\mu_1\left(X, \beta_1\right)\right\} \\ &amp;\quad - E\left\{\frac{(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}}{1-e(X)}+\mu_0\left(X, \beta_0\right)\right\} 
\end{aligned}
$$

&lt;p&gt;$\mu_1 - \mu_2$ gives us the AIPW estimator.&lt;/p&gt;
&lt;h3 id=&#34;reducing-the-bias&#34;&gt;Reducing the Bias&lt;/h3&gt;

$$
\mu_1 =E\left\{\frac{Z\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}\right\}+E\left\{\mu_1\left(X, \beta_1\right)\right\} 
$$

&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;: We can view &lt;mark&gt;$Y-\mu_1\left(X, \beta_1\right)$&lt;/mark&gt; as the regression residuals, from which we apply IPW to extract useful signals. Alternatively, we can view &lt;mark&gt;$\mu_1\left(X, \beta_1\right) - Y$&lt;/mark&gt; as the &lt;strong&gt;bias&lt;/strong&gt;, which we then use IPW to &lt;strong&gt;correct the bias&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&#34;connecting-to-ddml-riesz-representation-for-bias-correction&#34;&gt;Connecting to DDML: Riesz Representation for Bias Correction&lt;/h3&gt;
&lt;p&gt;In the generic debiased framework (Chernozhukov, Newey, and Singh 2022), we leverage the &lt;mark&gt;&lt;strong&gt;Riesz representer to correct for bias&lt;/strong&gt;&lt;/mark&gt;, similar to the approach described above. This method parallels our use of IPW to extract signals from residuals, but specifically employs the Riesz representer for bias correction.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Data is $Z = \{Y, D, X\}$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $g$ be outcome regression, $g(D, X)=E[Y \mid D, X]$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Suppose moment is of the form: for some moment $m()$ that is &lt;strong&gt;linear in $g$&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

$$\theta-E[m(Z ; g)]=0,$$
$$\theta = E[m(Z ; g)] :=E[g(1, X)-g(0, X)]$$

&lt;p&gt;Then debiased version of the moment is of the form:&lt;/p&gt;
&lt;p&gt;$$
\theta-E[m(Z ; g)+a(X) \cdot(Y-g(X))]=0,
$$&lt;/p&gt;
&lt;p&gt;where&lt;/p&gt;
&lt;p&gt;$a(X)$ is the Riesz Representer of the linear functional $L(g) := E[m(Z ; g)]$. The existence of $a(X)$ is guaranteed by the &lt;a href=&#34;https://en.wikipedia.org/wiki/Riesz_representation_theorem&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Riesz representation theorem&lt;/a&gt;, for all square integrable $g(\cdot)$, we have&lt;/p&gt;
&lt;p&gt;$$
E[m(Z ; g)]=E[a(X) \cdot g(X)].
$$&lt;/p&gt;
&lt;p&gt;As we consider the ATE, the Riesz Representer are just some &amp;ldquo;inverse propensity score terms&amp;rdquo;,&lt;/p&gt;
&lt;p&gt;$$
a(D, X):=\frac{D}{e(X)}-\frac{(1-D)}{1-e(X)},
$$&lt;/p&gt;
&lt;p&gt;Consider the expression&lt;/p&gt;
&lt;p&gt;$$
E[m(Z ; g)+a(X) \cdot(Y-g(X))]
$$&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    The key intuition here is that &lt;mark&gt;$Y−g(X)$&lt;/mark&gt; represents the residual part, and we use the Riesz Representer $a(X)$ to &lt;strong&gt;correct this bias/residual term&lt;/strong&gt;.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;double-robust-estimation-of-att&#34;&gt;Double robust estimation of ATT&lt;/h3&gt;
&lt;p&gt;How about the double robust estimator for ATT? Can we also use the same idea to derive it? Yes!&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;( &amp;#34;one-sided&amp;#34; unconfoundedness and overlap)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$
D \indep Y(0) \mid X \text { and } e(X)&lt;1 .
$$

  &lt;/div&gt;
&lt;/div&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under the &#34;one-sided&#34; unconfoundedness and overlap assumption, 

$$
\E\{Y(0) \mid D=1\}=\frac{1}{e}\E\left\{\frac{e(X)}{1-e(X)} (1-D) Y\right\} \tag{1}
$$

and

$$
\tau_{\mathrm{T}}=\E(Y \mid D=1)-\E\left\{\frac{e(X)}{e} \frac{1-D}{1-e(X)} Y\right\},
$$

where $e=\P(D=1)$ is the marginal probability of the treatment.

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;We also have a doubly robust estimator for $\E\{Y(0) \mid D=1\}$  which combines the propensity score and the outcome models.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 2&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
     Define
$$
\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}} := \frac{1}{e} \E\left[ \frac{e(X, \alpha)}{1 - e(X, \alpha)}(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}+D \mu_0\left(X, \beta_0\right)\right] ,
$$

Under above Assumption, if either $e(X, \alpha)=e(X)$ or $\mu_0\left(X, \beta_0\right)=\mu_0(X)$, then $\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}}= \E\{Y(0) \mid D=1\}$.

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;How to come up with $\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}}$ ? Exactly the same idea as before!&lt;/p&gt;

$$
\E\{Y(0) \mid D=1\} = \E\{Y(0) - \mu_0\left(X, \beta_0\right) \mid D=1\} + \E\{\mu_0\left(X, \beta_0\right) \mid D=1\} 
$$

&lt;p&gt;Now, we can view &lt;mark&gt;$Y(0) - \mu_0\left(X, \beta_0\right)$&lt;/mark&gt; as a &amp;ldquo;pseudo potential outcome&amp;rdquo; under the control and apply eqn (1) to weight it, then we can get the form of $\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}}$ .&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Ding, Peng (2024) A first course in causal inference. &lt;a href=&#34;https://www.routledge.com/A-First-Course-in-Causal-Inference/Ding/p/book/9781032758626&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Chapman &amp;amp; Hall&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, and Rahul Singh (2022), “Automatic Debiased Machine Learning of Causal and Structural Effects,” &lt;i&gt;Econometrica&lt;/i&gt;, 90 (3), 967–1027.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Big Picture of Debiased Machine Learning</title>
      <link>https://chenxing.space/blog/big-picture-of-debiased-machine-learning/</link>
      <pubDate>Tue, 25 Mar 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/big-picture-of-debiased-machine-learning/</guid>
      <description>&lt;p&gt;Debiased machine learning (DML) is a generic recipe. The idea behind it is adding a correction term to the plug-in estimator of the functional, which leads to properties such as semi-parametric efﬁciency, double robustness, and Neyman orthogonality.&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250325151559540.png&#34; alt=&#34;image-20250325151559540&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;p&gt;(Auto)-DML is a &lt;strong&gt;Method-of-Moments estimator&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;debiased/orthogonal moment scores&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Why it matters?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;try to solve: model selection and/or &lt;strong&gt;regularization bias&lt;/strong&gt; from ML learners (e.g. Lasso)&lt;/li&gt;
&lt;li&gt;Neyman orthogonality: ensure the parameter of interest insentitive to first order perturbation of nuisance estimation&lt;/li&gt;
&lt;li&gt;double robustness&lt;/li&gt;
&lt;li&gt;asymptotic normality&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Idea: Debiasing is achieved by adding a correction term to the plug-in estimator of the functional&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Three representations: $\theta = \mathbb{E}[m(W,g)] = \mathbb{E}[Y\alpha(W)] = \mathbb{E}[g(W)\alpha(W)]$, where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$g()$ is outcome regression;&lt;/li&gt;
&lt;li&gt;$\alpha()$ is Rieze Representer (RR);&lt;/li&gt;
&lt;li&gt;$m()$ is a continuous linear functional;&lt;/li&gt;
&lt;li&gt;$W = (D, X)$ is data containing treatment $D$ and covariates $X$;&lt;/li&gt;
&lt;li&gt;$Y(d)$ is  potential outcome&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Correct the residual using RR&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;mark&gt;$\mathbb{E}\{m(W,g) - \theta + \alpha(W)[Y-g(W)]\} = 0$&lt;/mark&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to construct orthogonal moment function?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;orthogonal moment function = identifying moment function + first step influence function (FSIF)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;identifying moment function: $m(W,g) - \theta$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;involving &lt;strong&gt;outcome regression&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;FSIF: $\alpha(W)[Y-g(W)]$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;correct the residual using Rieze Representer (RR)&lt;/li&gt;
&lt;li&gt;Rieze Representer (RR)
&lt;ul&gt;
&lt;li&gt;In the case of ATE with binary treatment, RR are inverse propensity score terms&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RR can be automatically characterized&lt;/strong&gt;; NO NEED to know its analytical form&lt;/li&gt;
&lt;li&gt;Can use random forests and NNet learners of RR&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Double Robustness&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;mark&gt;$\mathbb{E}[m(W ; g) -\theta_0  \left.+\alpha(W)(Y-g(W))\right] =-\mathbb{E}\left[\left(\alpha-\alpha_0\right)\left(g-g_0\right)\right]$&lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The score will be zero in expectation when &lt;strong&gt;either&lt;/strong&gt; $\alpha(W) = \alpha_0(W)$ &lt;strong&gt;or&lt;/strong&gt; $g(W) = g_0(W)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Cross-fitting&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Why it matters?
&lt;ul&gt;
&lt;li&gt;Reduce overfitting bias&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Causal Survival Forest 🌲⏳</title>
      <link>https://chenxing.space/blog/causal-survival-forest-notes/</link>
      <pubDate>Thu, 05 Sep 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/causal-survival-forest-notes/</guid>
      <description>&lt;p&gt;In this post, I provide summary notes on the paper &amp;ldquo;Estimating Heterogeneous Treatment Effects with Right-Censored Data via Causal Survival Forests&amp;rdquo; by Cui et al. (2023).&lt;/p&gt;
&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;p&gt;&lt;mark&gt;How to estimate heterogeneous treatment effects with right-censored data?&lt;/mark&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous treatment effect (HTE) estimation&lt;/strong&gt; plays a central role in data-driven personalization&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Existing methods often can&amp;rsquo;t handle &lt;strong&gt;&lt;mark&gt;censored survival outcomes&lt;/mark&gt;&lt;/strong&gt;, common in medical/business applications&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;causal-survival-forests-csf&#34;&gt;Causal Survival Forests (CSF)&lt;/h2&gt;
&lt;p&gt;To address this challenge, the paper proposes causal survival forests (CSF)&lt;/p&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/1233.jpg&#34; alt=&#34;1233&#34; style=&#34;zoom:50%;&#34; /&gt;&lt;/center&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;An adaptation of the causal forest algorithm of Athey et al. (2019)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;It adjusts for censoring using doubly robust estimating equations developed in the survival analysis literature&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Advantages&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;mark&gt;&lt;strong&gt;Doubly robust&lt;/strong&gt;&lt;/mark&gt;, computationally tractable, and outperforms available baselines in our experiments&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Good statistical properties &amp;ndash; UCAN&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;setting-and-notation&#34;&gt;Setting and notation&lt;/h2&gt;
&lt;p&gt;Assume i.i.d tuples $\{X_i, T_i, C_i, W_i\}$, where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$X_i \in \mathcal{X}$ denote &lt;strong&gt;covariates&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;$T_i \in \mathbb{R}_{+}$ is the &lt;strong&gt;survival time&lt;/strong&gt; for $i$th unit&lt;/li&gt;
&lt;li&gt;$C_i \in \mathbb{R}_{+}$ is the &lt;strong&gt;censoring time&lt;/strong&gt; (the time at which $i$th unit gets censored)&lt;/li&gt;
&lt;li&gt;$W_i \in \{0,1\}$  denotes a &lt;strong&gt;binary treatment&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Using potential outcome framework, posit potential outcomes $\{T_i(1), T_i(0)\}$ s.t. $T_i = T_i(W_i)$, we need to estimate the conditional average treatment effect (CATE)&lt;/p&gt;
&lt;p&gt;$$
\tau(x)=\mathbb{E}\left[y(T_i(1)) - y(T_i(0)) \mid X_i=x\right],
$$&lt;/p&gt;
&lt;p&gt;where $y()$ is the outcome transformation. For example,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$y(T) = T$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$y(T) = T \wedge h$ for the &lt;strong&gt;restricted mean survival time (RMST)&lt;/strong&gt;; here $h$ is some chosen maximum considered time&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$y(T) = \1\{T \ge h\}$ for the &lt;strong&gt;survival probability&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;mark&gt;To estimate $\tau(x)$, the main challenge is that $T_i$ is not always observable. We &lt;strong&gt;can only observe&lt;/strong&gt;&lt;/mark&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;censored survival time&lt;/strong&gt;: $U_i = T_i \wedge C_i$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;non-censoring indicator&lt;/strong&gt;: $\Delta_i = 1\{T_i \le C_i\}$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Based on Assumption 1 (see later), we define the &lt;strong&gt;effective non-censoring indicator&lt;/strong&gt; as follows:&lt;/p&gt;

\begin{align}
\Delta_i^h &amp;= 1\{T_i \wedge h \le C_i\} \\
           &amp;\overset{(2)}{=} \Delta_i \vee 1\{U_i \ge h\}
\end{align}

&lt;p&gt;Note that, for the eqn (2), everything is observed. We can regard an observation with $\Delta_i^h = 1$ as a &lt;strong&gt;complete observation&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;assumptions&#34;&gt;Assumptions&lt;/h2&gt;
&lt;p&gt;In order to identify treatment effects, we need to rely on two sets of assumptions.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Assumption 2-4 enable us to identify the causal effect of $W_i$ on $T_i$ without censoring&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Assumption 5-6 is to guarantee that censoring due to $C_i$ does not break identification results&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Assumption 1 (Finite Horizon)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;$$
y(t) = y(h), \quad \forall \ t \ge h, \ 0&amp;lt;h&amp;lt;\infty
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 2 (Potential Outcomes)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;$$
\{T_i(1), T_i(0)\} \quad s.t. \quad T_i = T_i(W_i) \quad a.s.
$$ 
&lt;strong&gt;Assumption 3 (Ignorability)&lt;/strong&gt;&lt;/p&gt;
$$
\{T_i(1), T_i(0)\} \perp W_i \mid X_i
$$ 
&lt;p&gt;&lt;strong&gt;Assumption 4 (Overlap)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Propensity score $e(x) = \mathbb{P}(W_i = 1\mid X_i = x)$ is uniformly bounded away from 0 and 1,&lt;/p&gt;
&lt;p&gt;$$
\eta_e \le e(x)\le 1- \eta_e, \quad  0&amp;lt; \eta_e \le \frac{1}{2}
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 5 (Ignorable censoring)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Censoring is independent of survival time conditionally on treatment and covariates,&lt;/p&gt;
&lt;p&gt;$$
T_i \perp C_i \mid W_i, X_i
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 6 (Positivity)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;$$
\mathbb{P}(C_i &amp;lt; h | W_i, X_i) \le 1- \eta_c, \quad 0&amp;lt;\eta_c\le1
$$&lt;/p&gt;
&lt;h2 id=&#34;causal-forests-without-censoring&#34;&gt;Causal Forests Without Censoring&lt;/h2&gt;
&lt;p&gt;How does causal forest work?&lt;/p&gt;
&lt;p&gt;Essentially, we are running a “forest”-localized version of Robinson’s regression&lt;/p&gt;
&lt;p&gt;$$
\tau(x):=\operatorname{lm}\left(Y_i-\hat{m}^{(-i)}\left(X_i\right) \sim W_i-\hat{e}^{(-i)}\left(X_i\right), \text { weights }=\textcolor{blue}{\alpha_i(x)}\right),
$$&lt;/p&gt;
&lt;p&gt;where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$\textcolor{blue}{\alpha_i(x)}$ capture how &amp;ldquo;similar&amp;rdquo; a target sample $x$ is to each of the training samples $X_i$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\hat{m}$ and $\hat{e}$ are machine learning estimates for the outcome and propensity score models with cross-fitting&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Using notations in previous section, we estimate $\tau(x)$ by solving the following equation,&lt;/p&gt;
&lt;p&gt;$$
\sum_{i=1}^n \alpha_i(x) \psi_{\hat{\tau}(x)}^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right)=0 \tag{cf}
$$&lt;/p&gt;
&lt;p&gt;where,&lt;/p&gt;

\begin{aligned}
\psi_\tau^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right) &amp;= \left[ W_i - \hat{e}\left(X_i\right) \right] \quad \times\\
&amp; \left[ y\left(T_i\right) - \hat{m}\left(X_i\right) - \tau \left( W_i - \hat{e}\left(X_i\right) \right) \right]
\end{aligned}

&lt;p&gt;is the orthogonal &lt;strong&gt;complete&lt;/strong&gt; score function (shown as up-script $^{(c)}$),&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$e(x) = \mathbb{P}(W_i = 1\mid X_i = x)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$m(x) = \mathbb{E}(y(T_i) \mid X_i)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\hat{e}(X_i)$ and $\hat{m}(X_i)$ are estimates derived via cross-fitting&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;adjusting-for-censoring-via-weighting&#34;&gt;Adjusting for Censoring via Weighting&lt;/h2&gt;
&lt;p&gt;In the presence of censoring, the $T_i$ in equation (cf) is no longer observable.&lt;/p&gt;
&lt;p&gt;Simply ignoring censoring and building models on with complete observations (i.e. $\Delta_i^h = 1$) would lead to bias.&lt;/p&gt;
&lt;h2 id=&#34;simple-censoring-adjustment-via-ipcw&#34;&gt;Simple Censoring Adjustment via IPCW&lt;/h2&gt;
&lt;p&gt;Define the &lt;strong&gt;conditional survival function for censoring process&lt;/strong&gt; as
$$
S_w^C(s \mid x)=\mathbb{P}\left[C_i \geq s \mid W_i=w, X_i=x\right]
$$
We have,
$$
\mathbb{P}\left[\Delta_i^h=1 \mid X_i, W_i, T_i\right]=S_{W_i}^C\left(T_i \wedge h \mid X_i\right)
$$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;the LHS is the conditional probability of observing a complete observations (i.e. $\Delta_i^h = 1$)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the RHS is the conditional probability that censoring time is greater than survival time&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Does the above $\mathbb{P}(\Delta_i^h = 1 \mid \cdots)$ look like propensity score function?&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The main idea of IPCW estimation is to &lt;strong&gt;only consider complete cases&lt;/strong&gt;, but &lt;strong&gt;up-weight all complete observations&lt;/strong&gt; by $1/S_{W_i}^C\left(T_i \wedge h \mid X_i\right)$ to compensate for censoring.&lt;/p&gt;
&lt;p&gt;As a result, IPCW estimators succeed in eliminating censoring bias.&lt;/p&gt;
&lt;p&gt;With IPCW, we estimate $\tau(x)$ by solving the following equation,&lt;/p&gt;
&lt;p&gt;$$
\sum_{\left\{i: \Delta_i^h=1\right\}} \frac{\alpha_i(x)}{\hat{S}_{W_i}^C\left(T_i \wedge h \mid X_i\right)} \psi_{\hat{\tau}(x)}^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right)=0 \tag{IPCW}
$$ 
Let&amp;rsquo;s compare the equation (cf) v.s (IPCW),&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;For eqn (cf), we sum over all observations; for eqn (IPCW), we only sum over complete observations&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;For eqn (IPCW), we add $\frac{1}{1/S_{W_i}^C\left(T_i \wedge h \mid X_i\right)}$ as a part of weight&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For more details on IPCW, please check:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Chapter 8 and 12 in the textbook &lt;i&gt;Causal Inference: What If&amp;quot;&lt;/i&gt; (Hernán and Robins, 2020). In particular, &amp;ldquo;Ch 12.6 Censoring and missing data&amp;rdquo; is very helpful.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Chapter 21 &amp;ldquo;Treatment Heterogeneity with Survival Outcomes&amp;rdquo; in the textbook &lt;i&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/i&gt; (Zubizarreta et al., 2023)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;a-doubly-robust-correction&#34;&gt;A Doubly Robust Correction&lt;/h2&gt;
&lt;p&gt;Two limitations of IPCW approach:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Only use complete observations; throw away all observations with $\Delta_i^h = 0$, and this may hurt efficiency&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;IPCW-type methods are generally not robust to estimation errors; Neyman orthogonality condition (Chernozhukov et al. 2018) does not hold&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&#34;csf-method&#34;&gt;CSF Method&lt;/h3&gt;
&lt;p&gt;CSF method does not rely on IPCW. Instead, it relies on a more robust approach to making estimating equations robust to censoring.&lt;/p&gt;
&lt;p&gt;Recall the simplest case (without censoring), we have,&lt;/p&gt;
&lt;p&gt;$$
\sum_{i=1}^n \alpha_i(x) \psi_{\hat{\tau}(x)}^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right)=0 \tag{cf}
$$&lt;/p&gt;
&lt;p&gt;where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt; $\psi_{\hat{\tau}(x)}^{(c)} (\cdot)$  is the score function with &lt;mark&gt;&lt;strong&gt;the complete data&lt;/strong&gt;&lt;/mark&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;mark&gt;The idea of CSF is to convert the above (cf) equation into a &lt;strong&gt;censoring robust estimating equation&lt;/strong&gt; by using estimates of the survival and censoring processes.&lt;/p&gt;
&lt;p&gt;Now, we estimate the $\tau(x)$ by solving the following equation,&lt;/p&gt;
$$
\sum_{i=1}^n \alpha_i(x) \psi_{\hat{\tau}(x)}\left(X_i, y\left(U_i\right), U_i \wedge h, W_i, \Delta_i^h ; \hat{e}, \hat{m}, \hat{\lambda}_w^C, \hat{S}_w^C, \hat{Q}_w\right)=0,
$$ 
&lt;p&gt;where the score function is,&lt;/p&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20240919213740735.png&#34; alt=&#34;image-20240919213740735&#34; style=&#34;zoom:50%;&#34; /&gt;&lt;/center&gt;
&lt;p&gt;the &lt;strong&gt;conditional expectation of the transformed survival time&lt;/strong&gt; is defined as:&lt;/p&gt;
&lt;p&gt;$$
Q_w(s \mid x)=\mathbb{E}\left[y\left(T_i\right) \mid X_i=x, W_i=w, T_i \wedge h&amp;gt;s\right]
$$&lt;/p&gt;
&lt;p&gt;and the associated &lt;strong&gt;conditional hazard function&lt;/strong&gt; is defined as:&lt;/p&gt;
&lt;p&gt;$$
\lambda_w^{\mathrm{C}}(s \mid x)=-\frac{d}{d s} \log S_w^{\mathrm{C}}(s \mid x)
$$&lt;/p&gt;
&lt;p&gt;$\hat{Q}_w(s \mid x), \hat{S}_w^C(s \mid x)$ and $\hat{\lambda}_w^C(s \mid x)$ are cross-fit nuisance parameter estimates.&lt;/p&gt;
&lt;mark&gt;How to understand the above score function?&lt;/mark&gt;
&lt;blockquote&gt;
&lt;p&gt;The short answer is that that functional form emerges for the math (i.e., the desire for a doubly robust adjustment); and, unlike with the basic AIPW formula, it&amp;rsquo;s not as immediately intuitive.&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a href=&#34;#fn:1&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&#34;-key-points&#34;&gt;💡 KEY Points&lt;/h2&gt;
&lt;p&gt;We should think about the &lt;strong&gt;Neyman-orthogonal property&lt;/strong&gt;. In summary, CSF alleviates the drawbacks of IPCW so by taking the (complete-data) causal forest estimating equation $\psi_{\tau(x)}^{(c)}(T, W, \ldots)$ (the &amp;ldquo;R-learner&amp;rdquo;) and turn it into a censoring robust estimating equation $\psi_{\tau(x)}(Y, W, \ldots)$ by using estimates of the &lt;strong&gt;survival&lt;/strong&gt; and &lt;strong&gt;censoring processes&lt;/strong&gt;&lt;sup id=&#34;fnref:2&#34;&gt;&lt;a href=&#34;#fn:2&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;2&lt;/a&gt;&lt;/sup&gt;. (&amp;quot;&amp;hellip;&amp;quot; refers to additional nuisance parameters):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;censoring process: $P\left[C_i&amp;gt;t \mid X_i=x, W_i=w\right]$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;survival process: $P\left[T_i&amp;gt;t \mid X_i=x, W_i=w\right]$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;mark&gt;The upshot of this &amp;ldquo;orthogonal&amp;rdquo; estimating equation is that it will be &lt;strong&gt;consistent if either the survival or censoring process is correctly specified&lt;/strong&gt;,&lt;/mark&gt; which is very beneficial when we want to estimate these by modern ML tools, such as random survival forests.&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    CSF approach is doubly robust in the sense that we can obtain the consistent estimator either the survival or censoring process is correctly specified.
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;For more details, &lt;a href=&#34;https://www.degruyter.com/document/doi/10.2202/1557-4679.1052/html?lang=en&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Rubin &amp;amp; van der Laan (2007)&lt;/a&gt; and the chapter on RCTs with time-to-event data in &lt;a href=&#34;https://link.springer.com/chapter/10.1007/978-1-4419-9782-1_17&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Targeted Learning (2011)&lt;/a&gt; gives some more digestible details on doubly robust estimation with survival data.&lt;sup id=&#34;fnref:3&#34;&gt;&lt;a href=&#34;#fn:3&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;I also found the &lt;a href=&#34;https://gist.github.com/erikcs/cb8325fefe8bdfad6fc230015ddfe9cb&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;code&lt;/a&gt; in &lt;code&gt;grf&lt;/code&gt; GitHub repo helpful to understand the implement procedure. Specifically, check the following lines:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# The conditional survival function for the survival process.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;sf.survival&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;do.call&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;grf&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;::&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;survival_forest&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;c&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;list&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;cbind&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;W&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;args.nuisance&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# The conditional survival function for the censoring process.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;sf.censor&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;do.call&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;grf&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;::&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;survival_forest&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;c&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;list&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;cbind&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;W&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;args.nuisance&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;p&gt;Athey, Susan, Julie Tibshirani, and Stefan Wager. 2019. “Generalized Random Forests.” &lt;i&gt;The Annals of Statistics&lt;/i&gt; 47 (2): 1148–78. &lt;a href=&#34;https://doi.org/10.1214/18-AOS1709&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1214/18-AOS1709&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. 2018. “Double/Debiased Machine Learning for Treatment and Structural Parameters.” &lt;i&gt;The Econometrics Journal&lt;/i&gt; 21 (1): C1–68. &lt;a href=&#34;https://doi.org/10.1111/ectj.12097&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1111/ectj.12097&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Cui, Yifan, Michael R Kosorok, Erik Sverdrup, Stefan Wager, and Ruoqing Zhu. 2023. “Estimating Heterogeneous Treatment Effects with Right-Censored Data via Causal Survival Forests.” &lt;i&gt;Journal of the Royal Statistical Society Series B: Statistical Methodology&lt;/i&gt; 85 (2): 179–211. &lt;a href=&#34;https://doi.org/10.1093/jrsssb/qkac001&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1093/jrsssb/qkac001&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Hernán MA, Robins JM (2020). Causal Inference: What If. Boca Raton: Chapman &amp;amp; Hall/CRC.&lt;/p&gt;
&lt;p&gt;Zubizarreta, J. R., Stuart, E. A., Small, D. S., &amp;amp; Rosenbaum, P. R. (2023). &lt;i&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/i&gt;. CRC Press.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;This was suggested by Professor Wager in an email conversation.&amp;#160;&lt;a href=&#34;#fnref:1&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:2&#34;&gt;
&lt;p&gt;Check more on &lt;a href=&#34;https://grf-labs.github.io/grf/articles/survival.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;grf tutorial: Causal forest with time-to-event data&lt;/a&gt;&amp;#160;&lt;a href=&#34;#fnref:2&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:3&#34;&gt;
&lt;p&gt;Suggested by &lt;a href=&#34;https://sites.google.com/view/erikcs#h.uqgrvlridx4y&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Erik Sverdrup&lt;/a&gt;. Many thanks!&amp;#160;&lt;a href=&#34;#fnref:3&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>A walkthrough of how Causal Forest 🌲 works</title>
      <link>https://chenxing.space/blog/a-walkthrough-of-how-causal-forest-works/</link>
      <pubDate>Tue, 20 Aug 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/a-walkthrough-of-how-causal-forest-works/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;In this post, I will go over how causal forest works based on the &lt;a href=&#34;https://grf-labs.github.io/grf/articles/grf_guide.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;tutorial in grf R package&lt;/a&gt;. Causal Forests offer a flexible, data-driven approach to estimating varied treatment effects, bridging machine learning and causal inference techniques.&lt;/p&gt;
&lt;h2 id=&#34;common-setting&#34;&gt;Common Setting&lt;/h2&gt;
&lt;p&gt;If we are working on an &lt;strong&gt;observational&lt;/strong&gt; study, we have the following data:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Outcome variable: $Y_i$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Binary treatment indicator: $W_i = \{0, 1\}$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;A set of covariates: $X_i$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Let&amp;rsquo;s assume that the following conditions hold:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;mark&gt;(Assumption 1) $W_i$ is unconfounded given $X_i$ (i.e. treatment is as good as random given covariates).&lt;/mark&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$\{Y_i(0), Y_i(1)\} \perp W_i | X_i$$&lt;/p&gt;
&lt;ol start=&#34;2&#34;&gt;
&lt;li&gt;
&lt;mark&gt;(Assumption 2) The confounders $X_i$ have a linear effect on $Y_i$.&lt;/mark&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;mark&gt;(Assumption 3) The treatment effect $\tau$ is constant.&lt;/mark&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Then we could run a regression of the type&lt;/p&gt;
&lt;p&gt;$$
Y_i = \tau W_i + \beta X_i + \epsilon_i
$$&lt;/p&gt;
&lt;p&gt;and interpret the estimate of $\hat{\tau}$ as the average treatment effect (ATE) $\tau = \mathbb{E}(Y_i(1) - Y_i(0))$.&lt;/p&gt;
&lt;h2 id=&#34;relaxing-assumptions&#34;&gt;Relaxing Assumptions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Assumption 1&lt;/strong&gt; is an &amp;ldquo;identifying&amp;rdquo; assumption we have to live with&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Assumption 2&lt;/strong&gt; and &lt;strong&gt;Assumption 3&lt;/strong&gt; are modeling assumptions that we can question.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;relaxing-assumption-2-partially-linear-model-plr&#34;&gt;Relaxing Assumption 2: Partially Linear Model (PLR)&lt;/h3&gt;
&lt;p&gt;Assumption 2 is a strong parametric modeling assumption that requires the confounders to have a linear effect on the outcome, and that we should be able to relax by relying on semi-parametric statistics.&lt;/p&gt;
&lt;p&gt;We can instead posit the partially linear model:&lt;/p&gt;
&lt;p&gt;$$
Y_i = \tau W_i + f(X_i) + \epsilon_i, \ \ \ \mathbb{E}(\epsilon_i | X_i, W_i) = 0
$$&lt;/p&gt;
&lt;p&gt;How do we get around estimating $\tau$ when we do not know $f(X_i)$?&lt;/p&gt;
&lt;p&gt;Define the propensity score as
$$e(x) = \mathbb{E}(W_i | X_i = x),$$&lt;/p&gt;
&lt;p&gt;and the conditional mean of $Y$ as&lt;/p&gt;
&lt;p&gt;$$m(x) = \mathbb{E}(Y_i | X_i = x) = f(x) + \tau e(x).$$&lt;/p&gt;
&lt;p&gt;By Robinson (1988), we can rewrite the above equation in &lt;strong&gt;&amp;ldquo;centered&amp;rdquo;&lt;/strong&gt; form:&lt;/p&gt;
&lt;p&gt;$$Y_i - m(x) = \tau \cdot [W_i - e(x)]  + \epsilon_i$$&lt;/p&gt;
&lt;p&gt;This formulation has great practical appeal, as it means $\tau$ can be estimated by &lt;mark&gt;&lt;strong&gt;residual-on-residual regression&lt;/strong&gt;&lt;/mark&gt;.&lt;/p&gt;
&lt;p&gt;Good properties 😀:  Robinson (1988) shows that this approach yields root-n consistent estimates of $\tau$, even if estimates of $m(x)$ and $e(x)$ converge at a slower rate (&amp;ldquo;4-th root&amp;rdquo; in particular). This property is often referred to as &lt;strong&gt;orthogonality&lt;/strong&gt; and &lt;mark&gt;is a desirable property that essentially tells you that given noisy &amp;ldquo;nuisance&amp;rdquo; estimates ($m(x)$ and $e(x)$) you can still recover &amp;ldquo;good estimates of your target parameter ($\tau$).&lt;/mark&gt; For more details, please refer to Wager, Stefan “STATS 361: Causal Inference” Lecture 3.&lt;/p&gt;
&lt;p&gt;But how to estimate $m(x)$ and $e(x)$?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Use modern machine learning models!&lt;/strong&gt; One could use boosting, random forest, and etc to estimate $m(x)$ and $e(x)$ because what we need is just &amp;ldquo;reasonable accurate&amp;rdquo; predictions, i.e.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;$$\mathbb{E}\left[(\hat{m}(X)-m(X))^2\right]^{\frac{1}{2}}, \mathbb{E}\left[(\hat{e}(X)-e(X))^2\right]^{\frac{1}{2}}=o_P\left(\frac{1}{n^{1 / 4}}\right)$$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Issue with Direct Plug-in of Estimates&lt;/strong&gt;: Directly plugging in $\hat{m}(x)$ and $\hat{e}(x)$  into the residual-on-residual regression typically leads to bias because modern ML methods regularize to trade off bias and variance.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Solution via Cross-Fitting&lt;/strong&gt;: cross-fitting, where the prediction for observation $i$  is obtained without using unit  $i$  for estimation, can help overcome this bias (Chernozhukov et al. 2018).&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    Recap: We have a way to adopt the modern ML toolkit to &lt;em&gt;non-parametrically control for confounding&lt;/em&gt; when estimating an ATE, and still retain desirable statistical properties such as unbiased-ness and consistency.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;relaxing-assumption-3-non-constant-treatment-effects&#34;&gt;Relaxing Assumption 3: Non-constant treatment effects&lt;/h3&gt;
&lt;p&gt;Non-constant treatment effects occur when the impact of a treatment varies across different subgroups or individuals. This concept relaxes the assumption of homogeneous treatment effects, where the treatment is assumed to have the same impact on all units.&lt;/p&gt;
&lt;p&gt;We could specify certain subgroups and run separate regressions for each subgroup and obtain different estimates of $\tau$. To avoid false discoveries, &lt;mark&gt;this approach would require us to specify potential subgroups &lt;strong&gt;without&lt;/strong&gt; looking at the data.&lt;/mark&gt;
How can we use the data to inform us of potential subgroups?&lt;/p&gt;
&lt;p&gt;Let&amp;rsquo;s define,&lt;/p&gt;
&lt;p&gt;$$
Y_i=\textcolor{red}{ \tau\left(X_i\right) } W_i+f\left(X_i\right)+\varepsilon_i, \quad E\left[\varepsilon_i \mid X_i, W_i\right]=0,
$$&lt;/p&gt;
&lt;p&gt;where $\textcolor{red}{ \tau\left(X_i\right) }$ is the conditional ATE, i.e., $\textcolor{red}{ \tau\left(X_i\right) } :=\mathbb{E}\left[Y_i(1)-Y_i(0) \mid X_i\right]$. How do we estimate this?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;: If we imagine we had access to some neighborhood $\mathcal{N}(x)$ where $\tau$ was constant, we could proceed exactly as before, by doing a residual-on-residual regression on the samples belonging to $\mathcal{N}(x)$, i.e.:&lt;/p&gt;
&lt;p&gt;$$
\tau(x) := \operatorname{lm}\left(Y_i - \hat{m}^{(-i)}(X_i) \sim W_i - \hat{e}^{(-i)}(X_i), \text{ weights } = \mathbf{1}\{X_i \in \mathcal{N}(x)\} \right)
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;This is conceptually what Causal Forest does&lt;/strong&gt;, &lt;mark&gt;it estimates the treatment effect $\tau(x)$ for a target sample $X_i = x$ by running a weighted residual-on-residual regression on samples that have similar treatment effects.&lt;/mark&gt;&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    Recap: Causal Forest is running a a &amp;ldquo;forest&amp;rdquo;-localized version of Robinson’s regression.
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;These &lt;em&gt;weights&lt;/em&gt; play a crucial role, so how does &lt;code&gt;grf&lt;/code&gt; 📦 find them?&lt;/p&gt;
&lt;h2 id=&#34;random-forest-as-an-adaptive-neighborhood-finder&#34;&gt;Random forest as an adaptive neighborhood finder&lt;/h2&gt;
&lt;p&gt;Breiman’s random forest for predicting the conditional mean $m(x) = \mathbb{E}(Y_i | X_i = x)$ can be briefly summarized in two steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Building phase: Build $B$ trees which greedily place covariate splits that &lt;mark&gt;maximize the squared difference in subgroups means&lt;/mark&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$
n_L \cdot n_R \cdot ( \bar{y}_L - \bar{y}_R )^{2}
$$&lt;/p&gt;
&lt;ol start=&#34;2&#34;&gt;
&lt;li&gt;Prediction phase: Aggregate each tree’s prediction to form the final point estimate by averaging the outcomes $Y_i$ that fall into the same terminal leaf $L_b(X_i)$ as the targets sample $x$:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$\begin{align}
\hat{m}(x) &amp;amp;= \frac{1}{B} \sum_{b=1}^B \sum_{i=1}^n   \frac{Y_i \mathbf{1}\left\{X_i \in L_b(x)\right\}}{\left|L_b(x)\right|} \\
&amp;amp;= \sum_{i=1}^n \frac{1}{B} \sum_{b=1}^B Y_i \frac{\mathbf{1}\left\{X_i \in L_b(x)\right\}}{\left|L_b(x)\right|} \tag{1} \\
&amp;amp;= \sum_{i=1}^n Y_i \textcolor{blue}{\frac{1}{B} \sum_{b=1}^B  \frac{\mathbf{1}\left\{X_i \in L_b(x)\right\}}{\left|L_b(x)\right|}} \\
&amp;amp; =\sum_{i=1}^n Y_i \textcolor{blue}{\alpha_i(x)} \tag{2},
\end{align}$$&lt;/p&gt;
&lt;p&gt;Note that, this procedure is a double summation, first over trees, then over training samples (see equation (1)). We can swap the order of summation and obtain $\textcolor{blue}{\alpha_i(x)}$ in the equation (2). &lt;mark&gt;We have defined $\textcolor{blue}{\alpha_i(x)}$ as the frequency with which the $i$-th training sample falls into the same leaf as $x$.&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;Does above $\textcolor{blue}{\alpha_i(x)}$ remind you the traditional deterministic kernel function and bandwidth story?&lt;/p&gt;
&lt;p&gt;The following image illustrates how the $\textcolor{blue}{\alpha_i(x)}$ are calculated: Some dots are larger because they are used by all trees, while some are smaller because they are only used by a few trees.&lt;/p&gt;
&lt;figure&gt;
    &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/Screenshot%202024-08-20%20at%2020.38.30.png&#34; alt=&#34;weights from causal forests&#34; style=&#34;zoom:50%;&#34; /&gt;
    &lt;figcaption&gt;Figure 1: Weights from Causal Forests&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id=&#34;causal-forest&#34;&gt;Causal Forest&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Causal Forest&lt;/strong&gt; essentially combines Breiman (2001) and Robinson (1988) by modifying the steps above to:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Building phase: Greedily places covariate splits that maximize the squared difference in subgroup treatment effects $$n_L \cdot n_R \cdot ( \hat{\tau}_L - \hat{\tau}_R )^{2},$$ where $\hat{\tau}$ is obtained by running Robinson’s residual-on-residual regression for each possible split point.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use the resulting forest weights $\textcolor{blue}{\alpha_i(x)}$ to estimate $$\tau(x):=\operatorname{lm}\left(Y_i-\hat{m}^{(-i)}\left(X_i\right) \sim W_i-\hat{e}^{(-i)}\left(X_i\right), \text { weights }=\textcolor{blue}{\alpha_i(x)}\right),$$ where $\textcolor{blue}{\alpha_i(x)}$ capture how &amp;ldquo;similar&amp;rdquo; a target sample $x$ is to each of the training samples $X_i$&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That is, &lt;mark&gt;&lt;em&gt;Causal Forest&lt;/em&gt; is running a “forest”-localized version of Robinson’s regression.&lt;/mark&gt; This adaptive weighting (instead of leaf-averaging) coupled with some other forest construction details known as &lt;strong&gt;“honesty” and “subsampling”&lt;/strong&gt; can be used to give asymptotic guarantees for estimation and inference with random forests (Wager &amp;amp; Athey, 2018)&lt;/p&gt;
&lt;h2 id=&#34;efficiently-estimating-summaries-of-the-cates&#34;&gt;Efficiently estimating summaries of the CATEs&lt;/h2&gt;
&lt;p&gt;What about estimating summaries of $\tau(x)$, in terms of estimands like the average treatment effect (ATE), or the best linear projection (BLP), that have guaranteed $\sqrt{n}$ rate of convergence along with exact confidence intervals?&lt;/p&gt;
&lt;p&gt;For estimating ATE, the most intuitive approach is to average the CATE, i.e. $\frac{1}{n}\sum_{i=1}^n \tau(X_i)$, right? However, there are more efficient methods than simply averaging individual CATE estimates.&lt;/p&gt;
&lt;p&gt;Robins, Rotnitzky &amp;amp; Zhao (1994) showed that the so-called &lt;mark&gt;&lt;strong&gt;Augmented Inverse Probability Weighted (AIPW) estimator is asymptotically optimal&lt;/strong&gt; for $\tau$ (meaning that among all non-parametric estimators, it has the lowest variance).&lt;/mark&gt;&lt;/p&gt;
$$\begin{gathered}
\hat{\tau}_{AIPW}=\frac{1}{n} \sum_{i=1}^n\left(\hat{\mu}_{(1)}\left(X_i\right)-\hat{\mu}_{(0)}\left(X_i\right)+\frac{W_i}{\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(1)}\left(X_i\right)\right)\right. \\ \left.-\frac{1-W_i}{1-\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(0)}\left(X_i\right)\right)\right), 
\end{gathered}$$
&lt;p&gt;where $$\mu_{(w)}(x):=\mu(x, w) := \mathbb{E}\left[Y_i \mid X_i=x, W_i=w\right]$$  and $$e(x)=\mathbb{P}\left[W_i=1 \mid X_i=x\right]$$&lt;/p&gt;
&lt;p&gt;To interpret the AIPW estimator $\hat{\tau}_{AIPW}$ , it is helpful to decompose it into two components: Let $\hat{\tau}_{AIPW} = A + B$ , where&lt;/p&gt;
$$
A=\frac{1}{n} \sum_{i=1}^n\left(\hat{\mu}_{(1)}\left(X_i\right)-\hat{\mu}_{(0)}\left(X_i\right)\right)
$$
$$
B=\frac{1}{n} \sum_{i=1}^n\left(\frac{W_i}{\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(1)}\left(X_i\right)\right)-\frac{1-W_i}{1-\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(0)}\left(X_i\right)\right)\right)
$$
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$A$ represents the outcome regression adjustment estimator using $\hat{\mu}_{(w)}$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$B$ is an inverse propensity score weighting (IPW) estimator applied to the residuals $Y_i-\hat{\mu}_{\left(W_i\right)}\left(X_i\right)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The AIPW estimator utilizes propensity score weighting on the residuals to debias the direct estimate&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;One key property of the AIPW estimator is its &lt;strong&gt;&amp;ldquo;double robustness&amp;rdquo;&lt;/strong&gt;, which means that the estimator remains consistent and asymptotically normal even if either the outcome model or the propensity score model is misspecified. For proof, please refer to Stefan Wager&amp;rsquo;s Lecture 3 notes in “STATS 361: Causal Inference”.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The expression for above AIPW estimator can be rearranged and expressed as&lt;/p&gt;
$$
\begin{aligned}
&amp;\frac{1}{n} \sum_{i=1}^n\left(\tau\left(X_i\right) +
\textcolor{red}{\left[\frac{W_i}{e(X_i)} - \frac{1-W_i}{1-e(X_i)}\right]} 
\textcolor{violet}{\left[Y_i-\mu\left(X_i, W_i\right)\right]}

\right) \\

&amp; = \frac{1}{n} \sum_{i=1}^n\left(\tau\left(X_i\right)+\frac{W_i-e\left(X_i\right)}{e\left(X_i\right)\left[1-e\left(X_i\right)\right]}\left(Y_i-\mu\left(X_i, W_i\right)\right)\right) \\

&amp;\triangleq \frac{1}{n} \sum_{i=1}^n \Gamma_i 
\end{aligned}
$$
&lt;p&gt;We can understand above terms as:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;$\tau(X_i)$ is an initial treatment effect estimate&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\textcolor{red}{\left[\frac{W_i}{e(X_i)} - \frac{1-W_i}{1-e(X_i)}\right]}$ is the &lt;strong&gt;Riesz Representer&lt;/strong&gt; (Chernozhukov et al., 2022), which is used to &amp;ldquo;correct&amp;rdquo; the bias&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\textcolor{violet}{\left[Y_i-\mu\left(X_i, W_i\right)\right]}$ is the residual part&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\Gamma_i$ is called double robust score&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Chernozhukov, Victor, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. 2018. “Double/Debiased Machine Learning for Treatment and Structural Parameters.” &lt;i&gt;The Econometrics Journal&lt;/i&gt; 21 (1): C1–68. &lt;a href=&#34;https://doi.org/10.1111/ectj.12097&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1111/ectj.12097&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Athey, Susan, Julie Tibshirani, and Stefan Wager. 2019. “Generalized Random Forests.” &lt;i&gt;The Annals of Statistics&lt;/i&gt; 47 (2): 1148–78. &lt;a href=&#34;https://doi.org/10.1214/18-AOS1709&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1214/18-AOS1709&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Robinson, Peter M. “Root-N-consistent semiparametric regression.” Econometrica: Journal of the Econometric Society (1988): 931-954.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Wager, Stefan, and Susan Athey. “Estimation and inference of heterogeneous treatment effects using random forests.” Journal of the American Statistical Association 113.523 (2018): 1228-1242.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Wager, S. (2022). STATS 361: Causal Inference Lecture notes. Stanford University. &lt;a href=&#34;https://web.stanford.edu/~swager/stats361.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/stats361.pdf&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Chernozhukov, V., Newey, W. K., &amp;amp; Singh, R. (2022). Automatic Debiased Machine Learning of Causal and Structural Effects. &lt;i&gt;Econometrica&lt;/i&gt;, &lt;i&gt;90&lt;/i&gt;(3), 967–1027. &lt;a href=&#34;https://doi.org/10.3982/ECTA18515&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.3982/ECTA18515&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
  </channel>
</rss>
