<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>semiparametric efficiency | Chen Xing</title>
    <link>https://chenxing.space/tag/semiparametric-efficiency/</link>
      <atom:link href="https://chenxing.space/tag/semiparametric-efficiency/index.xml" rel="self" type="application/rss+xml" />
    <description>semiparametric efficiency</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Mon, 16 Feb 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>semiparametric efficiency</title>
      <link>https://chenxing.space/tag/semiparametric-efficiency/</link>
    </image>
    
    <item>
      <title>A Road Map of Nonparametric Efficiency in Causal Inference</title>
      <link>https://chenxing.space/blog/big/</link>
      <pubDate>Mon, 16 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/big/</guid>
      <description>&lt;p&gt;Causal inference promises to answer &amp;ldquo;what if&amp;rdquo; questions from observational data. But the standard approach — specify a parametric model, fit it, report the estimate — breaks down under misspecification. And misspecification is the norm, not the exception. This post draws from my notes on (Kennedy 2023). I walk through the core ideas of nonparametric efficiency theory — the framework that lets us pair flexible machine learning with rigorous inference. The central character is the &lt;mark&gt;&lt;strong&gt;Efficient Influence Function&lt;/strong&gt;&lt;/mark&gt;, a single mathematical object that tells us how to build estimators, why they work, and when we can trust their confidence intervals.&lt;/p&gt;
&lt;h2 id=&#34;1-motivation&#34;&gt;1. Motivation&lt;/h2&gt;
&lt;p&gt;First let&amp;rsquo;s keep the &lt;strong&gt;casual&lt;/strong&gt; and &lt;strong&gt;statistical&lt;/strong&gt; issues separate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Causal to Statistical:&lt;/strong&gt; The target causal parameter (e.g., ATE) can often be identified as a statistical &lt;strong&gt;functional&lt;/strong&gt; of the observed data distribution.&lt;/p&gt;
&lt;p&gt;$$\psi: \mathcal{P} \mapsto \mathbb{R}$$&lt;/p&gt;
&lt;p&gt;Once identified, the causal problem becomes a &lt;mark&gt;pure functional estimation problem&lt;/mark&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Problem with Parametrics:&lt;/strong&gt; Parametric models are likely misspecified, leading to biased estimates of the functional.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Problem with Naive Nonparametrics:&lt;/strong&gt; While we should use flexible machine learning (nonparametric) methods to avoid misspecification, simple &amp;ldquo;plug-in&amp;rdquo; estimators (fitting models and plugging predictions into the formula) fail because:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;They are generally not $\sqrt{n}$-consistent (converge too slowly).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;They do not yield valid Confidence Intervals.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;The Solution:&lt;/strong&gt; We need Efficiency Theory to understand the theoretical limit of performance and Influence Functions to construct estimators that reach that limit.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;2-efficiency-theory-lower-bounds&#34;&gt;2. Efficiency Theory (Lower Bounds)&lt;/h2&gt;
&lt;p&gt;Now let&amp;rsquo;s build the &amp;ldquo;best&amp;rdquo; estimator.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Benchmark:&lt;/strong&gt; We want to find the &lt;mark&gt;&lt;strong&gt;best possible performance&lt;/strong&gt;&lt;/mark&gt; (lowest Mean Squared Error) any estimator can achieve.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parametric Intuition:&lt;/strong&gt; In parametric models, the Cramer-Rao (CR) lower bound sets this benchmark: no unbiased estimator can have a variance lower than the inverse Fisher information.&lt;/p&gt;
&lt;p&gt;But now we are in the &amp;ldquo;nonparametric world&amp;rdquo;, how to exploit CR lower bound?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The &amp;ldquo;Submodel&amp;rdquo; Trick:&lt;/strong&gt; To apply CR bounds to infinite-dimensional nonparametric models:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;We define a &lt;em&gt;Parametric Submodel&lt;/em&gt;: a smooth, one-dimensional path through the complex nonparametric space that passes through the true distribution.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Logic:&lt;/em&gt; If we cannot estimate the parameter well in this simple &amp;ldquo;parametric&amp;rdquo; slice, we certainly cannot do it in the full nonparametric model.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Result:&lt;/em&gt; The hardest submodel defines the Local Minimax Lower Bound.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id=&#34;3-the-unified-approach-central-role-of-the-eif&#34;&gt;3. The Unified Approach: Central Role of the EIF&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;Efficient Influence Function (EIF)&lt;/strong&gt; is the &amp;ldquo;master key.&amp;rdquo; It is the derivative term in a Distributional Taylor Expansion&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a href=&#34;#fn:1&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt; (von Mises expansion):&lt;/p&gt;







&lt;figure class=&#34;blog-figure&#34;&gt;
   &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20260216113746294.png&#34; alt=&#34;image-20260216113746294&#34; style=&#34;zoom:50%;&#34; /&gt; 
  &lt;figcaption&gt;
    
      Figure 1: Distributional Taylor Expansion
    
  &lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This single expansion solves all three statistical tasks:&lt;/p&gt;
&lt;h3 id=&#34;a-how-to-construct-the-estimator-the-recipe&#34;&gt;A. How to construct the estimator? (The Recipe)&lt;/h3&gt;
&lt;p&gt;The expansion reveals that the bias of a naive plug-in estimator is roughly the expected value of the EIF.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Recipe:&lt;/em&gt; We construct a One-Step Estimator by taking the naive plug-in estimator and adding the empirical average of the estimated EIF to &amp;ldquo;de-bias&amp;rdquo; it.&lt;/p&gt;
&lt;p&gt;$$
\hat{\psi}_{\text{one-step}} = \text{Plug-in} + \frac{1}{n}\sum \text{EIF}(\text{Data})
$$&lt;/p&gt;
&lt;details&gt;
&lt;summary&gt;Click to expand: Deconstructing the One-Step Estimator&lt;/summary&gt;
&lt;h4 id=&#34;the-intuition-a-newton-raphson-correction&#34;&gt;The Intuition: A &amp;ldquo;Newton-Raphson&amp;rdquo; Correction&lt;/h4&gt;
&lt;p&gt;Think of it like Newton&amp;rsquo;s method in calculus: if you want to find the root of a function, you make an initial guess, calculate the derivative (slope) at that point, and use it to take a &amp;ldquo;step&amp;rdquo; closer to the true answer.&lt;/p&gt;
&lt;p&gt;In this context:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;Plug-in&lt;/em&gt; is your initial &amp;ldquo;guess&amp;rdquo; at the parameter.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;EIF&lt;/em&gt; acts as the &amp;ldquo;derivative&amp;rdquo; (gradient) that tells you which direction to move to fix the error.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;Formula&lt;/em&gt; represents taking that single &amp;ldquo;step&amp;rdquo; to correct the bias of your initial guess.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&#34;deconstructing-the-formula&#34;&gt;Deconstructing the Formula&lt;/h4&gt;
 
$$
\hat{\psi}_{\text{one-step}} = \underbrace{\psi(\hat{P})}_{\text{Plug-in}} + \underbrace{\frac{1}{n}\sum_{i=1}^n \phi(Z_i; \hat{P})}_{\text{Bias Correction}}
$$ 
&lt;p&gt;&lt;strong&gt;Part A: The Naive Plug-in&lt;/strong&gt; $\psi(\hat{P})$&lt;/p&gt;
&lt;p&gt;This is the estimate you get if you train your machine learning models (like regression or propensity scores), estimating the whole distribution $\hat{P}$, and simply &amp;ldquo;plug them in&amp;rdquo; to the formula for your parameter.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Why it fails:&lt;/em&gt; Flexible machine learning models trade bias for variance (regularization). They &amp;ldquo;smooth&amp;rdquo; the data to avoid overfitting. While this is good for predicting individual outcomes, it creates a first-order bias in the target parameter that does not vanish fast enough (slower than $1/\sqrt{n}$). If you stop here, your confidence intervals will be wrong.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Part B: The Bias Correction&lt;/strong&gt; $\frac{1}{n}\sum \text{EIF}$&lt;/p&gt;
&lt;p&gt;This term estimates the bias of the plug-in and removes it. The logic relies on the Distributional Taylor Expansion, which tells us that the error of the plug-in estimator is approximately equal to the negative expectation of the Influence Function:&lt;/p&gt;
&lt;p&gt;$$
\psi(\hat{P}) - \psi(P_{\text{true}}) \approx -\mathbb{E}[\text{EIF}]
$$&lt;/p&gt;
&lt;p&gt;Since the error is roughly $-\mathbb{E}[\text{EIF}]$, we can &amp;ldquo;cancel out&amp;rdquo; this error by adding the empirical average of the EIF estimated from our data. Note: if our initial estimate $\hat{P}$ were perfect (the truth), the average of the EIF would be exactly zero. The fact that this term is &lt;em&gt;not&lt;/em&gt; zero reflects the bias in our initial model that needs to be corrected.&lt;/p&gt;
&lt;h4 id=&#34;concrete-example-average-treatment-effect-ate&#34;&gt;Concrete Example: Average Treatment Effect (ATE)&lt;/h4&gt;
&lt;p&gt;To make this concrete, consider the ATE. This specific &amp;ldquo;One-Step&amp;rdquo; estimator is mathematically identical to the famous AIPW (Augmented Inverse Probability Weighting) or Doubly Robust estimator.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;Plug-in:&lt;/em&gt; You train a regression model to predict outcomes for treated vs. untreated groups. You calculate the average difference. The problem? If your regression is slightly wrong (which it always is), your effect estimate is biased.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;EIF Correction:&lt;/em&gt; You look at the residuals — the difference between what your model predicted and what actually happened, weighted by the propensity score (the probability of treatment). If your model consistently under-predicts outcomes for the treated group, the EIF term will be positive. Adding this average EIF to the plug-in &amp;ldquo;bumps&amp;rdquo; the estimate up, correcting the bias.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&#34;summary&#34;&gt;Summary&lt;/h4&gt;
&lt;p&gt;The &amp;ldquo;One-Step Estimator&amp;rdquo; recipe acknowledges that modern Machine Learning is great at learning patterns (the Plug-in) but bad at getting the total volume/area right (Bias). By adding the average of the Efficient Influence Function (the derivative), you explicitly calculate that missing &amp;ldquo;volume&amp;rdquo; and add it back in, ensuring the final estimate is unbiased and efficient.&lt;/p&gt;
&lt;/details&gt;
&lt;h3 id=&#34;b-how-to-analyze-the-estimator-the-proof&#34;&gt;B. How to analyze the estimator? (The Proof)&lt;/h3&gt;
&lt;p&gt;The expansion includes a &lt;em&gt;Remainder Term&lt;/em&gt; ($R_2$). To prove the estimator is $\sqrt{n}$-consistent, we must show this remainder is negligible (second-order).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Sample Splitting:&lt;/em&gt; We need to use cross-fitting to handle the &amp;ldquo;empirical process term&amp;rdquo; (preventing overfitting when estimating nuisance parameters).&lt;/p&gt;
&lt;h3 id=&#34;c-how-to-correct-the-bias-double-robustness&#34;&gt;C. How to correct the bias? (Double Robustness)&lt;/h3&gt;
&lt;p&gt;The math of the EIF reveals that the error depends on the &lt;em&gt;product of errors&lt;/em&gt; of the nuisance parameters (e.g., error in outcome regression $\times$ error in propensity score).&lt;/p&gt;
&lt;p&gt;This yields &lt;strong&gt;Double Robustness&lt;/strong&gt;: even if individual machine learning models converge slowly (e.g., $n^{-1/4}$), their product converges fast enough ($n^{-1/2}$) to allow for valid inference.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;4-the-link-between-eif-and-riesz-representer&#34;&gt;4. The Link Between EIF and Riesz Representer&lt;/h2&gt;
&lt;p&gt;The connection between the Efficient Influence Function (EIF) and the Riesz Representer (RR) is the bridge between &lt;em&gt;theoretical statistics&lt;/em&gt; and &lt;em&gt;practical machine learning&lt;/em&gt;.&lt;/p&gt;
&lt;h3 id=&#34;the-deconstruction-of-the-eif&#34;&gt;The Deconstruction of the EIF&lt;/h3&gt;
&lt;p&gt;We already know the EIF is the &amp;ldquo;correction term&amp;rdquo; we add to a naive plug-in estimator to fix bias. The Riesz Representer is the key component &lt;em&gt;inside&lt;/em&gt; that EIF.&lt;/p&gt;
&lt;p&gt;For a vast class of problems (linear functionals), the EIF always takes this specific structure:&lt;/p&gt;
 $$
\underbrace{\phi(Z)}_{\text{EIF}} = \underbrace{\psi(P)}_{\text{Plug-in}} + \underbrace{\alpha(X, W)}_{\text{Riesz Representer}} \times \underbrace{(Y - \mu(X, W))}_{\text{Outcome Residual}} 
$$ 
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$\mu(X, W)$: The conditional expectation (the outcome model, e.g., regression of $Y$ on $X, W$).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\alpha(X, W)$: The Riesz Representer. It tells us &lt;em&gt;how much weight&lt;/em&gt; to give each residual to correct the bias.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The EIF says: &amp;ldquo;To fix bias, look at your regression errors ($Y - \mu$).&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The RR says: &amp;ldquo;Here is exactly how to &lt;em&gt;weight&lt;/em&gt; those errors for this specific causal problem.&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;takeaway&#34;&gt;Takeaway&lt;/h2&gt;
&lt;p&gt;Here is the takeaway. Nonparametric efficiency theory is not an abstraction separate from practice — it is the practice. The EIF tells us what to estimate, the Riesz Representer tells us how to weight the errors, and double robustness explains why the whole thing holds together even when our models are imperfect. If we use AIPW, DML, or any doubly robust estimator, we are already using these ideas. The theory simply makes explicit what those estimators are doing and why they deserve our trust.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Kennedy, Edward H. (2023), “Semiparametric doubly robust targeted double machine learning: a review.”&lt;/p&gt;
&lt;p&gt;Williams, Nicholas T., Oliver J. Hines, and Kara E. Rudolph (2026), “Riesz representers for the rest of us.”&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, Victor Quintas-Martinez, and Vasilis Syrgkanis (2024), “Automatic debiased machine learning via riesz regression.”&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, and Rahul Singh (2022), “Automatic Debiased Machine Learning of Causal and Structural Effects,” &lt;i&gt;Econometrica&lt;/i&gt;, 90 (3), 967–1027.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;The Distributional Taylor Expansion is essentially the standard Taylor expansion applied at the distribution scale rather than the real-number scale. It is exactly equivalent to the concept of pathwise differentiability — both formalize the idea of taking a derivative of a statistical functional with respect to perturbations of the underlying distribution.&amp;#160;&lt;a href=&#34;#fnref:1&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>AIPW vs. Residual-on-Residual regression: Non-Parametric Flexibility or Efficiency?</title>
      <link>https://chenxing.space/blog/aipw-vs-residual-on-residual-regression-non-parametric-flexibility-or-efficiency/</link>
      <pubDate>Wed, 11 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/aipw-vs-residual-on-residual-regression-non-parametric-flexibility-or-efficiency/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Two powerful tools in causal inference are the &lt;strong&gt;Augmented Inverse Propensity Weighting (AIPW)&lt;/strong&gt; estimator and the &lt;strong&gt;Residual-on-Residual regression&lt;/strong&gt; estimator for partially linear models. Drawing from Wager’s notes (2024), this post breaks down how these estimators work, compares their strengths and weaknesses, and offers tips for when to use each.&lt;/p&gt;
&lt;h2 id=&#34;residual-on-residual-regression&#34;&gt;Residual-on-Residual regression&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;semiparamteric partially linear model&lt;/strong&gt; (PLM) assumes the outcome $ Y $ can be written as:&lt;/p&gt;
&lt;p&gt;$$ Y = \theta D + g(X) + \epsilon \tag{1}$$&lt;/p&gt;
&lt;p&gt;Here, $ D $ is a binary treatment, $ \theta $ is the causal parameter of interest, $ g(X) $ is some unknown function of covariates $ X $, and $ \epsilon $ is random noise with $\E[\epsilon \mid D, X] = 0$. The residual-on-residual estimator (or Robinson (1988) estimator) isolates $ \theta $ in two steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Partialing out covariates&lt;/strong&gt;: &amp;ldquo;Partialing out&amp;rdquo; $ X $ from $D$ and $Y$ using nonparametric regression and cross-fitting: $ \tilde{D} = D - \hat{\E}[D \mid X] $ and $ \tilde{Y} = Y - \hat{\E}[Y \mid X] $.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Estimate the effect&lt;/strong&gt;: Use linear regression $ \tilde{Y} \sim \tilde{D} $ to get $ \hat{\theta}$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    However, the model is &lt;strong&gt;not fully general&lt;/strong&gt;, because it  imposes a &lt;strong&gt;parametric specification&lt;/strong&gt; on the key component of interest. It imposes &lt;strong&gt;additivity&lt;/strong&gt; in $g(X)$ and $D$.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;augmented-inverse-propensity-weighting-aipw-estimator&#34;&gt;Augmented Inverse Propensity Weighting (AIPW) Estimator&lt;/h3&gt;
&lt;p&gt;AIPW takes a &lt;strong&gt;fully non-parametric&lt;/strong&gt; approach, aiming to estimate ATE, $ \tau = E[Y(1) - Y(0)] $, where $ Y(1) $ and $ Y(0) $ are potential outcomes under treatment and control.&lt;/p&gt;
 $$ \hat{\tau}_{\text{AIPW}} = \frac{1}{n} \sum_{i=1}^n \left[ \hat{\mu}_1(X_i) - \hat{\mu}_0(X_i) + D_i \frac{Y_i - \hat{\mu}_1(X_i)}{\hat{e}(X_i)} - (1 - D_i) \frac{Y_i - \hat{\mu}_0(X_i)}{1 - \hat{e}(X_i)} \right] $$ 
&lt;p&gt;Here, $ \hat{\mu}_1(X) $ and $ \hat{\mu}_0(X) $ are outcome regression estimators for treated and untreated units, and $ \hat{e}(X) $ is the estimator of propensity score. AIPW combines outcome modeling with inverse propensity weighting, making it &lt;strong&gt;doubly robust&lt;/strong&gt;: it’s consistent if &lt;em&gt;either&lt;/em&gt; the outcome model or propensity score is consistent.&lt;/p&gt;
&lt;h2 id=&#34;the-key-difference-non-parametric-vs-partially-linear&#34;&gt;The Key Difference: Non-Parametric vs. Partially Linear&lt;/h2&gt;
&lt;p&gt;&lt;mark&gt;&lt;strong&gt;AIPW is fully non-parametric&lt;/strong&gt;, imposing no specific parametric form on the treatment effect, while &lt;strong&gt;residual-on-residual regression estimator assumes a partially linear structure&lt;/strong&gt;&lt;/mark&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;AIPW’s flexibility&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;AIPW does not assume a specific parametric form!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AIPW is &lt;strong&gt;efficient&lt;/strong&gt; in the generic &lt;strong&gt;non-parametric&lt;/strong&gt; setting.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250611092857556.png&#34; alt=&#34;image-20250611092857556&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;/blockquote&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Residual-on-residual regression’s structure&lt;/strong&gt;: What partially linear assumption buys us is that residual-on-residual estimators that exploit this constraint can have &lt;strong&gt;smaller variance than AIPW&lt;/strong&gt;. In other words, adding this additional structure makes the residual-on-residual estimator more efficient than AIPW.&lt;/p&gt;
&lt;blockquote&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250611092145259.png&#34; alt=&#34;image-20250611092145259&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;/blockquote&gt;
&lt;div class=&#34;alert alert-warning&#34;&gt;
&lt;div&gt;
  A risk   of using the residual-on-residual estimator is that &lt;strong&gt;constant treatment effect&lt;/strong&gt; model (1) may be misspecified.
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
&lt;h2 id=&#34;why-this-matters&#34;&gt;Why This Matters&lt;/h2&gt;
&lt;p&gt;The choice between AIPW and residual-on-residual regression reflects a deeper trade-off in causal inference: &lt;strong&gt;flexibility versus efficiency&lt;/strong&gt;. AIPW’s non-parametric nature makes it a Swiss Army knife for complex data, while PLM structure is like a precision tool—effective when conditions are right.&lt;/p&gt;
&lt;p&gt;As Wager’s notes highlight:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Both AIPW and residual-on-residual regression are &lt;strong&gt;Neyman-orthogonal&lt;/strong&gt;, making them robust to first‐stage errors.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;However, their assumptions shape their performance. AIPW attains the lowest possible asymptotic variance for ATE under unconfoundedness. The residual-on-residual estimator, by imposing extra structure, can go beyond that bound when its structure is correct but at the cost of vulnerability to misspecification.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;conclusion&#34;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;AIPW’s fully non-parametric approach offers robustness and flexibility, while residual-on-residual regression’s partially linear structure prioritizes efficiency when assumptions hold.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Wager, S. (2024). &lt;em&gt;Causal inference: A statistical learning approach&lt;/em&gt;. &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/causal_inf_book.pdf&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Big Picture of Debiased Machine Learning</title>
      <link>https://chenxing.space/blog/big-picture-of-debiased-machine-learning/</link>
      <pubDate>Tue, 25 Mar 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/big-picture-of-debiased-machine-learning/</guid>
      <description>&lt;p&gt;Debiased machine learning (DML) is a generic recipe. The idea behind it is adding a correction term to the plug-in estimator of the functional, which leads to properties such as semi-parametric efﬁciency, double robustness, and Neyman orthogonality.&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250325151559540.png&#34; alt=&#34;image-20250325151559540&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;p&gt;(Auto)-DML is a &lt;strong&gt;Method-of-Moments estimator&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;debiased/orthogonal moment scores&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Why it matters?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;try to solve: model selection and/or &lt;strong&gt;regularization bias&lt;/strong&gt; from ML learners (e.g. Lasso)&lt;/li&gt;
&lt;li&gt;Neyman orthogonality: ensure the parameter of interest insentitive to first order perturbation of nuisance estimation&lt;/li&gt;
&lt;li&gt;double robustness&lt;/li&gt;
&lt;li&gt;asymptotic normality&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Idea: Debiasing is achieved by adding a correction term to the plug-in estimator of the functional&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Three representations: $\theta = \mathbb{E}[m(W,g)] = \mathbb{E}[Y\alpha(W)] = \mathbb{E}[g(W)\alpha(W)]$, where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$g()$ is outcome regression;&lt;/li&gt;
&lt;li&gt;$\alpha()$ is Rieze Representer (RR);&lt;/li&gt;
&lt;li&gt;$m()$ is a continuous linear functional;&lt;/li&gt;
&lt;li&gt;$W = (D, X)$ is data containing treatment $D$ and covariates $X$;&lt;/li&gt;
&lt;li&gt;$Y(d)$ is  potential outcome&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Correct the residual using RR&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;mark&gt;$\mathbb{E}\{m(W,g) - \theta + \alpha(W)[Y-g(W)]\} = 0$&lt;/mark&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to construct orthogonal moment function?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;orthogonal moment function = identifying moment function + first step influence function (FSIF)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;identifying moment function: $m(W,g) - \theta$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;involving &lt;strong&gt;outcome regression&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;FSIF: $\alpha(W)[Y-g(W)]$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;correct the residual using Rieze Representer (RR)&lt;/li&gt;
&lt;li&gt;Rieze Representer (RR)
&lt;ul&gt;
&lt;li&gt;In the case of ATE with binary treatment, RR are inverse propensity score terms&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RR can be automatically characterized&lt;/strong&gt;; NO NEED to know its analytical form&lt;/li&gt;
&lt;li&gt;Can use random forests and NNet learners of RR&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Double Robustness&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;mark&gt;$\mathbb{E}[m(W ; g) -\theta_0  \left.+\alpha(W)(Y-g(W))\right] =-\mathbb{E}\left[\left(\alpha-\alpha_0\right)\left(g-g_0\right)\right]$&lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The score will be zero in expectation when &lt;strong&gt;either&lt;/strong&gt; $\alpha(W) = \alpha_0(W)$ &lt;strong&gt;or&lt;/strong&gt; $g(W) = g_0(W)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Cross-fitting&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Why it matters?
&lt;ul&gt;
&lt;li&gt;Reduce overfitting bias&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Orthogonal vs Non-orthogonal Learning</title>
      <link>https://chenxing.space/blog/orthogonal-vs-non-orthogonal-learning/</link>
      <pubDate>Sat, 22 Feb 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/orthogonal-vs-non-orthogonal-learning/</guid>
      <description>


&lt;p&gt;The following exercise is from &lt;a href=&#34;https://colab.research.google.com/github/CausalAIBook/MetricsMLNotebooks/blob/main/PM2/r_orthogonal_orig.irnb&#34;&gt;CausalAIBook Notebook&lt;/a&gt;. The goal is to illustrate the difference between Neyman orthogonal vs non-orthogonal approach.&lt;/p&gt;
&lt;div id=&#34;setting&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;Setting&lt;/h3&gt;
&lt;p&gt;We compare the performance of the naive and orthogonal methods in a computational experiment where &lt;span class=&#34;math inline&#34;&gt;\(p=n=100\)&lt;/span&gt;, &lt;span class=&#34;math inline&#34;&gt;\(\beta_j = 1/j^2\)&lt;/span&gt;, &lt;span class=&#34;math inline&#34;&gt;\((\gamma_{DW})_j = 1/j^2\)&lt;/span&gt; and &lt;span class=&#34;math display&#34;&gt;\[Y = 1 \cdot D + \beta&amp;#39; W + \epsilon_Y\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;where &lt;span class=&#34;math inline&#34;&gt;\(W \sim N(0,I)\)&lt;/span&gt;, &lt;span class=&#34;math inline&#34;&gt;\(\epsilon_Y \sim N(0,1)\)&lt;/span&gt;, and &lt;span class=&#34;math display&#34;&gt;\[D = \gamma&amp;#39;_{DW} W + \tilde{D}\]&lt;/span&gt; where &lt;span class=&#34;math inline&#34;&gt;\(\tilde{D} \sim N(0,1)/4\)&lt;/span&gt;.&lt;/p&gt;
&lt;p&gt;Note that: &lt;span class=&#34;math inline&#34;&gt;\(\tilde{D}\)&lt;/span&gt; is the residual from the projection &lt;span class=&#34;math inline&#34;&gt;\(D \sim W\)&lt;/span&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The true treatment effect here is 1&lt;/strong&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;From the plots produced in this post (estimate minus ground truth), we show that the naive single-selection estimator is heavily biased (lack of Neyman orthogonality in its estimation strategy), while the orthogonal estimator based on partialling out, is approximately unbiased and Gaussian.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
&lt;div id=&#34;simulation&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;Simulation&lt;/h3&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;library(hdm)
library(ggplot2)

# Initialize constants
B &amp;lt;- 10000  # Number of iterations
n &amp;lt;- 100  # Sample size
p &amp;lt;- 100  # Number of features

# Initialize arrays to store results
Naive &amp;lt;- rep(0, B)
Orthogonal &amp;lt;- rep(0, B)


lambdaYs &amp;lt;- rep(0, B)
lambdaDs &amp;lt;- rep(0, B)

for (i in 1:B) {
  # Generate parameters
  beta &amp;lt;- 1 / (1:p)^2
  gamma &amp;lt;- 1 / (1:p)^2

  # Generate covariates / random data
  X &amp;lt;- matrix(rnorm(n * p), n, p)
  D &amp;lt;- X %*% gamma + rnorm(n) / 4

  # Generate Y using DGP
  Y &amp;lt;- D + X %*% beta + rnorm(n)

  # Single selection method
  rlasso_result &amp;lt;- hdm::rlasso(Y ~ D + X)  # Fit lasso regression
  sx_ids &amp;lt;- which(rlasso_result$coef[-c(1, 2)] != 0)  # Selected covariates

  # Check if any Xs are selected
  if (sum(sx_ids) == 0) {
    Naive[i] &amp;lt;- lm(Y ~ D)$coef[2]  # Fit linear regression with only D if no Xs are selected
  } else {
    Naive[i] &amp;lt;- lm(Y ~ D + X[, sx_ids])$coef[2]  # Fit linear regression with selected X otherwise
  }

  # Partialling out / Double Lasso

  fitY &amp;lt;- hdm::rlasso(Y ~ X, post = TRUE)
  resY &amp;lt;- fitY$res

  fitD &amp;lt;- hdm::rlasso(D ~ X, post = TRUE)
  resD &amp;lt;- fitD$res

  Orthogonal[i] &amp;lt;- lm(resY ~ resD)$coef[2]  # Fit linear regression for residuals
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;div id=&#34;making-a-nice-plot&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;Making a Nice Plot&lt;/h3&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;# Specify ratio
img_width &amp;lt;- 15
img_height &amp;lt;- img_width / 2

# Create a data frame for the estimates
df &amp;lt;- data.frame(
  Method = rep(c(&amp;quot;Naive&amp;quot;, &amp;quot;Orthogonal&amp;quot;), each = B),
  Value = c(Naive - 1, Orthogonal - 1)
)

# Create the histogram using ggplot2
hist_plot &amp;lt;- ggplot(df, aes(x = Value, fill = Method)) +
  geom_histogram(binwidth = 0.1, color = &amp;quot;black&amp;quot;, alpha = 0.7) +
  facet_wrap(~Method, scales = &amp;quot;fixed&amp;quot;) +
  labs(
    title = &amp;quot;Distribution of Estimates (Centered around Ground Truth)&amp;quot;,
    x = &amp;quot;Bias&amp;quot;,
    y = &amp;quot;Frequency&amp;quot;
  ) +
  geom_vline(xintercept = 0, color = &amp;quot;red&amp;quot;, linetype = &amp;quot;dashed&amp;quot;) +
  scale_x_continuous(breaks = seq(-2, 1.5, 0.5)) +
  theme_minimal() +
  theme(
    plot.title = element_text(hjust = 0.5),  # Center the plot title
    strip.text = element_text(size = 10),  # Increase text size in facet labels
    legend.position = &amp;quot;none&amp;quot;, # Remove the legend
    panel.grid.major = element_blank(),  # Make major grid lines invisible
    # panel.grid.minor = element_blank(),  # Make minor grid lines invisible
    strip.background = element_blank()  # Make the strip background transparent
  ) +
  theme(panel.spacing = unit(2, &amp;quot;lines&amp;quot;))  # Adjust the ratio to separate subplots wider

# Set a wider plot size
options(repr.plot.width = img_width, repr.plot.height = img_height)

# Display the histogram
print(hist_plot)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250222093132851.png&#34; alt=&#34;image-20250222093132851&#34; style=&#34;zoom:50%;&#34; /&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;: As we can see from the above bias plots (estimates minus the ground truth effect of 1), the double lasso procedure concentrates around zero whereas the naive estimator does not.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;reference&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;Reference&lt;/h3&gt;
&lt;p&gt;&lt;a href=&#34;https://colab.research.google.com/github/CausalAIBook/MetricsMLNotebooks/blob/main/PM2/r_orthogonal_orig.irnb&#34;&gt;CausalAIBook Notebook&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Semiparametric Models</title>
      <link>https://chenxing.space/blog/study-notes-on-semiparametric-models/</link>
      <pubDate>Wed, 05 Feb 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/study-notes-on-semiparametric-models/</guid>
      <description>&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Semiparametric models&lt;/strong&gt; contain both a finite-dimensional parameter of interest ($\theta$) and an infinite-dimensional nuisance parameter ($\eta$).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The goal is to estimate $\theta$ as efficiently as possible while filtering out the impact of $\eta$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;key-concepts&#34;&gt;Key Concepts&lt;/h2&gt;
&lt;h3 id=&#34;1-tangent-space&#34;&gt;1. Tangent Space&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The &lt;strong&gt;tangent space&lt;/strong&gt; consists of all possible local perturbations of the statistical model.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In parametric models, these directions are given by score functions (derivatives of the log-likelihood). In semiparametric models, the tangent space is typically an infinite-dimensional subspace of $L^2(P)$ (the space of square-integrable functions).&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;2-nuisance-tangent-space-mathcalt_eta&#34;&gt;2. Nuisance Tangent Space ($\mathcal{T}_{\eta}$)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This is the subset of the full tangent space that corresponds to variations in the nuisance parameter $\eta$, while holding the parameter of interest $\theta$ fixed.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;It represents all the directions in which the nuisance part of the model can change and potentially affect the estimation of $\theta$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;3-orthogonal-complement-of-the-nuisance-tangent-space-mathcalt_etaperp&#34;&gt;3. Orthogonal Complement of the Nuisance Tangent Space ($\mathcal{T}_{\eta}^\perp$)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Defined as:&lt;/p&gt;

  $$
  \mathcal{T}_{\eta}^\perp = \{ h \in L^2(P) : \langle h, g \rangle = 0 \quad \text{for all } g \in \mathcal{T}_{\eta} \}
  $$
  
&lt;p&gt;where the inner product $\langle \cdot, \cdot \rangle$ is typically given by covariance (or Fisher information).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;This space contains directions that are &amp;ldquo;free&amp;rdquo; of the influence of the nuisance parameter. &lt;mark&gt;In other words, any variation in this space does not get &amp;ldquo;contaminated&amp;rdquo; by changes in $\eta$.&lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;why-it-matters&#34;&gt;Why It Matters?&lt;/h2&gt;
&lt;h3 id=&#34;efficient-estimation&#34;&gt;Efficient Estimation&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;In semiparametric estimation, constructing an estimator with the smallest possible variance (i.e., achieving the efficiency bound) involves ensuring that its &lt;strong&gt;influence function&lt;/strong&gt; lies in $\mathcal{T}_{\eta}^\perp$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;strong&gt;influence function&lt;/strong&gt; describes how an estimator responds to small changes in the data distribution.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;By projecting any candidate influence function onto $\mathcal{T}_{\eta}^\perp$, one removes the component due to the nuisance parameter, yielding the &lt;strong&gt;efficient influence function&lt;/strong&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;practical-implication&#34;&gt;Practical Implication&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;This separation allows us to focus on the parameter of interest while systematically &amp;ldquo;filtering out&amp;rdquo; nuisance effects, leading to more precise (optimal) estimators.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;summary&#34;&gt;Summary&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Nuisance Tangent Space ($\mathcal{T}_{\eta}$)&lt;/strong&gt;: Captures all the directions of change due to the nuisance parameter $\eta$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Orthogonal Complement ($\mathcal{T}_{\eta}^\perp$)&lt;/strong&gt;: Contains directions free from nuisance effects, representing pure variations in $\theta$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Efficient Influence Function&lt;/strong&gt;: By projecting onto $\mathcal{T}_{\eta}^\perp$, one obtains an influence function that is optimal, meaning that the corresponding estimator achieves the semiparametric efficiency bound.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>From Donsker Classes to Neyman Orthogonality: The Power of DML</title>
      <link>https://chenxing.space/blog/from-donsker-classes-to-neyman-orthogonality-the-power-of-dml/</link>
      <pubDate>Fri, 26 Apr 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/from-donsker-classes-to-neyman-orthogonality-the-power-of-dml/</guid>
      <description>&lt;h3 id=&#34;motivation--intuition&#34;&gt;Motivation &amp;amp; Intuition&lt;/h3&gt;
&lt;p&gt;In classical semiparametric theory, we want to estimate a low‑dimensional target parameter (say, a treatment effect) while controlling for high‑dimensional nuisance functions (like nonparametric regressions). However, in order to use the central limit theorem (CLT) and to characterize the asymptotic behavior, classical results require that the space of functions in which these nuisance functions lie is “small” in a technical sense. &lt;mark&gt;In particular, they must form a &lt;strong&gt;Donsker class&lt;/strong&gt; — roughly speaking, a collection of functions whose complexity (measured via “entropy”) is bounded enough so that the empirical process converges to a Gaussian process. This condition, however, is too restrictive in modern applications where the nuisance functions are estimated by flexible machine learning methods&lt;/mark&gt; (e.g., random forests, boosting, deep neural nets) that may come from very large, high‑dimensional spaces.&lt;/p&gt;
&lt;p&gt;The double machine learning approach overcomes this problem by using &lt;strong&gt;Neyman orthogonal&lt;/strong&gt; scores. The key idea is that the moment functions used to estimate the target parameter are constructed in such a way that &lt;mark&gt;small errors in estimating the nuisance functions have only a second-order effect on the final estimator&lt;/mark&gt;. In plain language, even if your machine learning methods are “messy” or come from huge function classes (i.e. they do not satisfy the Donsker conditions), the estimation of your target parameter remains robust as long as the nuisance estimators converge at a certain rate. Moreover, the approach uses sample splitting (or cross‑fitting) to avoid overfitting biases.&lt;/p&gt;
&lt;h3 id=&#34;key-math-details&#34;&gt;Key Math Details&lt;/h3&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250220090505509.png&#34; alt=&#34;image-20250220090505509&#34; style=&#34;zoom:100%;&#34; /&gt;&lt;/center&gt;
&lt;h3 id=&#34;summary&#34;&gt;Summary&lt;/h3&gt;
&lt;p&gt;Flexibility in Nuisance Estimation: DML frees us from the need for nuisance estimators to lie in “small” Donsker classes. Thanks to Neyman orthogonality and cross-fitting, we can plug in flexible, machine-learning based nuisance estimates—even if they come from very rich function spaces—without contaminating the asymptotic distribution of the target parameter estimator.&lt;/p&gt;
&lt;p&gt;Thus, the contribution of DML is in allowing the use of complex, modern ML methods to estimate nuisance functions while still obtaining valid inference for the target parameter, bypassing the traditional, more restrictive Donsker conditions.&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
