<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>instrumental variables | Chen Xing</title>
    <link>https://chenxing.space/tag/instrumental-variables/</link>
      <atom:link href="https://chenxing.space/tag/instrumental-variables/index.xml" rel="self" type="application/rss+xml" />
    <description>instrumental variables</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Sat, 31 May 2025 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>instrumental variables</title>
      <link>https://chenxing.space/tag/instrumental-variables/</link>
    </image>
    
    <item>
      <title>Notes on causal inference with No Overlap – Regression Discontinuity</title>
      <link>https://chenxing.space/blog/notes-on-causal-inference-with-no-overlap-regression-discontinuity/</link>
      <pubDate>Sat, 31 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-causal-inference-with-no-overlap-regression-discontinuity/</guid>
      <description>&lt;p&gt;Here is my notes on regression discontinuity from Prof. Ding&amp;rsquo;s textbook (2024) and Prof. Imai&amp;rsquo;s lecture notes.&lt;/p&gt;
&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;p&gt;We often cannot run a randomized experiment and have to use/design observational studies to find a setting where credible causal inference is possible.&lt;/p&gt;
&lt;p&gt;The key is the knowledge of &lt;strong&gt;treatment assignment mechanism&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regression discontinuity design&lt;/strong&gt; (RD Design):&lt;/p&gt;
&lt;p&gt;RD Design is a simple and widely used &lt;strong&gt;quasi-experimental&lt;/strong&gt; design. The term “quasi experimental” is to emphasize that these approaches are still framed using concepts from randomized experiments but require econometric innovations to compensate for the lack of random treatment assignment.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Sharp RD Design&lt;/em&gt;: treatment assignment is based on a &lt;strong&gt;deterministic&lt;/strong&gt; rule (i.e. we have full knowledge of how treatment is assigned)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Fuzzy RD Design&lt;/em&gt;: &lt;strong&gt;encouragement to receive&lt;/strong&gt; treatment is based on a deterministic rule&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;setting&#34;&gt;Setting&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Binary treatment $Z\in \{0,1\}$  &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Potential outcomes $\{Y(0), Y(1)\}$  &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;There is a &lt;strong&gt;running variable&lt;/strong&gt; $X \in \R$ such that $Z=I\left(X \geq x_0\right)$, where $x_0$ is a pre-determined threshold. Note that, the treatment assignment is deterministic&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The unconfoundedness assumption holds automatically  $$
Z \indep \{Y(1), Y(0)\} \mid X
$$ &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The overlap assumption does not hold&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

$$
e(X)=\operatorname{pr}(Z=1 \mid X)=1\left(X \geq x_0\right) = \text{1 or 0}
$$

&lt;h2 id=&#34;identification&#34;&gt;Identification&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;RD can identify a &lt;strong&gt;local average causal effect&lt;/strong&gt; at the cutoff point $x_0$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Estimand:&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
  
$$
\tau\left(x_0\right)=E\left\{Y(1)-Y(0) \mid X=x_0\right\} .
$$

&lt;ul&gt;
&lt;li&gt;






&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(continuity assumption)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
1. $E\{Y(1) \mid X=x\}$ is continuous from the right at $x_0$ &lt;br&gt;
2. $E\{Y(0) \mid X=x\}$ is continuous from the left at $x_0$

  &lt;/div&gt;
&lt;/div&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250531113432572.png&#34; alt=&#34;image-20250531113432572&#34; style=&#34;zoom:20%;&#34; /&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;We have  $$
\begin{align}
E\left\{Y(1) \mid X=x_0\right\} &amp; =\lim _{\varepsilon \rightarrow 0+} E\left\{Y(1) \mid X=x_0+\varepsilon\right\} \tag{continuity} \\
&amp; =\lim _{\varepsilon \rightarrow 0+} E\left\{Y(1) \mid Z=1, X=x_0+\varepsilon\right\} \tag{def of Z}\\
&amp; =\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=1, X=x_0+\varepsilon\right),
\end{align}
$$  Similarly, $$
E\left\{Y(0) \mid X=x_0\right\}=\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=0, X=x_0-\varepsilon\right)
$$  So the local average causal effect at $x_0$ can be identified by the difference of the two limits&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Advantage: internal validity&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Disadvantage: external validity&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;key-theorem&#34;&gt;Key Theorem&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Assume that the treatment is determined by $Z=I\left(X \geq x_0\right)$ where $x_0$ is a predetermined threshold. Assume that $E\{Y(1) \mid X=x\}$ is continuous from the right at $x_0$ and $E\{Y(0) \mid X=x\}$ is continuous from the left at $x_0$. Then the local average treatment effect at $X=x_0$ is identified by

$$
\tau\left(x_0\right)=\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=1, X=x_0+\varepsilon\right)-\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=0, X=x_0-\varepsilon\right)
$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;$\tau\left(x_0\right)$ is nonparametrically identified.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Ding, P. (2024). A First Course in Causal Inference. CRC Press.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://imai.fas.harvard.edu/teaching/files/regression_discontinuity.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Lecture notes: Regression Discontinuity Design&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Control Function Method</title>
      <link>https://chenxing.space/blog/notes-on-control-function-method/</link>
      <pubDate>Sat, 31 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-control-function-method/</guid>
      <description>&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Endogeneity&lt;/strong&gt;: When an explanatory variable (treatment) is correlated with the unobserved error term (e.g., due to omitted variables, measurement error, or simultaneity).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consequence&lt;/strong&gt;: Standard regression (e.g., OLS) yields &lt;strong&gt;biased estimates&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Goal&lt;/strong&gt;: The control function (CF) approach &amp;ldquo;purges&amp;rdquo; endogeneity by modeling the correlation between the treatment and unobservables.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The control-function (CF) approach tackles this by &lt;strong&gt;explicitly modelling the source of endogeneity&lt;/strong&gt; and then partialling it out.  For linear models that modelling step turns out to be algebraically equivalent to 2SLS, but the real power of CF is that it &lt;strong&gt;extends seamlessly to nonlinear or limited‐dependent‐variable settings&lt;/strong&gt; where 2SLS cannot be applied directly.&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    Isolate the problematic part of the error term correlated with treatment, then add it as a control variable to &amp;rsquo;neutralize&amp;rsquo; endogeneity.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;how-it-works&#34;&gt;How It Works&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Outcome model&lt;/strong&gt;: $Y = \beta_0 + \beta_1 T + \beta_2 X + U$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$T$: Endogenous treatment&lt;/li&gt;
&lt;li&gt;$X$: Exogenous controls&lt;/li&gt;
&lt;li&gt;$U$: Unobservables (correlated with $T$).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;First-stage model&lt;/strong&gt;: $ T = \gamma_0 + \gamma_1 Z + \gamma_2 X + V $&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$Z$: Instrumental variable (IV)&lt;/li&gt;
&lt;li&gt;$V$: First-stage error.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;CF Insight&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;If $U$ and $V$ are correlated (i.e., $\text{Cov}(U,V) \neq 0$), we can decompose $U$ into:&lt;br&gt;
$$ U = \rho V + \epsilon $$
where $\epsilon$ is uncorrelated with $T$ and $V$ (by construction).&lt;/p&gt;
&lt;h2 id=&#34;estimation&#34;&gt;Estimation&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;First Stage&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Regress $T$ on $Z$ and $X$:&lt;br&gt;
$$ \hat{T} = \hat{\gamma}_0 + \hat{\gamma}_1 Z + \hat{\gamma}_2 X $$&lt;/li&gt;
&lt;li&gt;Obtain the residual: $\hat{V} = T - \hat{T}$.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Second Stage&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Add $\hat{V}$ to the outcome model:&lt;br&gt;
$$ Y = \beta_0 + \beta_1 T + \beta_2 X + \rho \hat{V} + \epsilon $$&lt;/li&gt;
&lt;li&gt;Estimate via OLS.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$\hat{V}$ &amp;ldquo;controls for&amp;rdquo; the part of $U$ correlated with $T$.&lt;/li&gt;
&lt;li&gt;Once $\hat{V}$ is included, $T$ becomes &lt;em&gt;exogenous&lt;/em&gt; in the modified model ($\text{Cov}(T, \epsilon) = 0$).&lt;/li&gt;
&lt;li&gt;$\hat{\beta}_1$ is consistent for the causal effect.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;key-assumptions&#34;&gt;Key Assumptions&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Instrument Validity&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$Z$ is relevant: $\text{Cov}(Z, T) \neq 0$ (strong first stage).&lt;/li&gt;
&lt;li&gt;$Z$ is exogenous: $\text{Cov}(Z, U) = 0$.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Correct Functional Form&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Linearity in the first stage and control function (e.g., $U = \rho V + \epsilon$).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Exclusion Restriction&lt;/strong&gt;: $Z$ affects $Y$ only through $T$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;cf-vs-other-methods&#34;&gt;CF vs. Other Methods&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Method&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Key Difference&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2SLS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Uses fitted values ($\hat{T}$); efficient but inconsistent under heteroscedasticity/nonlinearity.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Control Function&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Uses residuals ($\hat{V}$); more flexible for nonlinear models (e.g., probit).&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Advantage of CF&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Directly models the endogeneity structure (via $\hat{V}$).&lt;/li&gt;
&lt;li&gt;Extends to non-additive errors, discrete outcomes, and heteroscedastic settings.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;example&#34;&gt;Example&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Problem&lt;/strong&gt;: Estimate returns to education ($T$) on wages ($Y$), where ability ($U$) is unobserved and correlated with education.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;IV&lt;/strong&gt;: Distance to college ($Z$).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Steps&lt;/strong&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Regress: education $\sim$ distance + controls → get residuals $\hat{V}$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Regress: wages $\sim$ education + controls + $\hat{V}$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Result&lt;/strong&gt;: Coefficient on education is causal.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;generalization&#34;&gt;Generalization&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Control-function term&lt;/th&gt;
&lt;th&gt;Typical reference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Binary/probit&lt;/strong&gt; $Y$ with endogenous $T$&lt;/td&gt;
&lt;td&gt;Include $h(\hat v)$ where $h(\cdot)$ is the generalized residual (Smith &amp;amp; Blundell, 1986)&lt;/td&gt;
&lt;td&gt;Rivers &amp;amp; Vuong (1988)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Heckman sample selection&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Inverse Mills ratio $\lambda(\hat v)$ controls for non-random sample entry&lt;/td&gt;
&lt;td&gt;Heckman (1979)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Count models (Poisson, NB)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Nonparametric sieve $h(\hat v)$ or parametric polynomial&lt;/td&gt;
&lt;td&gt;Wooldridge (2015)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Semiparametric/ML partially-linear&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Learnt nuisance $m(X)$; CF gives Neyman-orthogonal moment for Double ML&lt;/td&gt;
&lt;td&gt;Chernozhukov et al. (2018)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The principle is identical: obtain a residual that captures unobserved heterogeneity driving $T$, then include it (or a flexible transformation) in the structural equation.&lt;/p&gt;
&lt;h2 id=&#34;why-this-matters&#34;&gt;Why This Matters&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;CF is &lt;strong&gt;essential&lt;/strong&gt; when:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;You have a valid IV and suspect omitted variable bias.&lt;/li&gt;
&lt;li&gt;You work with nonlinear models (e.g., binary/duration outcomes).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Software Implementation&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Stata: &lt;code&gt;ivregress 2sls&lt;/code&gt; (equivalent to CF in linear cases) or &lt;code&gt;cmp&lt;/code&gt; for nonlinear.&lt;/li&gt;
&lt;li&gt;R: &lt;code&gt;ivreg&lt;/code&gt; (linear), &lt;code&gt;controlfunction&lt;/code&gt; package.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;critical-caveats&#34;&gt;Critical Caveats&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Weak Instruments&lt;/strong&gt;: If $Z$ is weak, $\hat{V}$ is noisy → bias.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Functional Form Misspecification&lt;/strong&gt;: If $U \neq \rho V + \epsilon$, CF fails.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No Magic Bullet&lt;/strong&gt;: Validity of $Z$ is untestable and must be justified theoretically.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;bottom-line&#34;&gt;Bottom Line&lt;/h2&gt;
&lt;p&gt;The CF approach harnesses IV residuals to &amp;ldquo;control&amp;rdquo; for endogeneity, converting an endogenous variable into a conditionally exogenous one. It’s a blend of IV intuition and regression control — powerful when assumptions hold.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on IV Free Methods</title>
      <link>https://chenxing.space/blog/notes-on-iv-free-methods/</link>
      <pubDate>Fri, 30 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-iv-free-methods/</guid>
      <description>&lt;p&gt;Here’s a overview of instrument-free methods:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Latent Instrumental Variable (LIV)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Gaussian Copula (GC)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The goal is for dealing with endogeneity when no external instruments are available.&lt;/p&gt;
&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;p&gt;Finding good instruments is hard. Around 2010, marketing researchers recognized that the “IV cure can be worse than the endogeneity disease,” and began seeking ways to exploit features of the data (rather than external IVs) to identify causal effects when valid instruments aren’t available .&lt;/p&gt;
&lt;h2 id=&#34;intuition&#34;&gt;Intuition&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;LIV&lt;/strong&gt; treats the endogenous regressor as driven by a discrete, unobserved “latent instrument” that captures its exogenous variation, while the remaining variation is deemed endogenous. One then estimates this latent class structure (akin to a mixture model) alongside the main outcome equation under &lt;strong&gt;normal-error assumptions&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GC&lt;/strong&gt; builds a joint distribution of the endogenous regressor and the structural error via a &lt;strong&gt;Gaussian copula&lt;/strong&gt;. By assuming the regressor is non-normal and the error is normal, the copula decomposition isolates the exogenous component and yields consistent estimates without explicit instruments .&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;data-structure&#34;&gt;Data Structure&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Both methods were originally developed for &lt;strong&gt;cross-sectional&lt;/strong&gt; datasets.&lt;/li&gt;
&lt;li&gt;LIV has been extended to dynamic settings (e.g., panel data with time-varying latent classes) and to nonlinear models (e.g., binary logit) .&lt;/li&gt;
&lt;li&gt;Applying GC in &lt;strong&gt;panel&lt;/strong&gt; contexts typically requires a first-difference or within transformation to purge fixed effects, which alters the error covariance and complicates copula estimation.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;how-it-works&#34;&gt;How It Works&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;LIV Approach&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model setup&lt;/strong&gt;: Decompose the regressor $X$ into two parts: a discrete latent instrument $Z^*$ (exogenous) and residual $u$ (endogenous).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Assumptions&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;$Z^*$ has a finite number of states&lt;/li&gt;
&lt;li&gt;Structural errors are normally distributed&lt;/li&gt;
&lt;li&gt;$X$ is non-normal&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Estimation&lt;/strong&gt;: Use an EM algorithm (or maximum likelihood) to jointly recover $Z^*$, the mixture probabilities, and the outcome regression parameters.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gaussian Copula&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model setup&lt;/strong&gt;: Specify a joint distribution of $(X,;\varepsilon)$ via a Gaussian copula linking marginal distributions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Assumptions&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;$X$ is non-normal&lt;/li&gt;
&lt;li&gt;$\varepsilon$ is normal&lt;/li&gt;
&lt;li&gt;Dependence captured entirely by the copula correlation parameter&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Estimation&lt;/strong&gt;: Fit marginal distributions and copula correlation by maximization of the joint likelihood, then recover the causal effect from the structural equation conditional on the estimated copula.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;pros-and-cons&#34;&gt;Pros and Cons&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Pros&lt;/th&gt;
&lt;th&gt;Cons&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;LIV&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;• Avoids need for external IVs&lt;br&gt;• Leverages latent class structure&lt;br/&gt;• Extensions to dynamic and nonlinear models exist&lt;/td&gt;
&lt;td&gt;• Relies on an untestable assumption that exogenous variation is discrete&lt;br/&gt;• Normal-error and non-normal-regressor assumptions&lt;br/&gt;• Model complexity and identification hinge on number of latent states&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GC&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;• Simple implementation (standard likelihood techniques)&lt;br/&gt;• No need for instruments beyond distributional assumptions&lt;/td&gt;
&lt;td&gt;• Sensitive to skewness in $X$ and to deviations from normality in $\varepsilon$&lt;br/&gt;• Requires large samples for stable estimates&lt;br/&gt;• Panel data require transformations that complicate the copula structure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&#34;key-takeaways&#34;&gt;Key takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Instrument-free methods can be powerful when valid external instruments are unavailable, but they substitute one set of &lt;strong&gt;strong (often untestable) assumptions&lt;/strong&gt; for another.&lt;/li&gt;
&lt;li&gt;Careful diagnostic checks—testing residual normality, examining the distribution of $X$, and sensitivity analyses—are essential to ensure credible inference.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Park, Sungho and Sachin Gupta (2012), “Handling Endogenous Regressors by Joint Estimation Using Copulas,” &lt;i&gt;Marketing Science&lt;/i&gt;, 31 (4), 567–86.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.ama.org/marketing-news/a-review-of-copula-correction-methods-to-address-regressor-error-correlation/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;A Review of Copula Correction Methods to Address Regressor–Error Correlation&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Instrumental Variables</title>
      <link>https://chenxing.space/blog/notes-on-instrumental-variables/</link>
      <pubDate>Thu, 29 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-instrumental-variables/</guid>
      <description>&lt;p&gt;Here are my notes on instrumental variables from &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Stefan&amp;rsquo;s lecture materials&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;p&gt;How can we identify causal eﬀects $W \rightarrow Y$ when we are in the presence of unobserved confounding $U$?&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250529155826927.png&#34; alt=&#34;image-20250529155826927&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;p&gt;One popular way is to ﬁnd and use &lt;strong&gt;instrumental variables&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;partially-linear-iv-models&#34;&gt;Partially Linear IV Models&lt;/h2&gt;
&lt;p&gt;When instrumental variables are available, it becomes possible to point identify causal effects in &lt;strong&gt;partially linear models&lt;/strong&gt; and &lt;mark&gt;certain types of causal effects in nonlinear models&lt;/mark&gt;.&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    The keyword in this section is &lt;strong&gt;linear, constant effect IV&lt;/strong&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Here we begin with partially linear models.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(partially linear)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    

Assume the structural equation for $Y$ is linear: 
$$
Y=f_Y\left(W, U, \varepsilon_Y\right)=\alpha+W \tau+\varepsilon \tag{PLM},
$$

where $\epsilon$ is an error term that captures the contribution of both $U$ and $\epsilon_Y$. 

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;This is a semiparametric specification, in that we impose a &lt;mark&gt;&lt;strong&gt;linear relationship&lt;/strong&gt;&lt;/mark&gt; between $W$ and $Y$ but let the rest be non-parametric. Notice that, besides the linearity assumption, we also assume a &lt;mark&gt;&lt;strong&gt;constant treatment effect&lt;/strong&gt;&lt;/mark&gt; (i.e. $\tau$), which is also a strong assumption.&lt;/p&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(partialling out X)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
I ignore the role of covariates $X$ and the possible extension like $$
\varepsilon_i \indep Z_i \mid X_i
$$
because we can recover the same DAG after &lt;mark&gt;partialling out&lt;/mark&gt; observed confounder $X$. As an illustration, we can use $\tilde{Y}, \tilde{W}, \tilde{Z}$, where $$\tilde{V} = V - \E[V \mid X]$$

  &lt;/div&gt;
&lt;/div&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(relax linearity assumption)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
The key &lt;strong&gt;limitation&lt;/strong&gt; here is that we assume the &lt;strong&gt;linear structure and constant treatment effect&lt;/strong&gt;. Later, we will consider IV estimator without the linearity assumption.
&lt;br&gt;&lt;/br&gt;

When the constant treatment eﬀect model (PLM) doesn’t hold, the average treatment eﬀect $\tau_{ATE} = \E [Y_i (1) − Y_i (0)]$ is &lt;mark&gt;NOT identified&lt;/mark&gt; without more data, because we don’t have any observations on treated never takers, etc. Without linearity, the estimator $\tau_{IV}$ still converges to a large-sample limit $\tau_{LATE}$, the &lt;strong&gt;local average treatment effect (LATE)&lt;/strong&gt;. 

  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;identifying-assumptions&#34;&gt;Identifying assumptions&lt;/h3&gt;
&lt;p&gt;There are &lt;strong&gt;3 identification assumptions&lt;/strong&gt;. Or let&amp;rsquo;s say there are three main assumptions that must be satisfied for a variable $Z$ to be considered an instrument.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(exogeneity)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
The instrument $Z_i$ must be &lt;strong&gt;exogenous&lt;/strong&gt;, which here means $\epsilon_i \indep Z_i$

  &lt;/div&gt;
&lt;/div&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(relevance)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
The instrument $Z_i$ must be &lt;strong&gt;relevant&lt;/strong&gt;, such that $\Cov[W_i, Z_i] \neq 0$

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Graphically, the relevance assumption corresponds to the existence of an active edge from $Z$ to $W$ in the causal graph.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-4&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 4&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(exclusion restriction)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
The instrument $Z_i$ must satisfy  &lt;strong&gt;exclusion restriction&lt;/strong&gt;, meaning that any eﬀect of $Z_i$ on $Y_i$ must be fully mediated via the treatment $W_i$ 

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Graphically, this means that we’ve excluded enough potential edges between variables in the causal graph so that all causal paths from $Z$ to $Y$ go through $W$.&lt;/p&gt;
&lt;h3 id=&#34;case-i-easiest&#34;&gt;Case I (Easiest)&lt;/h3&gt;
&lt;p&gt;Consider the fully linear version,&lt;/p&gt;
 $$
\begin{align}
&amp; Y=\alpha+W \tau+\varepsilon, \quad \varepsilon \indep Z \tag{C1} \\
&amp; W=Z \gamma+\eta
\end{align}
$$ 
&lt;p&gt;Then,&lt;/p&gt;
 $$\operatorname{Cov}[Y, Z]=\operatorname{Cov}[\tau W+\varepsilon, Z]=\tau \operatorname{Cov}[W, Z] $$

$$\tau= \frac{\operatorname{Cov}[Y, Z]}{\operatorname{Cov}[W, Z]}$$


&lt;p&gt;It implies,&lt;/p&gt;
&lt;p&gt;$$
\hat{\tau}_{IV}= \frac{\widehat{\operatorname{Cov}}[Y_i, Z_i]}{\widehat{\operatorname{Cov}}[W_i, Z_i]}
$$&lt;/p&gt;
&lt;h3 id=&#34;case-ii-more-general-optimal-instruments&#34;&gt;Case II (More general: optimal instruments)&lt;/h3&gt;
&lt;p&gt;Case I assumes (1) linear relationship between $Y$ and $W$ and (2) linear relationship between $W$ and $Z$. This may be too restrictive. How to extend to &lt;mark&gt;more general specification&lt;/mark&gt;? What should we do if&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;we have &lt;strong&gt;multiple instruments&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;or we believe that the instrument may act &lt;strong&gt;non-linearly&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Consider the following,&lt;/p&gt;
&lt;p&gt;$$
Y=\tau W+\varepsilon, \quad \varepsilon \indep Z, \quad Y, W \in \mathbb{R}, \quad Z \in \mathcal{Z}, \tag{C2}
$$&lt;/p&gt;
&lt;p&gt;where $\mathcal{Z}$ can be a high-dimensional space. Define function $w$ that maps $Z_i$ to the real line $$w: \mathcal{Z} \rightarrow \R$$&lt;/p&gt;
&lt;p&gt;Then by the same argument as the Case I (note: we can regard $w(Z)$ as a &amp;ldquo;pseudo $Z$&amp;rdquo;),&lt;/p&gt;
&lt;p&gt;$$
\tau=\frac{\operatorname{Cov}[Y, w(Z)]}{\operatorname{Cov}[W, w(Z)]}
$$&lt;/p&gt;
&lt;p&gt;provided the denominator is non-zero, resulting a feasible estimator&lt;/p&gt;
 $$
\hat{\tau}_{I V}=\frac{\widehat{\operatorname{Cov}}\left[Y_i, w\left(Z_i\right)\right]}{\widehat{\operatorname{Cov}}\left[W_i, w\left(Z_i\right)\right]}=\frac{\frac{1}{n} \sum_{i=1}^n\left(Y_i-\bar{Y}\right)\left(w\left(Z_i\right)-\overline{w(Z)}\right)}{\frac{1}{n} \sum_{i=1}^n\left(W_i-\bar{W}\right)\left(w\left(Z_i\right)-\overline{w(Z)}\right)}
$$ 







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Suppose $\left(X_i, W_i, Y_i, Z_i\right)$ are IID draws from a distribution satisfying (C2), and let $w: \mathcal{Z} \rightarrow \mathbb{R}$ be such that $\operatorname{Cov}[W, w(Z)] \neq 0$. Then, $\hat{\tau}_{I V}$ as given above is consistent for $\tau$, and

$$
\sqrt{n}\left(\hat{\tau}_{I V}-\tau\right) \Rightarrow \mathcal{N}\left(0, V_w\right), \quad V_w=\frac{\operatorname{Var}\left[\varepsilon_i\right] \operatorname{Var}\left[w\left(Z_i\right)\right]}{\operatorname{Cov}\left[W_i, w\left(Z_i\right)\right]^2}
$$


  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;What is the &lt;strong&gt;best&lt;/strong&gt; function $w$, say $w^*(z)$, that minimizes the variance $V_w$? It turns out that the &lt;strong&gt;optimal instrument&lt;/strong&gt; is &lt;mark&gt;the best prediction of $W_i$ from $Z_i$&lt;/mark&gt;.&lt;/p&gt;
&lt;p&gt;$$
w^*(z)=\mathbb{E}\left[W_i \mid Z_i=z\right]
$$&lt;/p&gt;
&lt;p&gt;How to do we estimate? Do &lt;strong&gt;cross-fitting&lt;/strong&gt;! Why? Because&lt;/p&gt;
 $$
\epsilon_i \rightarrow W_i \rightarrow \hat{w}(Z_i) 
$$
&lt;p&gt;We no longer have $\hat{w}(Z_i) \indep \epsilon_i$. Therefore, we need cross-ﬁtting to address this issue.&lt;/p&gt;
&lt;h3 id=&#34;case-iii-much-more-general-non-parametric-iv-regression&#34;&gt;Case III (Much more general: non-parametric IV regression)&lt;/h3&gt;
&lt;p&gt;The more general version than Case II is the following,&lt;/p&gt;
&lt;p&gt;$$
Y_i=\alpha+g\left(W_i\right)+\varepsilon_i, \quad Z_i \indep \varepsilon_i, \quad Y_i, W_i \in \mathbb{R}, \quad Z_i \in \mathcal{Z} \tag{C3}
$$&lt;/p&gt;
&lt;p&gt;where $g(\cdot)$ is some generic smooth function we want to estimate. Note that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;(C3) still requires the eﬀect of $W_i$ on $Y_i$ to be &lt;strong&gt;additive&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;however, unlike (C2), it now allows this additive eﬀect to be modified by a &lt;strong&gt;non-linearity&lt;/strong&gt; $g(\cdot)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Now,&lt;/p&gt;
 $$
\begin{aligned}
\mathbb{E}\left[Y_i \mid Z_i=z\right] &amp; =\mathbb{E}\left[\alpha+g\left(W_i\right)+\varepsilon_i \mid Z_i=z\right] \\
&amp; =\alpha+\mathbb{E}\left[g\left(W_i\right) \mid Z_i=z\right] \\
&amp; =\alpha+\int_{\mathbb{R}} g(w) f(w \mid z) d w,
\end{aligned}
$$ 
&lt;p&gt;There are two steps for learning $g(\cdot)$:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;non-parametric model $\hat{f}(w \mid z)$ using cross-fitting&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;estimate $g(\cdot)$ using empirical minimization&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For more details, see Section 9.2 in  &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Stefan&amp;rsquo;s lecture&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&#34;local-average-treatment-effects&#34;&gt;Local Average Treatment Effects&lt;/h2&gt;
&lt;h3 id=&#34;motivation-1&#34;&gt;Motivation&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;IV without the linearity assumption&lt;/strong&gt;: One may doubt the validity the linearity and constant treatment eﬀect assumption in previous section. What about &lt;strong&gt;non-parametric identification&lt;/strong&gt; using IV?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Encouragement Design and Noncompliance&lt;/strong&gt;: Noncompliance is a common problem in encouragement designs involving human beings as experimental units. In those cases, the experimenters cannot force the units to take the treatment but rather only encourage them to do so. &lt;strong&gt;Heterogeneous effects&lt;/strong&gt; should be allowed.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;setup&#34;&gt;Setup&lt;/h3&gt;
&lt;p&gt;Consider an randomized experiment,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Let  $Z_i \in \{0, 1\}$  be the &lt;strong&gt;treatment assigned&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let  $W_i \in \{0, 1\}$  be the &lt;strong&gt;treatment received&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;When $Z_i \neq W_i$, the noncompliance problem arises&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Potential outcome $\{W_i(1), W_i(0)\}$   s.t. $W_i = W_i(Z_i)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Potential outcome $\{Y_i(w,z)\}_{(w,z) \in \{0,1\}}$   s.t. $Y_i = Y_i(W_i, Z_i)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;identifying-assumptions-1&#34;&gt;Identifying assumptions&lt;/h3&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250530095837904.png&#34; alt=&#34;image-20250530095837904&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;h3 id=&#34;late-theorem&#34;&gt;LATE Theorem&lt;/h3&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250530100412392.png&#34; alt=&#34;image-20250530100412392&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;p&gt;Idea of proof:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Start with $\Cov(Y, Z) = \E[YZ] - \E[Y]\E[Z]$, then apply LIE by conditioning on $Z$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Derive $\Cov(W, Z)$ similarly as above, then get the ratio&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Decompose the ATE on $Y$, $\E[Y(1) - Y(0)]$, into four terms (always-taker, compiler, defier, never-taker) using law of total probability&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;By exclusion restriction and monotonicity assumption, only &amp;ldquo;compiler&amp;rdquo; remains&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&#34;multiple-instruments&#34;&gt;Multiple instruments&lt;/h3&gt;
&lt;p&gt;We may have access to data from &lt;strong&gt;multiple randomized trials&lt;/strong&gt; that can be used to study a treatment effect via a non-compliance analysis.&lt;/p&gt;
&lt;p&gt;Marketing example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;goal: study the effect of subscription to a loyalty program ($W_i$) on long-term customer CLV ($Y_i$)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;randomized trial 1: offering discounts for joining the loyalty program $\left(Z_i=\1(\{\right.$ customer received a discount $\left.\})\right)$  &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;randomized trial 2: showing advertisements $\left(Z_i=\1(\{\right.$ customer was shown an ad for the program $\left.\})\right)$  &lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;Previously, under the &lt;strong&gt;linear treatment effect model&lt;/strong&gt;, multiple instruments could be combined into a &lt;strong&gt;single optimal instrument&lt;/strong&gt;, and the optimal instrument corresponds to the summary of all the instruments that best predicts the treatment. &lt;br&gt;&lt;/br&gt; Without the linear treatment effect model, however, we caution that no such result is available. &lt;mark&gt;Diﬀerent instruments may induce diﬀerence compliance patterns&lt;/mark&gt;, and so the LATEs identiﬁed diﬀerent instruments may not be the same.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In the marketing example, the ATE for customers who respond to a discount may be diﬀerent from the ATE for customers who respond to an advertisement.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Wager, S. (2024). Causal inference: A statistical learning approach. &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/causal_inf_book.pdf&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Hausman-Taylor estimator Notes</title>
      <link>https://chenxing.space/blog/hausman-taylor-estimator-notes/</link>
      <pubDate>Wed, 27 Mar 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/hausman-taylor-estimator-notes/</guid>
      <description>&lt;h2 id=&#34;hausman-taylor-in-r&#34;&gt;Hausman-Taylor in R&lt;/h2&gt;
&lt;p&gt;The following example is from &lt;a href=&#34;https://bookdown.org/ccolonescu/RPoE4/panel-data-models.html#the-random-effects-model&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Example: The fixed effects model, however, does not allow time-invariant variables such as educ or black. Since the problem of the random effects model is endogeneity, one can use instrumental variables methods when time-invariant regressors must be in the model. The &lt;strong&gt;Hausman-Taylor estimator&lt;/strong&gt; uses instrumental variables in a random effects model; it assumes four categories of regressors: time-varying exogenous, time-varying endogenous, time-invariant exogenous, and time-invariant endogenous. The number of time-varying variables must be at least equal to the number of time-invariant ones. In our wage model, suppose exper, tenure and union are time-varying exogenous, south is time-varying endogenous, black is time-invariant exogenous, and educ is time-invariant endogenous. The same &lt;code&gt;plm()&lt;/code&gt; function allows carrying out Hausman-Taylor estimation by setting &lt;code&gt;model=&lt;/code&gt; “ht”.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-R&#34; data-lang=&#34;R&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;wage.HT&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;plm&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;lwage&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;~&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;educ&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;exper&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;I&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;exper^2&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;      &lt;span class=&#34;n&#34;&gt;tenure&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;I&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;tenure^2&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;black&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;south&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;union&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;|&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;      &lt;span class=&#34;n&#34;&gt;exper&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;I&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;exper^2&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;tenure&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;I&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;tenure^2&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;union&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;black&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;      &lt;span class=&#34;n&#34;&gt;data&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;nlspd&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;model&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;s&#34;&gt;&amp;#34;ht&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;kable&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;tidy&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;wage.HT&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;digits&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;m&#34;&gt;5&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;caption&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;     &lt;span class=&#34;s&#34;&gt;&amp;#34;Hausman-Taylor estimates for the wage equation&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Note that, the instruments are specified at the end of the formula after a &lt;mark&gt;&lt;code&gt;|&lt;/code&gt; sign&lt;/mark&gt; (pipe). For more details, check&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;library(plm)
browseVignettes(&amp;#34;plm&amp;#34;)
&lt;/code&gt;&lt;/pre&gt;&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;check the &lt;code&gt;plm&lt;/code&gt; 📦 in R.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;code&gt;plm&lt;/code&gt; 📦 github repo: &lt;a href=&#34;https://github.com/ycroissant/plm?tab=readme-ov-file&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://github.com/ycroissant/plm?tab=readme-ov-file&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;plm&lt;/code&gt; 📦 cran: &lt;a href=&#34;https://cran.r-project.org/web/packages/plm/index.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://cran.r-project.org/web/packages/plm/index.html&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;additional-resources-for-panel-data-analysis&#34;&gt;Additional Resources for Panel Data Analysis&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&#34;https://zhuanlan.zhihu.com/p/356250433&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;计量经济学笔记（三）：面板数据分析&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
  </channel>
</rss>
