<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>estimation theory | Chen Xing</title>
    <link>https://chenxing.space/tag/estimation-theory/</link>
      <atom:link href="https://chenxing.space/tag/estimation-theory/index.xml" rel="self" type="application/rss+xml" />
    <description>estimation theory</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Sun, 27 Apr 2025 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>estimation theory</title>
      <link>https://chenxing.space/tag/estimation-theory/</link>
    </image>
    
    <item>
      <title>Notes for Variational Inference</title>
      <link>https://chenxing.space/blog/notes-for-variational-inference/</link>
      <pubDate>Sun, 27 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-for-variational-inference/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;In modern Bayesian statistics, we often face posterior distributions that are difficult to compute. Let $p(z)$ be prior density and $p(x \mid z)$ be likelihood. The standard approach to compute posterior $p(z \mid x)$ is to use MCMC (like Metropolis-Hastings, Gibbs sampling and HMC). But MCMC has &lt;strong&gt;downsides&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Slow&lt;/strong&gt; for big datasets or complex models.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Difficult to scale&lt;/strong&gt; in the era of massive data.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Variational inference (VI)&lt;/strong&gt; offers a &lt;strong&gt;faster&lt;/strong&gt; alternative. The key difference between MCMC and VI is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCMC sample a Markov chain&lt;/li&gt;
&lt;li&gt;VI solve an optimization problem&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The main idea behind VI is to use optimization. Specifically,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;step 1: we posit a &lt;em&gt;family&lt;/em&gt; of densities $Q$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;step 2: find a member in $Q$ to minimize the KL divergence $$
q^*(\mathbf{z})=\underset{q(\mathbf{z}) \in \mathcal{Q}}{\arg \min } \mathrm{KL}(q(\mathbf{z}) | p(\mathbf{z} \mid \mathbf{x})) .
$$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Key Idea&lt;/strong&gt;: Rather than sampling, VI &lt;em&gt;optimizes&lt;/em&gt; — it finds a best guess distribution by minimizing a divergence.&lt;/p&gt;
&lt;h3 id=&#34;vi-vs-mcmc-when-to-use-which&#34;&gt;VI vs MCMC: When to Use Which?&lt;/h3&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250427083257373.png&#34; alt=&#34;image-20250427083257373&#34; style=&#34;zoom:80%;&#34; /&gt;&lt;/center&gt;
&lt;h3 id=&#34;pros-and-cons-of-vi&#34;&gt;Pros and Cons of VI&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Pros&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Much &lt;strong&gt;faster&lt;/strong&gt; than MCMC.&lt;/li&gt;
&lt;li&gt;Easy to scale with &lt;strong&gt;stochastic optimization&lt;/strong&gt; and &lt;strong&gt;distributed computation&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cons&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;VI &lt;strong&gt;underestimates posterior variance&lt;/strong&gt; (it tends to be &amp;ldquo;overconfident&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;It &lt;strong&gt;does not guarantee&lt;/strong&gt; exact samples from the true posterior.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1 id=&#34;variational-inference&#34;&gt;Variational Inference&lt;/h1&gt;
&lt;p&gt;Recall that in the Bayesian framework, $$p(z \mid x) = \frac{p(z,x)}{p(x)} \propto p(x \mid z) p(z)$$ We try to avoid calculating the denominator (marginal likelihood), $p(x)$, also called &lt;em&gt;evidence&lt;/em&gt;, as it requires us to calculate high dimensional integrals.&lt;/p&gt;
&lt;p&gt;Variational inference turns Bayesian inference into an optimization problem by minimizing KL divergence within a simpler family of distributions, typically using coordinate ascent to maximize the evidence lower bound (ELBO).&lt;/p&gt;
&lt;p&gt;First, the optimization goal is&lt;/p&gt;
&lt;p&gt;$$
q^*(\mathbf{z})=\underset{q(\mathbf{z}) \in \mathcal{Q}}{\arg \min } KL(q(\mathbf{z}) | p(\mathbf{z} \mid \mathbf{x})),
$$&lt;/p&gt;
&lt;p&gt;where $q^*$ is the best approximation.&lt;/p&gt;
$$
\begin{aligned}
KL(q(\boldsymbol{z}) \| p(\boldsymbol{z} \mid \boldsymbol{x})) &amp; =\int_z q(\boldsymbol{z}) \log \left[\frac{q(\boldsymbol{z})}{p(\boldsymbol{z} \mid \boldsymbol{x})}\right] d \boldsymbol{z} \\
&amp; =\int_{\boldsymbol{z}}[q(\boldsymbol{z}) \log q(\boldsymbol{z})] d \boldsymbol{z}-\int_{\boldsymbol{z}}[q(\boldsymbol{z}) \log p(\boldsymbol{z} \mid \boldsymbol{x})] d \boldsymbol{z} \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q[\log p(\boldsymbol{z} \mid \boldsymbol{x})] \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q\left[\log \left[\frac{p(\boldsymbol{x}, \boldsymbol{z})}{p(\boldsymbol{x})}\right]\right] \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]+\mathbb{E}_q[\log p(\boldsymbol{x})] \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]+\log p(\boldsymbol{x})
\end{aligned}
$$
&lt;p&gt;Note that, $\log p(\boldsymbol{x})$ does not contain $q(\cdot)$, so we can ignore it in the optimization. We define &lt;strong&gt;evidence lower bound (ELBO)&lt;/strong&gt; as&lt;/p&gt;
&lt;p&gt;$$
\operatorname{ELBO}(q)=\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]-\mathbb{E}_q[\log q(\boldsymbol{z})]
$$&lt;/p&gt;
&lt;p&gt;This value is called evidence lower bound because it is the lower bound of “log evidence”.&lt;/p&gt;
$$
\begin{aligned}
\log p(\boldsymbol{x}) &amp; =\operatorname{ELBO}(q)+KL(q(\boldsymbol{z}) \| p(\boldsymbol{z} \mid \boldsymbol{x})) \\
&amp; \geq \operatorname{ELBO}(q)
\end{aligned}
$$
&lt;p&gt;The second line holds because $KL(\cdot | \cdot) \ge 0$ by Jensen’s inequality.&lt;/p&gt;
&lt;p&gt;Therefore, &lt;strong&gt;maximizing ELBO is equivalent to minimizing KL divergence&lt;/strong&gt;.&lt;/p&gt;
$$
\begin{aligned}
q^*(\boldsymbol{z}) &amp; =\underset{q(\boldsymbol{z}) \in \mathcal{Q}}{\operatorname{argmin}} KL(q(\boldsymbol{z}) \| p(\boldsymbol{z} \mid \boldsymbol{x})) \\
&amp; =\underset{q(\boldsymbol{z}) \in \mathcal{Q}}{\operatorname{argmax}} \operatorname{ELBO}(\mathrm{q}) \\
&amp; =\underset{q(\boldsymbol{z}) \in \mathcal{Q}}{\operatorname{argmax}}\left\{\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]-\mathbb{E}_q[\log q(\boldsymbol{z})]\right\}
\end{aligned}
$$
&lt;p&gt;What is the &lt;strong&gt;intuition&lt;/strong&gt; for $\operatorname{ELBO}(q)$?&lt;/p&gt;
$$
\begin{aligned}
\operatorname{ELBO}(q) &amp; \triangleq \mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]-\mathbb{E}_q[\log q(\boldsymbol{z})] \\
&amp; =\mathbb{E}[\log p(\mathbf{z})]+\mathbb{E}[\log p(\mathbf{x} \mid \mathbf{z})]-\mathbb{E}[\log q(\mathbf{z})] \\
&amp; =\mathbb{E}[\log p(\mathbf{x} \mid \mathbf{z})]-\mathrm{KL}(q(\mathbf{z}) \| p(\mathbf{z})) .
\end{aligned}
$$
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;first term is try to &amp;ldquo;maximize the likelihood&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;second term is try to encourage density $q(\cdot)$ close to prior&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;balance between likelihood and prior&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;mean-field-variational-family&#34;&gt;Mean-Field Variational Family&lt;/h2&gt;
&lt;p&gt;In mean field variational inference, we assume that the variational family &lt;strong&gt;factorizes&lt;/strong&gt;,&lt;/p&gt;
&lt;p&gt;$$
q(z_1, \cdots, z_m) = \prod_{j=1}^m q_j(z_j),
$$
Each variable is &lt;strong&gt;independent&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;coordinate-ascent-algorithm&#34;&gt;Coordinate ascent algorithm&lt;/h2&gt;
&lt;p&gt;We will use &lt;strong&gt;coordinate ascent inference&lt;/strong&gt;, iteratively optimizing each variational distribution holding the others fixed.&lt;/p&gt;
&lt;p&gt;The ELBO converges to a local minimum. Use the resulting $q$ is as a proxy for the true posterior.&lt;/p&gt;
&lt;p&gt;There is a strong relationship between this algorithm and Gibbs sampling.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;In Gibbs sampling we sample from the conditional&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In coordinate ascent variational inference, we iteratively set each factor to&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;$$
\text{distribution of } z_k \propto \exp {\mathbb{E}[\log(\text{conditional})]}
$$&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Blei, D. M., Kucukelbir ,Alp, &amp;amp; and McAuliffe, J. D. (2017). Variational Inference: A Review for Statisticians. &lt;i&gt;Journal of the American Statistical Association&lt;/i&gt;, &lt;i&gt;112&lt;/i&gt;(518), 859–877. &lt;a href=&#34;https://doi.org/10.1080/01621459.2017.1285773&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1080/01621459.2017.1285773&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Looking for a nice summary? Check this  first: &lt;a href=&#34;https://www.cs.princeton.edu/courses/archive/fall11/cos597C/lectures/variational-inference-i.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Variational Inference - Princeton CS tutorial&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;For derivation details, check this: &lt;a href=&#34;https://leimao.github.io/article/Introduction-to-Variational-Inference/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Introduction to Variational Inference&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>A Story Behind Maximum Likelihood</title>
      <link>https://chenxing.space/blog/a-story-behind-mle/</link>
      <pubDate>Wed, 31 Aug 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/a-story-behind-mle/</guid>
      <description>&lt;h2 id=&#34;tldr&#34;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;What are we actually doing on &lt;strong&gt;MLE&lt;/strong&gt;? What is the motivation for the Maximum Likelihood Estimation? In this post, we will go over the story behind the maximum likelihood.&lt;/p&gt;
&lt;p&gt;We&amp;rsquo;ll first talk about the &lt;strong&gt;total variation distance&lt;/strong&gt;, which is a very intuitive measure to tell you how &amp;ldquo;close&amp;rdquo; our estimator is from the true parameter.&lt;/p&gt;
&lt;p&gt;Next, let&amp;rsquo;s move to &lt;strong&gt;KL-divergence&lt;/strong&gt; used to replace the &lt;strong&gt;total variation distance&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In the end, we&amp;rsquo;ll &lt;strong&gt;minimize the KL-divergence&lt;/strong&gt; to get the &amp;ldquo;good&amp;rdquo; estimator. In this process, the &lt;strong&gt;maximum likelihood principle&lt;/strong&gt; will come in.&lt;/p&gt;
&lt;h2 id=&#34;total-variation-distance&#34;&gt;Total Variation Distance&lt;/h2&gt;
&lt;h4 id=&#34;how-to-define-a-good-estimator&#34;&gt;How to define a &amp;ldquo;good&amp;rdquo; estimator?&lt;/h4&gt;
&lt;p&gt;A &amp;ldquo;good&amp;rdquo; estimator should be very &amp;ldquo;close&amp;rdquo; to the true parameter isn&amp;rsquo;t it?&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-motivation&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831072120650.png&#34; alt=&#34;image-20220831072120650&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Motivation
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;details class=&#34;spoiler &#34;  id=&#34;spoiler-0&#34;&gt;
  &lt;summary&gt;Click to view the formal setting&lt;/summary&gt;
  &lt;p&gt;&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831072454087.png&#34; alt=&#34;image-20220831072454087&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;&lt;/p&gt;
&lt;/details&gt;
&lt;h5 id=&#34;definition-of-total-variation-distance&#34;&gt;Definition of Total Variation Distance&lt;/h5&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-def-tv-distance&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831072725776.png&#34; alt=&#34;image-20220831072725776&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Def (TV distance)
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;This definition is very intuitive but &lt;em&gt;pretty strong&lt;/em&gt; and &lt;strong&gt;hard to calculate&lt;/strong&gt; right? Because you have to find the maximum over all possible sets. The good news is, we also have this formulation:&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831082216550.png&#34; alt=&#34;image-20220831082216550&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;details class=&#34;spoiler &#34;  id=&#34;spoiler-1&#34;&gt;
  &lt;summary&gt;Click to view the formal statement.&lt;/summary&gt;
  &lt;p&gt;&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831074119183.png&#34; alt=&#34;image-20220831074119183&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;&lt;/p&gt;
&lt;/details&gt;
&lt;p&gt;Before the proof, let&amp;rsquo;s use the graph to get some feeling for the equation. &lt;strong&gt;Where does the 1/2 come from?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-area-between-the-curves&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831074556512.png&#34; alt=&#34;image-20220831074556512&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Area between the curves
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;&lt;u&gt;Proof&lt;/u&gt;:&lt;/p&gt;
&lt;details class=&#34;spoiler &#34;  id=&#34;spoiler-2&#34;&gt;
  &lt;summary&gt;Click to view the proof&lt;/summary&gt;
  &lt;p&gt;&lt;figure  id=&#34;figure-proof-of-tv-equation&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831075931046.png&#34; alt=&#34;image-20220831075931046&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Proof of TV equation
    &lt;/figcaption&gt;&lt;/figure&gt;&lt;/p&gt;
&lt;/details&gt;
&lt;p&gt;Note that, the key part in the proof is the observation: $$\int f = \int g = 1 \implies \int f - g = 0,$$ Then we have $$\int_{\{x: f-g &gt; 0\}} f - g = \int_{\{x: f-g &lt; 0\}} g - f.$$&lt;/p&gt;
&lt;h5 id=&#34;properties-of-tv&#34;&gt;Properties of TV&lt;/h5&gt;
&lt;p&gt;Total variation is &lt;strong&gt;symmetric&lt;/strong&gt;, &lt;strong&gt;non-negative&lt;/strong&gt;, &lt;strong&gt;definite&lt;/strong&gt; and satisfies the triangle inequality.&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-properties-of-total-variation&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831082810899.png&#34; alt=&#34;image-20220831082810899&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Properties of Total Variation
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id=&#34;unclear-how-to-estimate-tv&#34;&gt;Unclear how to estimate TV!&lt;/h2&gt;
&lt;p&gt;Our goal is to find the &amp;ldquo;good&amp;rdquo; estimator.&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220901074618037.png&#34; alt=&#34;image-20220901074618037&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;If using Total Variation distance to describe the &amp;ldquo;close&amp;rdquo;, we need to :&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Build an estimator $\widehat{TV}(\mathbb{P}_{\theta}, \mathbb{P}_{\theta^*})$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Find $\hat{\theta}$ that &lt;em&gt;minimize&lt;/em&gt; the function.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;However, it is &lt;strong&gt;unclear how to build&lt;/strong&gt; $\widehat{TV}(\mathbb{P}_{\theta}, \mathbb{P}_{\theta^*})$!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;We don&amp;rsquo;t know the $f_{\theta^*}$  (the true parameter $\theta^*$  is unknown), and it is very hard to manipulate the integral of density difference.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The common strategy is to replace the &lt;strong&gt;expectation&lt;/strong&gt; ($E(\cdot)$ ) with the &lt;strong&gt;average&lt;/strong&gt; ($\frac{1}{n}\sum_n(\cdot)$ ), but there is no clear expectation in $TV(\cdot)$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Due to above difficulties, we need a more convenient distance between probability measure to &lt;strong&gt;replace&lt;/strong&gt; total variation. This is the &lt;strong&gt;motivation for DL divergence&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;u&gt;REMARK&lt;/u&gt;:&lt;/p&gt;
&lt;p&gt;The total variation distance, describing the &amp;ldquo;worst&amp;rdquo; scenario, is very intuitive and has a clear interpretation. However, it is hard to build an estimator. That&amp;rsquo;s why we move to KL divergence.&lt;/p&gt;
&lt;h2 id=&#34;kl-divergence&#34;&gt;KL divergence&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s check the definition first,&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-kl-divergence&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831135444913.png&#34; alt=&#34;image-20220831135444913&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      KL divergence
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-properties-of-kl-divergence&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831140735027.png&#34; alt=&#34;image-20220831140735027&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Properties of KL-divergence
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;For the second property (non-negative), we can use Jensen&amp;rsquo;s inequality to show it.&lt;/p&gt;
&lt;h2 id=&#34;kl--mle&#34;&gt;KL 🤝 MLE&lt;/h2&gt;
&lt;p&gt;Now we are ready to introduce &lt;strong&gt;maximum likelihood principle&lt;/strong&gt; using the KL-divergence.&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-maximum-likelihood-connection-with-kl-divergence&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831143953327.png&#34; alt=&#34;image-20220831143953327&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Maximum Likelihood Connection with KL-divergence
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id=&#34;summary&#34;&gt;Summary&lt;/h2&gt;
&lt;p&gt;What are we actually doing on &lt;strong&gt;MLE&lt;/strong&gt;?&lt;/p&gt;
&lt;p&gt;The short answer is &lt;strong&gt;we are minimizing the KL-divergence&lt;/strong&gt;.&lt;/p&gt;
&lt;div class=&#34;alert alert-note&#34;&gt;
  &lt;div&gt;
    Remember: &lt;strong&gt;Maximum Likelihood Estimation&lt;/strong&gt; is just the empirical version of trying to &lt;strong&gt;minimize the KL-divergence!&lt;/strong&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;YouTube video: &lt;a href=&#34;https://www.youtube.com/watch?v=rLlZpnT02ZU&amp;amp;list=PLUl4u3cNGP60uVBMaoNERc6knT_MgPKS0&amp;amp;index=4&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;MIT Parametric Inference (cont.) and Maximum Likelihood Estimation&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
  </channel>
</rss>
