<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Probability &amp; Bayesian Statistics | Chen Xing</title>
    <link>https://chenxing.space/category/probability-bayesian-statistics/</link>
      <atom:link href="https://chenxing.space/category/probability-bayesian-statistics/index.xml" rel="self" type="application/rss+xml" />
    <description>Probability &amp; Bayesian Statistics</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Sun, 27 Apr 2025 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>Probability &amp; Bayesian Statistics</title>
      <link>https://chenxing.space/category/probability-bayesian-statistics/</link>
    </image>
    
    <item>
      <title>Notes for Variational Inference</title>
      <link>https://chenxing.space/blog/notes-for-variational-inference/</link>
      <pubDate>Sun, 27 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-for-variational-inference/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;In modern Bayesian statistics, we often face posterior distributions that are difficult to compute. Let $p(z)$ be prior density and $p(x \mid z)$ be likelihood. The standard approach to compute posterior $p(z \mid x)$ is to use MCMC (like Metropolis-Hastings, Gibbs sampling and HMC). But MCMC has &lt;strong&gt;downsides&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Slow&lt;/strong&gt; for big datasets or complex models.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Difficult to scale&lt;/strong&gt; in the era of massive data.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Variational inference (VI)&lt;/strong&gt; offers a &lt;strong&gt;faster&lt;/strong&gt; alternative. The key difference between MCMC and VI is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCMC sample a Markov chain&lt;/li&gt;
&lt;li&gt;VI solve an optimization problem&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The main idea behind VI is to use optimization. Specifically,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;step 1: we posit a &lt;em&gt;family&lt;/em&gt; of densities $Q$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;step 2: find a member in $Q$ to minimize the KL divergence $$
q^*(\mathbf{z})=\underset{q(\mathbf{z}) \in \mathcal{Q}}{\arg \min } \mathrm{KL}(q(\mathbf{z}) | p(\mathbf{z} \mid \mathbf{x})) .
$$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Key Idea&lt;/strong&gt;: Rather than sampling, VI &lt;em&gt;optimizes&lt;/em&gt; — it finds a best guess distribution by minimizing a divergence.&lt;/p&gt;
&lt;h3 id=&#34;vi-vs-mcmc-when-to-use-which&#34;&gt;VI vs MCMC: When to Use Which?&lt;/h3&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250427083257373.png&#34; alt=&#34;image-20250427083257373&#34; style=&#34;zoom:80%;&#34; /&gt;&lt;/center&gt;
&lt;h3 id=&#34;pros-and-cons-of-vi&#34;&gt;Pros and Cons of VI&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Pros&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Much &lt;strong&gt;faster&lt;/strong&gt; than MCMC.&lt;/li&gt;
&lt;li&gt;Easy to scale with &lt;strong&gt;stochastic optimization&lt;/strong&gt; and &lt;strong&gt;distributed computation&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cons&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;VI &lt;strong&gt;underestimates posterior variance&lt;/strong&gt; (it tends to be &amp;ldquo;overconfident&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;It &lt;strong&gt;does not guarantee&lt;/strong&gt; exact samples from the true posterior.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1 id=&#34;variational-inference&#34;&gt;Variational Inference&lt;/h1&gt;
&lt;p&gt;Recall that in the Bayesian framework, $$p(z \mid x) = \frac{p(z,x)}{p(x)} \propto p(x \mid z) p(z)$$ We try to avoid calculating the denominator (marginal likelihood), $p(x)$, also called &lt;em&gt;evidence&lt;/em&gt;, as it requires us to calculate high dimensional integrals.&lt;/p&gt;
&lt;p&gt;Variational inference turns Bayesian inference into an optimization problem by minimizing KL divergence within a simpler family of distributions, typically using coordinate ascent to maximize the evidence lower bound (ELBO).&lt;/p&gt;
&lt;p&gt;First, the optimization goal is&lt;/p&gt;
&lt;p&gt;$$
q^*(\mathbf{z})=\underset{q(\mathbf{z}) \in \mathcal{Q}}{\arg \min } KL(q(\mathbf{z}) | p(\mathbf{z} \mid \mathbf{x})),
$$&lt;/p&gt;
&lt;p&gt;where $q^*$ is the best approximation.&lt;/p&gt;
$$
\begin{aligned}
KL(q(\boldsymbol{z}) \| p(\boldsymbol{z} \mid \boldsymbol{x})) &amp; =\int_z q(\boldsymbol{z}) \log \left[\frac{q(\boldsymbol{z})}{p(\boldsymbol{z} \mid \boldsymbol{x})}\right] d \boldsymbol{z} \\
&amp; =\int_{\boldsymbol{z}}[q(\boldsymbol{z}) \log q(\boldsymbol{z})] d \boldsymbol{z}-\int_{\boldsymbol{z}}[q(\boldsymbol{z}) \log p(\boldsymbol{z} \mid \boldsymbol{x})] d \boldsymbol{z} \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q[\log p(\boldsymbol{z} \mid \boldsymbol{x})] \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q\left[\log \left[\frac{p(\boldsymbol{x}, \boldsymbol{z})}{p(\boldsymbol{x})}\right]\right] \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]+\mathbb{E}_q[\log p(\boldsymbol{x})] \\
&amp; =\mathbb{E}_q[\log q(\boldsymbol{z})]-\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]+\log p(\boldsymbol{x})
\end{aligned}
$$
&lt;p&gt;Note that, $\log p(\boldsymbol{x})$ does not contain $q(\cdot)$, so we can ignore it in the optimization. We define &lt;strong&gt;evidence lower bound (ELBO)&lt;/strong&gt; as&lt;/p&gt;
&lt;p&gt;$$
\operatorname{ELBO}(q)=\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]-\mathbb{E}_q[\log q(\boldsymbol{z})]
$$&lt;/p&gt;
&lt;p&gt;This value is called evidence lower bound because it is the lower bound of “log evidence”.&lt;/p&gt;
$$
\begin{aligned}
\log p(\boldsymbol{x}) &amp; =\operatorname{ELBO}(q)+KL(q(\boldsymbol{z}) \| p(\boldsymbol{z} \mid \boldsymbol{x})) \\
&amp; \geq \operatorname{ELBO}(q)
\end{aligned}
$$
&lt;p&gt;The second line holds because $KL(\cdot | \cdot) \ge 0$ by Jensen’s inequality.&lt;/p&gt;
&lt;p&gt;Therefore, &lt;strong&gt;maximizing ELBO is equivalent to minimizing KL divergence&lt;/strong&gt;.&lt;/p&gt;
$$
\begin{aligned}
q^*(\boldsymbol{z}) &amp; =\underset{q(\boldsymbol{z}) \in \mathcal{Q}}{\operatorname{argmin}} KL(q(\boldsymbol{z}) \| p(\boldsymbol{z} \mid \boldsymbol{x})) \\
&amp; =\underset{q(\boldsymbol{z}) \in \mathcal{Q}}{\operatorname{argmax}} \operatorname{ELBO}(\mathrm{q}) \\
&amp; =\underset{q(\boldsymbol{z}) \in \mathcal{Q}}{\operatorname{argmax}}\left\{\mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]-\mathbb{E}_q[\log q(\boldsymbol{z})]\right\}
\end{aligned}
$$
&lt;p&gt;What is the &lt;strong&gt;intuition&lt;/strong&gt; for $\operatorname{ELBO}(q)$?&lt;/p&gt;
$$
\begin{aligned}
\operatorname{ELBO}(q) &amp; \triangleq \mathbb{E}_q[\log p(\boldsymbol{x}, \boldsymbol{z})]-\mathbb{E}_q[\log q(\boldsymbol{z})] \\
&amp; =\mathbb{E}[\log p(\mathbf{z})]+\mathbb{E}[\log p(\mathbf{x} \mid \mathbf{z})]-\mathbb{E}[\log q(\mathbf{z})] \\
&amp; =\mathbb{E}[\log p(\mathbf{x} \mid \mathbf{z})]-\mathrm{KL}(q(\mathbf{z}) \| p(\mathbf{z})) .
\end{aligned}
$$
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;first term is try to &amp;ldquo;maximize the likelihood&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;second term is try to encourage density $q(\cdot)$ close to prior&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;balance between likelihood and prior&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;mean-field-variational-family&#34;&gt;Mean-Field Variational Family&lt;/h2&gt;
&lt;p&gt;In mean field variational inference, we assume that the variational family &lt;strong&gt;factorizes&lt;/strong&gt;,&lt;/p&gt;
&lt;p&gt;$$
q(z_1, \cdots, z_m) = \prod_{j=1}^m q_j(z_j),
$$
Each variable is &lt;strong&gt;independent&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;coordinate-ascent-algorithm&#34;&gt;Coordinate ascent algorithm&lt;/h2&gt;
&lt;p&gt;We will use &lt;strong&gt;coordinate ascent inference&lt;/strong&gt;, iteratively optimizing each variational distribution holding the others fixed.&lt;/p&gt;
&lt;p&gt;The ELBO converges to a local minimum. Use the resulting $q$ is as a proxy for the true posterior.&lt;/p&gt;
&lt;p&gt;There is a strong relationship between this algorithm and Gibbs sampling.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;In Gibbs sampling we sample from the conditional&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In coordinate ascent variational inference, we iteratively set each factor to&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;$$
\text{distribution of } z_k \propto \exp {\mathbb{E}[\log(\text{conditional})]}
$$&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Blei, D. M., Kucukelbir ,Alp, &amp;amp; and McAuliffe, J. D. (2017). Variational Inference: A Review for Statisticians. &lt;i&gt;Journal of the American Statistical Association&lt;/i&gt;, &lt;i&gt;112&lt;/i&gt;(518), 859–877. &lt;a href=&#34;https://doi.org/10.1080/01621459.2017.1285773&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1080/01621459.2017.1285773&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Looking for a nice summary? Check this  first: &lt;a href=&#34;https://www.cs.princeton.edu/courses/archive/fall11/cos597C/lectures/variational-inference-i.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Variational Inference - Princeton CS tutorial&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;For derivation details, check this: &lt;a href=&#34;https://leimao.github.io/article/Introduction-to-Variational-Inference/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Introduction to Variational Inference&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Using Stan to do MCMC</title>
      <link>https://chenxing.space/blog/using-stan-to-do-mcmc/</link>
      <pubDate>Thu, 20 Apr 2023 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/using-stan-to-do-mcmc/</guid>
      <description>&lt;h2 id=&#34;what-is-stan&#34;&gt;What is Stan?&lt;/h2&gt;
&lt;p&gt;Stan is an intuitive, yet sophisticated, probabilistic programming language that provides an interface to a recently proposed extension to Hamiltonian Monte Carlo (&lt;strong&gt;HMC&lt;/strong&gt;), known as No U-Turn Sampler (&lt;strong&gt;NUTS&lt;/strong&gt;).&lt;/p&gt;
&lt;p&gt;Stan is what is known as an &lt;strong&gt;imperative&lt;/strong&gt; programming language. This is not true for BUGS and JAGS, which are known as &lt;strong&gt;declarative&lt;/strong&gt; languages.&lt;/p&gt;
&lt;h2 id=&#34;why-choose-stan&#34;&gt;Why choose Stan?&lt;/h2&gt;
&lt;p&gt;Stan is popular and has huge range of resources generated by users and developers, so when you get stuck, you can find someone to get help.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Bayesian Statistics Self-Study Resources</title>
      <link>https://chenxing.space/blog/bayesian-statistics-self-study-resources/</link>
      <pubDate>Sat, 11 Mar 2023 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/bayesian-statistics-self-study-resources/</guid>
      <description>&lt;p&gt;Here are some good resources for Bayesian Statistics self-learning.&lt;/p&gt;
&lt;h2 id=&#34;textbooks&#34;&gt;Textbooks&lt;/h2&gt;
&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303211423906.png&#34; alt=&#34;img2023-03-21 14.22.43&#34; style=&#34;zoom:33%;&#34; /&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;A Student&amp;rsquo;s Guide to Bayesian Statistics&amp;rdquo;&lt;/strong&gt;: Easy to understand, also with problem and solution available on &lt;a href=&#34;https://ben-lambert.com/a-students-guide-to-bayesian-statistics/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;the author&amp;rsquo;s personal blog&lt;/a&gt;. There are also lots of other useful Bayesian resources on his blog, for example, &lt;a href=&#34;https://ben-lambert.com/bayesian-lecture-slides/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;the Bayesian lecture slides&lt;/a&gt;, definitely deserve to check!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Bayesian Data Analysis&amp;rdquo;&lt;/strong&gt;: Sort of &amp;ldquo;Bayesian Bible&amp;rdquo;. Comprehensive coverage of Bayesian methods.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Statistical Rethinking A Bayesian Course with Examples in R and STAN&amp;rdquo;&lt;/strong&gt;: Very Special flavor &amp;hellip; includes a strong focus on using the statistical software R and the modeling language Stan.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.bayesrulesbook.com&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Bayes Rules! An Introduction to Applied Bayesian Modeling&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;video-lectures&#34;&gt;Video Lectures&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;A Student&amp;rsquo;s Guide to Bayesian Statistics&lt;/strong&gt; textbook also provides a series of &lt;a href=&#34;https://www.youtube.com/watch?v=P_og8H-VkIY&amp;amp;list=PLwJRxp3blEvZ8AKMXOy0fc0cqT61GsKCG&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;video lectures&lt;/a&gt; on YouTube.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Statistical Rethinking&lt;/strong&gt; &lt;a href=&#34;https://github.com/rmcelreath/stat_rethinking_2023&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;with video lectures&lt;/a&gt;. &amp;ldquo;Statistical Rethinking: A Bayesian Course with Examples in R and Stan&amp;rdquo; by Richard McElreath. This book is very practical and focuses on applied Bayesian statistics, and would be a good choice if you want to get your hands dirty with real-world problems and data. It uses the R and Stan software tools, which are popular among statisticians and data scientists.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/playlist?list=PLFDbGp5YzjqXQ4oE4w9GVWdiokWB9gEpm&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Bayesian statistics: a comprehensive course&lt;/a&gt;: This playlist provides a complete introduction to the field of Bayesian statistics and covers a wide range of topics, including Bayesian inference, prior and posterior distributions, conjugate priors, hierarchical models, and more.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;MIT Open Course Ware: &amp;ldquo;Statistics for Applications&amp;rdquo; &lt;a href=&#34;https://www.youtube.com/watch?v=bFZ-0FH5hfs&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Lecture 17&lt;/a&gt; and lecture 18 give a brief introduction to Bayesian statistics.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/playlist?list=PLvcbYUQ5t0UEkf2NUEo7XSsyVTyeEk3Gq&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Bayesian Statistics&lt;/a&gt; YouTube playlist provides us some short videos, around 10 mins per video, to illustrate some key concepts in Bayesian stats.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/playlist?list=PL_lWxa4iVNt2GBPOVZMVKD4jYl9Q7hs2K&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;S22 MATH 347 Bayesian Statistics&lt;/a&gt; at Vassar College video lectures with &lt;a href=&#34;https://github.com/monika76five/Undergrad-Bayesian-Course&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;GitHub repo&lt;/a&gt; containing the course material.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;miscellaneous&#34;&gt;Miscellaneous&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;http://www2.stat.duke.edu/~fl35/BayesianCausalInference.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;STA 790 (Special Topics): Bayesian Causal Inference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://bayesf22.classes.andrewheiss.com&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Bayesian Statistics -  Independent readings course on Bayesian statistics with R and Stan - Andrew Heiss and Meng Ye - Fall 2022&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=fNk_zzaMoSs&amp;amp;list=PLZHQObOWTQDPD3MizzM2xVFitgF8hE_ab&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;3Blue1Brown&amp;rsquo;s &amp;ldquo;Essence of Linear Algebra&amp;rdquo; series video lectures&lt;/a&gt; to review linear algebra.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>A Story Behind Maximum Likelihood</title>
      <link>https://chenxing.space/blog/a-story-behind-mle/</link>
      <pubDate>Wed, 31 Aug 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/a-story-behind-mle/</guid>
      <description>&lt;h2 id=&#34;tldr&#34;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;What are we actually doing on &lt;strong&gt;MLE&lt;/strong&gt;? What is the motivation for the Maximum Likelihood Estimation? In this post, we will go over the story behind the maximum likelihood.&lt;/p&gt;
&lt;p&gt;We&amp;rsquo;ll first talk about the &lt;strong&gt;total variation distance&lt;/strong&gt;, which is a very intuitive measure to tell you how &amp;ldquo;close&amp;rdquo; our estimator is from the true parameter.&lt;/p&gt;
&lt;p&gt;Next, let&amp;rsquo;s move to &lt;strong&gt;KL-divergence&lt;/strong&gt; used to replace the &lt;strong&gt;total variation distance&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In the end, we&amp;rsquo;ll &lt;strong&gt;minimize the KL-divergence&lt;/strong&gt; to get the &amp;ldquo;good&amp;rdquo; estimator. In this process, the &lt;strong&gt;maximum likelihood principle&lt;/strong&gt; will come in.&lt;/p&gt;
&lt;h2 id=&#34;total-variation-distance&#34;&gt;Total Variation Distance&lt;/h2&gt;
&lt;h4 id=&#34;how-to-define-a-good-estimator&#34;&gt;How to define a &amp;ldquo;good&amp;rdquo; estimator?&lt;/h4&gt;
&lt;p&gt;A &amp;ldquo;good&amp;rdquo; estimator should be very &amp;ldquo;close&amp;rdquo; to the true parameter isn&amp;rsquo;t it?&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-motivation&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831072120650.png&#34; alt=&#34;image-20220831072120650&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Motivation
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;details class=&#34;spoiler &#34;  id=&#34;spoiler-0&#34;&gt;
  &lt;summary&gt;Click to view the formal setting&lt;/summary&gt;
  &lt;p&gt;&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831072454087.png&#34; alt=&#34;image-20220831072454087&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;&lt;/p&gt;
&lt;/details&gt;
&lt;h5 id=&#34;definition-of-total-variation-distance&#34;&gt;Definition of Total Variation Distance&lt;/h5&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-def-tv-distance&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831072725776.png&#34; alt=&#34;image-20220831072725776&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Def (TV distance)
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;This definition is very intuitive but &lt;em&gt;pretty strong&lt;/em&gt; and &lt;strong&gt;hard to calculate&lt;/strong&gt; right? Because you have to find the maximum over all possible sets. The good news is, we also have this formulation:&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831082216550.png&#34; alt=&#34;image-20220831082216550&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;details class=&#34;spoiler &#34;  id=&#34;spoiler-1&#34;&gt;
  &lt;summary&gt;Click to view the formal statement.&lt;/summary&gt;
  &lt;p&gt;&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831074119183.png&#34; alt=&#34;image-20220831074119183&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;&lt;/p&gt;
&lt;/details&gt;
&lt;p&gt;Before the proof, let&amp;rsquo;s use the graph to get some feeling for the equation. &lt;strong&gt;Where does the 1/2 come from?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-area-between-the-curves&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831074556512.png&#34; alt=&#34;image-20220831074556512&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Area between the curves
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;&lt;u&gt;Proof&lt;/u&gt;:&lt;/p&gt;
&lt;details class=&#34;spoiler &#34;  id=&#34;spoiler-2&#34;&gt;
  &lt;summary&gt;Click to view the proof&lt;/summary&gt;
  &lt;p&gt;&lt;figure  id=&#34;figure-proof-of-tv-equation&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831075931046.png&#34; alt=&#34;image-20220831075931046&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Proof of TV equation
    &lt;/figcaption&gt;&lt;/figure&gt;&lt;/p&gt;
&lt;/details&gt;
&lt;p&gt;Note that, the key part in the proof is the observation: $$\int f = \int g = 1 \implies \int f - g = 0,$$ Then we have $$\int_{\{x: f-g &gt; 0\}} f - g = \int_{\{x: f-g &lt; 0\}} g - f.$$&lt;/p&gt;
&lt;h5 id=&#34;properties-of-tv&#34;&gt;Properties of TV&lt;/h5&gt;
&lt;p&gt;Total variation is &lt;strong&gt;symmetric&lt;/strong&gt;, &lt;strong&gt;non-negative&lt;/strong&gt;, &lt;strong&gt;definite&lt;/strong&gt; and satisfies the triangle inequality.&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-properties-of-total-variation&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831082810899.png&#34; alt=&#34;image-20220831082810899&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Properties of Total Variation
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id=&#34;unclear-how-to-estimate-tv&#34;&gt;Unclear how to estimate TV!&lt;/h2&gt;
&lt;p&gt;Our goal is to find the &amp;ldquo;good&amp;rdquo; estimator.&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220901074618037.png&#34; alt=&#34;image-20220901074618037&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;If using Total Variation distance to describe the &amp;ldquo;close&amp;rdquo;, we need to :&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Build an estimator $\widehat{TV}(\mathbb{P}_{\theta}, \mathbb{P}_{\theta^*})$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Find $\hat{\theta}$ that &lt;em&gt;minimize&lt;/em&gt; the function.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;However, it is &lt;strong&gt;unclear how to build&lt;/strong&gt; $\widehat{TV}(\mathbb{P}_{\theta}, \mathbb{P}_{\theta^*})$!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;We don&amp;rsquo;t know the $f_{\theta^*}$  (the true parameter $\theta^*$  is unknown), and it is very hard to manipulate the integral of density difference.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The common strategy is to replace the &lt;strong&gt;expectation&lt;/strong&gt; ($E(\cdot)$ ) with the &lt;strong&gt;average&lt;/strong&gt; ($\frac{1}{n}\sum_n(\cdot)$ ), but there is no clear expectation in $TV(\cdot)$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Due to above difficulties, we need a more convenient distance between probability measure to &lt;strong&gt;replace&lt;/strong&gt; total variation. This is the &lt;strong&gt;motivation for DL divergence&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;u&gt;REMARK&lt;/u&gt;:&lt;/p&gt;
&lt;p&gt;The total variation distance, describing the &amp;ldquo;worst&amp;rdquo; scenario, is very intuitive and has a clear interpretation. However, it is hard to build an estimator. That&amp;rsquo;s why we move to KL divergence.&lt;/p&gt;
&lt;h2 id=&#34;kl-divergence&#34;&gt;KL divergence&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s check the definition first,&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-kl-divergence&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831135444913.png&#34; alt=&#34;image-20220831135444913&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      KL divergence
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-properties-of-kl-divergence&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831140735027.png&#34; alt=&#34;image-20220831140735027&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Properties of KL-divergence
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;For the second property (non-negative), we can use Jensen&amp;rsquo;s inequality to show it.&lt;/p&gt;
&lt;h2 id=&#34;kl--mle&#34;&gt;KL 🤝 MLE&lt;/h2&gt;
&lt;p&gt;Now we are ready to introduce &lt;strong&gt;maximum likelihood principle&lt;/strong&gt; using the KL-divergence.&lt;/p&gt;
&lt;p&gt;















&lt;figure  id=&#34;figure-maximum-likelihood-connection-with-kl-divergence&#34;&gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/image-20220831143953327.png&#34; alt=&#34;image-20220831143953327&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;figcaption&gt;
      Maximum Likelihood Connection with KL-divergence
    &lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id=&#34;summary&#34;&gt;Summary&lt;/h2&gt;
&lt;p&gt;What are we actually doing on &lt;strong&gt;MLE&lt;/strong&gt;?&lt;/p&gt;
&lt;p&gt;The short answer is &lt;strong&gt;we are minimizing the KL-divergence&lt;/strong&gt;.&lt;/p&gt;
&lt;div class=&#34;alert alert-note&#34;&gt;
  &lt;div&gt;
    Remember: &lt;strong&gt;Maximum Likelihood Estimation&lt;/strong&gt; is just the empirical version of trying to &lt;strong&gt;minimize the KL-divergence!&lt;/strong&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;YouTube video: &lt;a href=&#34;https://www.youtube.com/watch?v=rLlZpnT02ZU&amp;amp;list=PLUl4u3cNGP60uVBMaoNERc6knT_MgPKS0&amp;amp;index=4&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;MIT Parametric Inference (cont.) and Maximum Likelihood Estimation&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Beta Distribution — Intuition, Derivation, and Examples</title>
      <link>https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/</link>
      <pubDate>Wed, 03 Aug 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/</guid>
      <description>&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;h4 id=&#34;model-probabilities&#34;&gt;Model probabilities&lt;/h4&gt;
&lt;p&gt;The Beta distribution is &lt;strong&gt;a probability distribution on probabilities&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The Beta distribution can be understood as representing a distribution &lt;em&gt;of probabilities&lt;/em&gt;, that is, it represents all the possible values of a probability when we don&amp;rsquo;t know what that probability is. For example,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;the Click-Through Rate of your advertisement&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the conversion rate of customers actually purchasing in your store&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;how likely the customer will become &amp;ldquo;inactive&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Because the Beta distribution models a probability, its domain is bounded between &lt;strong&gt;0&lt;/strong&gt; and &lt;strong&gt;1&lt;/strong&gt;.&lt;/p&gt;
&lt;h4 id=&#34;generalization-of-uniform-distribution&#34;&gt;Generalization of Uniform Distribution&lt;/h4&gt;
&lt;p&gt;Give me a &lt;strong&gt;continuous&lt;/strong&gt; and &lt;strong&gt;bounded&lt;/strong&gt; random variable except the &lt;em&gt;Uniform Distribution&lt;/em&gt;. This is another way to look at &lt;em&gt;beta distribution&lt;/em&gt;, continuous and bounded between 0 and 1; also the density is not flat.&lt;/p&gt;

$$
X \sim Beta(a, b), \text{ where } a&gt;0, \ b&gt;0.
$$



$$
f_X(x) = c \cdot x ^{a-1}(1-x)^{b-1}, \text{ where } x&gt;0.
$$


&lt;p&gt;What is $c$ ? Just a normalization constant! We&amp;rsquo;ll find the value of $c$ later.&lt;/p&gt;
&lt;h4 id=&#34;conjugate-prior&#34;&gt;Conjugate Prior&lt;/h4&gt;
&lt;p&gt;The Beta distribution is the &lt;strong&gt;conjugate prior&lt;/strong&gt; for the Bernoulli, binomial, negative binomial and geometric distributions (seems like those are the distributions that involve success &amp;amp; failure) in Bayesian inference.&lt;/p&gt;
&lt;mark&gt;Computing a posterior using a conjugate prior is very convenient, because you can avoid expensive numerical computation involved in Bayesian Inference.&lt;/mark&gt;
&lt;blockquote&gt;
&lt;p&gt;Conjugate prior = Convenient prior&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For example, the beta distribution is a conjugate prior to the binomial. &lt;strong&gt;If we choose to use the beta distribution Beta(α, β) as a prior, during the modeling phase, we already know the posterior will also be a beta distribution.&lt;/strong&gt; Therefore, after carrying out more experiments, &lt;strong&gt;you can compute the posterior simply by adding the number of successes (x), and failures (n-x) to the existing parameters α, β respectively&lt;/strong&gt;, instead of multiplying the likelihood with the prior distribution. The posterior also becomes a Beta distribution with parameters &lt;strong&gt;(x+α, n-x+β).&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id=&#34;what-is-the-intuition&#34;&gt;What is the Intuition?&lt;/h2&gt;
&lt;p&gt;The intuition for the beta distribution comes into play when we look at it from the lens of the binomial distribution.&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/1*eZKUz1_Jyvt6DNj8Tcv20Q.png&#34; alt=&#34;img&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;The difference between the binomial and the beta is that the &lt;strong&gt;former models the number of successes (x), while the latter models the probability (p) of success.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In other words, the probability is a &lt;strong&gt;parameter&lt;/strong&gt; in binomial; In the Beta, the probability is a &lt;strong&gt;random variable&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&#34;interpretation-of-α-β&#34;&gt;Interpretation of &lt;strong&gt;α, β&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;You can think of &lt;strong&gt;α-1 as the number of successes&lt;/strong&gt; and &lt;strong&gt;β-1 as the number of failures,&lt;/strong&gt; just like &lt;strong&gt;n&lt;/strong&gt; &amp;amp; &lt;strong&gt;n-x&lt;/strong&gt; terms in binomial.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;You can choose the α and β parameters however you think they are supposed to be&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If you think the probability of success is very high, let’s say 90%, &lt;strong&gt;set 90 for α&lt;/strong&gt; and &lt;strong&gt;10 for β.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;If you think otherwise, 90 for β and 10 for α.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;As &lt;strong&gt;α&lt;/strong&gt; becomes larger (more successful events), the bulk of the probability distribution will shift towards the right, whereas an increase in &lt;strong&gt;β&lt;/strong&gt; moves the distribution towards the left (more failures).&lt;/p&gt;
&lt;p&gt;Also, the distribution will narrow if both &lt;strong&gt;α&lt;/strong&gt; and &lt;strong&gt;β&lt;/strong&gt; increase, for we are more certain.&lt;/p&gt;
&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/%E6%88%AA%E5%B1%8F2022-08-03%2010.48.34.png&#34; alt=&#34;截屏2022-08-03 10.48.34&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;p&gt;Dr. Bognar at the University of Iowa built &lt;a href=&#34;https://homepage.divms.uiowa.edu/~mbognar/applets/beta.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;the calculator for Beta distribution&lt;/a&gt;, which I found useful and beautiful. You can experiment with different values of &lt;strong&gt;α&lt;/strong&gt; and &lt;strong&gt;β&lt;/strong&gt; and visualize how the shape changes.&lt;/p&gt;
&lt;h2 id=&#34;derivation&#34;&gt;Derivation&lt;/h2&gt;
&lt;p&gt;In this section, we&amp;rsquo;ll derive Beta distribution using the Beta-Gamma Connections.&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    We can regard Beta distribution as the &lt;strong&gt;fraction of waiting time&lt;/strong&gt;.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;fraction-of-waiting-time&#34;&gt;Fraction of Waiting Time&lt;/h3&gt;
&lt;p&gt;Let $X$  be the waiting time at Bank,&lt;/p&gt;

$$
X \sim Gamma(n_1, \lambda)
$$

&lt;p&gt;Let $Y$  be the waiting time at Post Office,&lt;/p&gt;

$$
Y \sim Gamma(n_2, \lambda)
$$

&lt;p&gt;Assume $X$  and $Y$  are independent. &lt;strong&gt;What is the distribution of the proportion&lt;/strong&gt; $\frac{X}{X+Y}$ ?&lt;/p&gt;
&lt;p&gt;Solution:&lt;/p&gt;
&lt;p&gt;Let $T := X+Y$ be the total waiting time. Clearly, $T \sim Gamma(n_1+n_2, \lambda)$, you can prove it by MGF.&lt;/p&gt;
&lt;p&gt;Let $W =: \frac{X}{X+Y}$ be the proportion of waiting time at Bank to the total waiting time. We need to find the PDF of $W$.&lt;/p&gt;
&lt;p&gt;The idea is to find the joint PDF $f_{T,W}(t,w)$ at first, and then get the marginal distribution.&lt;/p&gt;

$$
\begin{aligned}f_{T,W}(t,w) &amp;= f_{X,Y}(x,y) \left | \frac{\partial(x,y)}{\partial(t,w)} \right|\\
    &amp;= \frac{1}{\Gamma(n_1)}\lambda^{n_1}x^{n_1 - 1}e^{-\lambda x} \frac{1}{\Gamma(n_2)}\lambda^{n_2}x^{n_2 - 1}e^{-\lambda y} \left|-t\right|\\
    &amp;= \lambda^{n_1+n_2}t^{n_1+n_2-1}e^{-\lambda t} \frac{1}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\\
    &amp;= \frac{\lambda^{n_1+n_2}t^{n_1+n_2-1}e^{-\lambda t}}{\Gamma(n_1+n_2)} \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\\
    &amp;= f_T(t) \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\end{aligned}
$$

&lt;p&gt;Integrating $t$ out to get the marginal:&lt;/p&gt;

$$
\begin{aligned}f_W(w) &amp;= \int_0^\infty f_{T,W}(t,w) dt \\&amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \cdot\int_0^\infty f_T(t)dt \\&amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \end{aligned}
$$

&lt;p&gt;Here we have Beta,&lt;/p&gt;
$$W \sim Beta(n_1, n_2)$$
&lt;mark&gt;REMARK: the above result also proves &lt;strong&gt;W and T and independent&lt;/strong&gt;!&lt;/mark&gt;
&lt;h3 id=&#34;beta-function-as-a-normalizing-constant&#34;&gt;Beta Function as a normalizing constant&lt;/h3&gt;
&lt;p&gt;Note that, $f_W(w)$  is a PDF needed to be integrated to 1,&lt;/p&gt;

$$
\int_0^1\frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} dw \equiv 1
$$

&lt;p&gt;So the normalization constant should be,&lt;/p&gt;

$$
c = \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)} := \frac{1}{B(n_1, n_2)}
$$

&lt;h3 id=&#34;mean-of-beta-distribution&#34;&gt;Mean of Beta distribution&lt;/h3&gt;
&lt;p&gt;As a byproduct in the above derivation, we get that fact that &lt;strong&gt;W and T are independent&lt;/strong&gt;. Then, we can use this to derive the mean of Beta distribution,&lt;/p&gt;
$$E(WT) = E(W)E(T) \implies E(X) =E\left(\frac{X}{X+Y}\right)E(X+Y),$$
&lt;p&gt;rearrange to get,&lt;/p&gt;
$$E\left(\frac{X}{X+Y}\right) = \frac{E(X)}{E(X+Y)}$$
&lt;p&gt;This result is clear Not True in general, but under our setting, we have this interesting result.&lt;/p&gt;
&lt;p&gt;We can use this result to find the mean of $W \sim Beta(a, b)$ without the slightest trace of calculus.&lt;/p&gt;
$$E(W)=E\left(\frac{X}{X+Y}\right)=\frac{E(X)}{E(X+Y)}=\frac{a / \lambda}{a / \lambda+b / \lambda}=\frac{a}{a+b}$$
&lt;h3 id=&#34;getting-beta-parameters-in-practice&#34;&gt;Getting Beta parameters in practice&lt;/h3&gt;
&lt;p&gt;The Beta distribution is the conjugate prior for many common distributions. We use it a lot. But in practice, &lt;mark&gt;how to figure out its parameters?&lt;/mark&gt; It is sometimes useful to estimate quickly the parameters of the Beta distribution using the method of moments:&lt;/p&gt;
$$X \sim Beta(\alpha, \beta),$$
$$\alpha + \beta = \frac{E(X)(1-E(X))}{Var(X)} - 1,$$
$$\alpha = (\alpha + \beta)E(X),$$
$$\beta = (\alpha + \beta)(1 - E(X))$$
&lt;p&gt;Here is the R code:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# calculate beta params using method of moments&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;cal_beta_params&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;function&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sdX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;varX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sdX^2&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;sum_ab&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;*&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;/&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;varX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;a&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sum_ab&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;*&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;b&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sum_ab&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;*&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# return&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;c&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;s&#34;&gt;&amp;#34;shape&amp;#34;&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;a&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;scale&amp;#34;&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;b&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;p&#34;&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# example&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;cal_beta_params&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0.136&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sdX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0.103&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;##    shape    scale 
## 1.370320 8.705559
&lt;/code&gt;&lt;/pre&gt;&lt;h3 id=&#34;summary&#34;&gt;Summary&lt;/h3&gt;
&lt;p&gt;To summarize, the bank–post office story tells us that: when we add independent Gamma r.v.s $X$  and $Y$  with the same rate $\lambda$ ,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;the total $X+Y$ has a Gamma distribution;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the fraction $\frac{X}{X+Y}$ has a Beta distribution;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the total is independent of the fraction.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;examples&#34;&gt;Examples&lt;/h2&gt;
&lt;p&gt;The PDF of Beta distribution can be U-shaped with asymptotic ends, bell-shaped, strictly increasing/decreasing or even straight lines. As you change &lt;strong&gt;α&lt;/strong&gt; or &lt;strong&gt;β&lt;/strong&gt;, the shape of the distribution changes.&lt;/p&gt;
&lt;h3 id=&#34;i-bell-shape&#34;&gt;I. Bell-Shape&lt;/h3&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/bellshape-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;The PDF of a beta distribution is approximately normal if &lt;strong&gt;α&lt;/strong&gt; + &lt;strong&gt;β&lt;/strong&gt; is large enough and α &amp;amp; β are approximately equal.&lt;/p&gt;
&lt;h4 id=&#34;intuition-behind-bell-shape&#34;&gt;Intuition behind Bell-Shape&lt;/h4&gt;
&lt;p&gt;&lt;strong&gt;Why would Beta(2,2) be bell-shaped?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If you think:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;α-1 as the number of successes&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;β-1 as the number of failures&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Beta(2,2)&lt;/strong&gt; means you got 1 success and 1 failure&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So it makes sense that the probability of the success is highest at 0.5.&lt;/p&gt;
&lt;p&gt;Also, &lt;strong&gt;Beta(1,1)&lt;/strong&gt; would mean you got zero for the head and zero for the tail. Then, your guess about the probability of success should be the same throughout [0,1]. The horizontal straight line confirms it.&lt;/p&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/beta11-1.png&#34; width=&#34;672&#34; /&gt;
&lt;h3 id=&#34;ii-straight-lines&#34;&gt;II. Straight Lines&lt;/h3&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/unnamed-chunk-2-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;&lt;strong&gt;α = 1 or β = 1&lt;/strong&gt;, the beta PDF can be a straight line.&lt;/p&gt;
&lt;h3 id=&#34;iii-u-shape&#34;&gt;III. U-Shape&lt;/h3&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/unnamed-chunk-3-1.png&#34; width=&#34;672&#34; /&gt;
When $\alpha &lt; 1, \beta&lt;1$ the PDF of the Beta is U-shaped.
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://en.wikipedia.org/wiki/Beta_distribution&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;wiki Beta distribution&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://towardsdatascience.com/beta-distribution-intuition-examples-and-derivation-cf00f4db57af&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Beta Distribution — Intuition, Examples, and Derivation&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://stats.stackexchange.com/questions/47771/what-is-the-intuition-behind-beta-distribution&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;What is the intuition behind beta distribution?&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=UZjlBQbV1KU&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;youtube video Lecture 23: Beta distribution | Statistics 110&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=v1uUgTcInQk&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;another youtube video: Beta distribution - an introduction&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>Tutorial Notes for Bayesian Statistics</title>
      <link>https://chenxing.space/blog/notes-for-bayesian-statistics/</link>
      <pubDate>Tue, 26 Jul 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-for-bayesian-statistics/</guid>
      <description>&lt;h2 id=&#34;bayesian-inference&#34;&gt;Bayesian Inference&lt;/h2&gt;
&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/%E6%88%AA%E5%B1%8F2022-08-02%2009.32.20.png&#34; alt=&#34;截屏2022-08-02 09.32.20&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/%E5%9B%BE%E7%89%87%E6%9D%A5%E8%87%AA%20chapter5%EF%BC%8C%E7%AC%AC%2056%20%E9%A1%B5.png&#34; alt=&#34;img&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;h2 id=&#34;posterior-distribution&#34;&gt;Posterior Distribution&lt;/h2&gt;
&lt;p&gt;The intuition behind the posterior distribution:&lt;/p&gt;
&lt;div class=&#34;alert alert-note&#34;&gt;
  &lt;div&gt;
    &amp;ldquo;It is a &lt;strong&gt;weighted version of likelihood&lt;/strong&gt;! &amp;hellip; just weighting the likelihood using my prior belief on theta &amp;hellip;&amp;rdquo;
  &lt;/div&gt;
&lt;/div&gt;
&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/%E5%9B%BE%E7%89%87%E6%9D%A5%E8%87%AA%20chapter4%EF%BC%8C%E7%AC%AC%2053%20%E9%A1%B5.png&#34; alt=&#34;img&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;h2 id=&#34;useful-resource&#34;&gt;Useful Resource&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=bFZ-0FH5hfs&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;17. Bayesian Statistics MIT 18.650 Statistics for Applications, Fall 2016&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://github.com/rmcelreath/stat_rethinking_2022&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Statistical Rethinking (2022 Edition) github&lt;/a&gt;, this is a recommended bayesian textbook with video lectures.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Beta distribution as a prior</title>
      <link>https://chenxing.space/blog/beta-distribution-as-a-prior/</link>
      <pubDate>Mon, 25 Jul 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/beta-distribution-as-a-prior/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;The Beta distribution is a useful probability distribution when you want model uncertainty over a parameter bounded between 0 and 1.&lt;/p&gt;
&lt;p&gt;In this post, we&amp;rsquo;ll explore how the two parameters of the Beta distribution determine its shape.&lt;/p&gt;
&lt;p&gt;One way to see how the shape parameters of the Beta distribution affect its shape is to generate a large number of random draws using the &lt;code&gt;rbeta(n, shape1, shape2)&lt;/code&gt; function and visualize these as a histogram.&lt;/p&gt;
&lt;h2 id=&#34;beta11&#34;&gt;Beta(1,1)&lt;/h2&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Explore using the rbeta function&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Visualize the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;hist&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-as-a-prior/index.en_files/figure-html/unnamed-chunk-1-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;A Beta(1,1) distribution is the same as a uniform distribution between 0 and 1. It is useful as a so-called &lt;em&gt;non-informative&lt;/em&gt; prior as it expresses than any value from 0 to 1 is equally likely.&lt;/p&gt;
&lt;h2 id=&#34;discover-shape-parameters&#34;&gt;Discover shape parameters&lt;/h2&gt;
&lt;p&gt;What happened if you set shape parameters to negative numbers?&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Modify the parameters&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;-1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Explore the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;head&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;## [1] NaN NaN NaN NaN NaN NaN
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Yes, &lt;code&gt;NaN&lt;/code&gt; stands for &lt;em&gt;not a number&lt;/em&gt; and the reason you got a lot of &lt;code&gt;NaN&lt;/code&gt;s is that the Beta distribution is only defined when its shape parameters are &lt;strong&gt;positive&lt;/strong&gt;.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Modify the parameters&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;100&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;100&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Visualize the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;hist&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-as-a-prior/index.en_files/figure-html/unnamed-chunk-3-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;So the larger the shape parameters are, the more concentrated the beta distribution becomes. When used as a prior, this Beta distribution encodes the information that the parameter is most likely close to 0.5 .&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Modify the parameters&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;100&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;20&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Visualize the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;hist&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-as-a-prior/index.en_files/figure-html/unnamed-chunk-4-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;So the larger the &lt;code&gt;shape1&lt;/code&gt; parameter is the closer the resulting distribution is to 1.0 and the larger the &lt;code&gt;shape2&lt;/code&gt; the closer it is to 0.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Some Note for Pareto Distribution</title>
      <link>https://chenxing.space/blog/some-note-for-pareto-distribution/</link>
      <pubDate>Sat, 16 Jul 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/some-note-for-pareto-distribution/</guid>
      <description>&lt;h2 id=&#34;power-law-distribution&#34;&gt;Power Law Distribution&lt;/h2&gt;
&lt;div class=&#34;alert alert-note&#34;&gt;
  &lt;div&gt;
    &lt;p&gt;log-log-scale of cCDF showing you a &lt;strong&gt;straight line&lt;/strong&gt; ?&lt;/p&gt;
&lt;p&gt;This is the &lt;em&gt;signature of the Power Law distribution&lt;/em&gt;.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;r-code&#34;&gt;R Code&lt;/h2&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;zetaEDA&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;zetaclv&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;enable_zeta_ggplot_theme&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# transactional data for cohort 2019&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;cohort19&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;eg_trans_data&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;with_groups&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;cust&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;mutate&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;fp_yr&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;min&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;lubridate&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;::&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;year&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;date&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;filter&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;fp_yr&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;==&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;2019&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;select&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;-&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;fp_yr&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# build cbs data&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;dcbs&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;generate_cbs&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;cohort19&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;timeUnit&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;weeks&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;## Note that: time unit is in &amp;lt; weeks &amp;gt;
&lt;/code&gt;&lt;/pre&gt;&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;head&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;dcbs&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;##      cust x      t.x     litt sales sales.x      first     T.cal
## 1 uid0001 1 20.00000 2.995732  4644    1174 2019-12-02  79.00000
## 2 uid0005 0  0.00000 0.000000  1169       0 2019-08-08  95.57143
## 3 uid0006 1 50.71429 3.926208  1430     922 2019-04-20 111.28571
## 4 uid0010 0  0.00000 0.000000  2820       0 2019-02-15 120.42857
## 5 uid0011 0  0.00000 0.000000  6460       0 2019-01-15 124.85714
## 6 uid0012 0  0.00000 0.000000   473       0 2019-10-07  87.00000
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Note that &lt;code&gt;t.x&lt;/code&gt; is the &lt;em&gt;Time between first and last transactions&lt;/em&gt;. This is the &amp;ldquo;observed&amp;rdquo; part of lifetime. Let&amp;rsquo;s look at the distribution of &lt;code&gt;t.x&lt;/code&gt;.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;dtmp&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;dcbs&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# remove single purchase customers&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;filter&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# get value of cdf, P(X &amp;lt;= x)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;mutate&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;cdf&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;ecdf&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# get ccef, P(X &amp;gt; x)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;mutate&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;ccdf&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;cdf&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;dtmp&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;ggplot&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;aes&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;x&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;y&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;ccdf&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;geom_point&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;color&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;red&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;geom_line&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/some-note-for-pareto-distribution/index.en_files/figure-html/unnamed-chunk-1-1.png&#34; width=&#34;672&#34; /&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=9JkWtaVCqs0&amp;amp;t=751s&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Youtube video: Network Analysis. Lecture 2. Power laws.&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Note for Beta Distribution</title>
      <link>https://chenxing.space/blog/note-for-beta-distribution/</link>
      <pubDate>Sat, 30 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/note-for-beta-distribution/</guid>
      <description>

&lt;div id=&#34;TOC&#34;&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#why-beta-distribution&#34; id=&#34;toc-why-beta-distribution&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1&lt;/span&gt; Why Beta Distribution?&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#model-probabilities&#34; id=&#34;toc-model-probabilities&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.1&lt;/span&gt; Model probabilities&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#generalization-of-uniform&#34; id=&#34;toc-generalization-of-uniform&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.2&lt;/span&gt; Generalization of uniform&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#construction&#34; id=&#34;toc-construction&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2&lt;/span&gt; Construction&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#bank-and-post-office-story&#34; id=&#34;toc-bank-and-post-office-story&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.1&lt;/span&gt; Bank and Post Office Story&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#summary&#34; id=&#34;toc-summary&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.1.1&lt;/span&gt; Summary&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#plots&#34; id=&#34;toc-plots&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.2&lt;/span&gt; plots&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#reference&#34; id=&#34;toc-reference&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3&lt;/span&gt; Reference&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;p&gt;This post is out of date, please check the new post named “&lt;strong&gt;Beta Distribution — Intuition, Derivation, and Examples&lt;/strong&gt;”.&lt;/p&gt;
&lt;div id=&#34;why-beta-distribution&#34; class=&#34;section level1&#34; number=&#34;1&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;1&lt;/span&gt; Why Beta Distribution?&lt;/h1&gt;
&lt;div id=&#34;model-probabilities&#34; class=&#34;section level2&#34; number=&#34;1.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.1&lt;/span&gt; Model probabilities&lt;/h2&gt;
&lt;p&gt;The short story is that the Beta distribution can be understood as representing a distribution &lt;em&gt;of probabilities&lt;/em&gt;, that is, it represents all the possible values of a probability when we don’t know what that probability is.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;generalization-of-uniform&#34; class=&#34;section level2&#34; number=&#34;1.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.2&lt;/span&gt; Generalization of uniform&lt;/h2&gt;
&lt;p&gt;Give me a &lt;strong&gt;continuous&lt;/strong&gt; and &lt;strong&gt;bounded&lt;/strong&gt; random variable, em, except the &lt;em&gt;Uniform distribution&lt;/em&gt;. That is another way to look at &lt;em&gt;beta distribution&lt;/em&gt;, continuous and bounded between 0, 1; also the density is not flat.
&lt;span class=&#34;math display&#34;&gt;\[
X \sim Beta(a, b), \text{ where } a&amp;gt;0, \ b&amp;gt;0.
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
f_X(x) = c \cdot x ^{a-1}(1-x)^{b-1}, \text{ where } x&amp;gt;0.
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;What is &lt;span class=&#34;math inline&#34;&gt;\(c\)&lt;/span&gt;? Just a normalization constant. We’ll get the value of &lt;span class=&#34;math inline&#34;&gt;\(c\)&lt;/span&gt; later.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;construction&#34; class=&#34;section level1&#34; number=&#34;2&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;2&lt;/span&gt; Construction&lt;/h1&gt;
&lt;div id=&#34;bank-and-post-office-story&#34; class=&#34;section level2&#34; number=&#34;2.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.1&lt;/span&gt; Bank and Post Office Story&lt;/h2&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(X\)&lt;/span&gt; be the waiting time at Bank,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
X \sim Gamma(n_1, \lambda)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt; be the waiting time at Post Office,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
Y \sim Gamma(n_2, \lambda)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Assume &lt;span class=&#34;math inline&#34;&gt;\(X\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt; are independent.&lt;/p&gt;
&lt;p&gt;Then, what is the distribution of the proportion &lt;span class=&#34;math inline&#34;&gt;\(\frac{X}{X+Y}\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;Define &lt;span class=&#34;math inline&#34;&gt;\(T = X+Y\)&lt;/span&gt; be the total waiting time.&lt;/p&gt;
&lt;p&gt;Clearly, &lt;span class=&#34;math inline&#34;&gt;\(T \sim Gamma(n_1+n_2, \lambda)\)&lt;/span&gt;, proved by MGF.&lt;/p&gt;
&lt;p&gt;Define &lt;span class=&#34;math inline&#34;&gt;\(W = \frac{X}{X+Y}\)&lt;/span&gt; , the proportion of waiting time at Bank to the total waiting time.&lt;/p&gt;
&lt;p&gt;What is the distribution of &lt;span class=&#34;math inline&#34;&gt;\(W\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;We need to derive &lt;span class=&#34;math inline&#34;&gt;\(f_W(w)\)&lt;/span&gt;, but first let’s find the joint pdf &lt;span class=&#34;math inline&#34;&gt;\(f_{T,W}(t,w)\)&lt;/span&gt;
&lt;span class=&#34;math display&#34;&gt;\[
\begin{align}
    f_{T,W}(t,w) &amp;amp;= f_{X,Y}(x,y) \left | \frac{\partial(x,y)}{\partial(t,w)} \right|\\
    &amp;amp;= \frac{1}{\Gamma(n_1)}\lambda^{n_1}x^{n_1 - 1}e^{-\lambda x} \frac{1}{\Gamma(n_2)}\lambda^{n_2}x^{n_2 - 1}e^{-\lambda y} \left|-t\right| \\
    &amp;amp;= \lambda^{n_1+n_2}t^{n_1+n_2-1}e^{\lambda t} \frac{1}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \\
    &amp;amp;= \frac{\lambda^{n_1+n_2}t^{n_1+n_2-1}e^{\lambda t}}{\Gamma(n_1+n_2)} \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\\
    &amp;amp;= f_T(t) \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Then we find the marginal,&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{aligned}
f_W(w) &amp;amp;= \int_0^\infty f_{T,W}(t,w) dt \\

&amp;amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \cdot\int_0^\infty f_T(t)dt \\

&amp;amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}
\end{aligned}\]&lt;/span&gt;
&lt;p&gt;Since &lt;span class=&#34;math inline&#34;&gt;\(f_W(w)\)&lt;/span&gt; is the pdf needed to be integrated to 1,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\int_0^1\frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} dw \equiv 1
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;so the normalization constant should be&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
c = \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)} := \frac{1}{B(n_1, n_2)}
\]&lt;/span&gt;&lt;/p&gt;
&lt;div id=&#34;summary&#34; class=&#34;section level3&#34; number=&#34;2.1.1&#34;&gt;
&lt;h3&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.1.1&lt;/span&gt; Summary&lt;/h3&gt;
&lt;p&gt;The &lt;em&gt;&lt;u&gt;connection between Gamma and Beta distribution&lt;/u&gt;&lt;/em&gt; helps us to find the normalization constant in Beta. In summary,&lt;/p&gt;
&lt;p&gt;If &lt;span class=&#34;math inline&#34;&gt;\(X \sim Gamma(\alpha, \lambda)\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(Y \sim Gamma(\beta, \lambda)\)&lt;/span&gt; are independent, then &lt;span class=&#34;math inline&#34;&gt;\(\frac{X}{X+Y} \sim Beta(\alpha, \beta)\)&lt;/span&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;plots&#34; class=&#34;section level2&#34; number=&#34;2.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.2&lt;/span&gt; plots&lt;/h2&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;library(zetaEDA)
library(ggfortify)
enable_zeta_ggplot_theme()&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Let’s check Beta density for some different parameters value.&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a = b = 1\)&lt;/span&gt;? The Beta(1,1) is just the Unif(0,1).&lt;/p&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;ggdistribution(func = dbeta, x = seq(0, 1, .01), shape1 = 1, shape2 = 1) +
  labs(title = &amp;quot;Beta Density with a = 1, b = 1&amp;quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-2-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a = b = \frac{1}{2}\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-3-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a = b = 2\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-4-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a= 2, \ b = 1\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-5-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;One more,&lt;/p&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;p &amp;lt;- ggdistribution(func = dbeta, x = seq(0, 1, .01), shape1 = 1.5, shape2 = 5, colour = &amp;quot;tomato&amp;quot;, linetype = &amp;quot;dashed&amp;quot;)

ggdistribution(func = dbeta, x = seq(0, 1, .01), shape1 = 5, shape2 = 1.5, colour = &amp;quot;blue&amp;quot;, p = p) +
  labs(title = &amp;quot;Red: a = 1.5, b = 5\n Blue: a = 5, b = 1.5&amp;quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-6-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;For more checking, click &lt;strong&gt;&lt;a href=&#34;https://homepage.divms.uiowa.edu/~mbognar/applets/beta.html&#34;&gt;this link&lt;/a&gt;&lt;/strong&gt; and try some parameters to check the density curve.&lt;/p&gt;
&lt;p&gt;Have fun!&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;reference&#34; class=&#34;section level1&#34; number=&#34;3&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;3&lt;/span&gt; Reference&lt;/h1&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://en.wikipedia.org/wiki/Beta_distribution&#34;&gt;wiki Beta distribution&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://stats.stackexchange.com/questions/47771/what-is-the-intuition-behind-beta-distribution&#34;&gt;What is the intuition behind beta distribution?&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=UZjlBQbV1KU&#34;&gt;youtube video Lecture 23: Beta distribution | Statistics 110&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=v1uUgTcInQk&#34;&gt;another youtube video: Beta distribution - an introduction&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Note for Gamma Distribution</title>
      <link>https://chenxing.space/blog/note-for-gamma-distribution/</link>
      <pubDate>Sat, 30 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/note-for-gamma-distribution/</guid>
      <description>

&lt;div id=&#34;TOC&#34;&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#motivation-for-gamma-function&#34;&gt;Motivation for Gamma Function&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#definition-of-gamma-function&#34;&gt;Definition of Gamma Function&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#gamma-distribution&#34;&gt;Gamma Distribution&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#how-to-remember-the-gamma-pdf&#34;&gt;How to remember the Gamma pdf?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#gamma-exponential-connection&#34;&gt;Gamma &amp;amp; Exponential Connection&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;div id=&#34;motivation-for-gamma-function&#34; class=&#34;section level2&#34;&gt;
&lt;h2&gt;Motivation for Gamma Function&lt;/h2&gt;
&lt;p&gt;We all know how to compute the factorial of integer. BUT what is the factorial of 1/2?&lt;/p&gt;
&lt;p&gt;In other words, how to interpolate the factorial function?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://tva1.sinaimg.cn/large/e6c9d24egy1h3r86q0n40j206y058a9z.jpg&#34; alt=&#34;img&#34; style=&#34;zoom:150%;&#34;/&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The gamma function can be seen as a solution to the following interpolation problem:&lt;/p&gt;
&lt;p&gt;“Find a smooth curve that connects the points (x, y) given by y = (x − 1)! at the positive integer values for x.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;More details, check the &lt;a href=&#34;https://en.wikipedia.org/wiki/Gamma_function&#34;&gt;wiki page&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;definition-of-gamma-function&#34; class=&#34;section level2&#34;&gt;
&lt;h2&gt;Definition of Gamma Function&lt;/h2&gt;
&lt;div class=&#34;definition&#34;&gt;
&lt;p&gt;&lt;span id=&#34;def:unnamed-chunk-2&#34; class=&#34;definition&#34;&gt;&lt;strong&gt;Definition 1  &lt;/strong&gt;&lt;/span&gt;(Gamma Function)
&lt;span class=&#34;math display&#34;&gt;\[\begin{align}
\Gamma(z) = \int_{0}^{\infty}x^{z-1}e^{-x}dx, \ \ z \in \mathbb{R}^+
\end{align}\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;For the the &lt;strong&gt;Gamma&lt;/strong&gt; &lt;strong&gt;function&lt;/strong&gt;, it is enough to know the following properties for now.&lt;/p&gt;
&lt;div class=&#34;lemma&#34;&gt;
&lt;p&gt;&lt;span id=&#34;lem:unnamed-chunk-3&#34; class=&#34;lemma&#34;&gt;&lt;strong&gt;Lemma 1  &lt;/strong&gt;&lt;/span&gt;&lt;span class=&#34;math display&#34;&gt;\[\begin{align}
&amp;amp; \Gamma(z+1) = z\Gamma(z), \ \ z \in \mathbb{R}^+ \\
&amp;amp; \Gamma(n) = (n-1)!, \  \ n = 1,2,3,...
\end{align}\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Easy to prove using integration by parts.&lt;/p&gt;
&lt;p&gt;Now, what is &lt;span class=&#34;math inline&#34;&gt;\(\Gamma(\frac{1}{2})\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\Gamma(\frac{1}{2}) = \int_{0}^{\infty}x^{-1/2}e^{-x}dx = ?
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Recall that, &lt;span class=&#34;math inline&#34;&gt;\(\int_{0}^{\infty} e^{-x^2} = \frac{1}{2}\sqrt{\pi}\)&lt;/span&gt;, let &lt;span class=&#34;math inline&#34;&gt;\(u = x^2\)&lt;/span&gt; and will get the result &lt;span class=&#34;math inline&#34;&gt;\(\Gamma(\frac{1}{2} )= \sqrt{\pi}\)&lt;/span&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;gamma-distribution&#34; class=&#34;section level2&#34;&gt;
&lt;h2&gt;Gamma Distribution&lt;/h2&gt;
&lt;p&gt;From the Gamma function, it is pretty natural to get Gamma pdf. JUST &lt;strong&gt;normalizing&lt;/strong&gt;!&lt;/p&gt;
&lt;p&gt;Clearly,
&lt;span class=&#34;math display&#34;&gt;\[
1 = \int_0^\infty \frac{x^{r-1}e^{-x}}{\Gamma(r)}dx = \int_0^\infty f_X(x)dx, \ \ \ X := Gamma(r, 1)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;What is the pdf for the general &lt;span class=&#34;math inline&#34;&gt;\(Gamma(r, \lambda)\)&lt;/span&gt; ? Let&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
Y = \frac{X}{\lambda}, \ Y \sim Gamma(r, \lambda)
\]&lt;/span&gt;
&lt;span class=&#34;math display&#34;&gt;\[
f_Y(y) = f_X(x)\frac{dx}{dy}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;We’ll get&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
f(y; r, \lambda ) = \frac{\lambda ^{r}y^{r-1}e^{-\lambda y}}{\Gamma(r)}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Here, &lt;span class=&#34;math inline&#34;&gt;\(r\)&lt;/span&gt; is called the &lt;strong&gt;shape&lt;/strong&gt; parameter and &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; is called the &lt;strong&gt;rate&lt;/strong&gt; parameter.&lt;/p&gt;
&lt;div id=&#34;how-to-remember-the-gamma-pdf&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;How to remember the Gamma pdf?&lt;/h3&gt;
&lt;p&gt;That’s my trick: exponential density times the power rise to (shape-1), then divided by normalizer.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exponential&lt;/strong&gt; density (very familiar): &lt;span class=&#34;math inline&#34;&gt;\(\lambda e^{-\lambda x}\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;power&lt;/strong&gt; rise to (shape-1): &lt;span class=&#34;math inline&#34;&gt;\((\lambda x)^{r-1}\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;normalizing&lt;/strong&gt; constant (using shape): &lt;span class=&#34;math inline&#34;&gt;\(\Gamma(r)\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{align}
f(x; r, \lambda ) &amp;amp;= \frac{\text{exp density} \cdot \text{power}^\text{shape-1} }{normalizer} \\
&amp;amp; = \frac{\lambda e^{-\lambda x}(\lambda x)^{r-1}}{\Gamma(r)} \\
&amp;amp; = \frac{\lambda ^{r}x^{r-1}e^{-\lambda x}}{\Gamma(r)}
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;Another way to remember is this:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exponential&lt;/strong&gt; key part:
&lt;span class=&#34;math display&#34;&gt;\[
e^{-\lambda x}
\]&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Add &lt;strong&gt;Power&lt;/strong&gt; part:
&lt;span class=&#34;math display&#34;&gt;\[
\lambda^{\square} x^{\square} e^{- \lambda x}
\]&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;multiply the &lt;em&gt;power part&lt;/em&gt; in exponential&lt;/li&gt;
&lt;li&gt;&lt;em&gt;&lt;u&gt;rate&lt;/u&gt;&lt;/em&gt; rises to &lt;em&gt;&lt;u&gt;shape&lt;/u&gt;&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;&lt;u&gt;variable&lt;/u&gt;&lt;/em&gt; rises to &lt;em&gt;&lt;u&gt;shape-1&lt;/u&gt;&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\lambda^{r} x^{r-1} e^{- \lambda x}
\]&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Add &lt;strong&gt;Normalizing&lt;/strong&gt; part:
&lt;span class=&#34;math display&#34;&gt;\[
\frac{\lambda^{r} x^{r-1} e^{- \lambda x}}{\Gamma(r)}
\]&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;div id=&#34;gamma-exponential-connection&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;Gamma &amp;amp; Exponential Connection&lt;/h3&gt;
&lt;p&gt;Let’s recall the Poisson Process,
&lt;span class=&#34;math display&#34;&gt;\[
N_t = \text{number of arrials up to time t} \sim Pois(\lambda t)
\]&lt;/span&gt;
The number of arrivals in the &lt;strong&gt;disjoint&lt;/strong&gt; intervals are &lt;strong&gt;independent&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(T_1\)&lt;/span&gt; be the time of &lt;strong&gt;1st&lt;/strong&gt; arrival,
&lt;span class=&#34;math display&#34;&gt;\[
P(T_1 &amp;gt; t) = P(N_t = 0) = e^{-\lambda t} \ \implies T_1 \sim Exp(\lambda)
\]&lt;/span&gt;
Now, let &lt;span class=&#34;math inline&#34;&gt;\(T_n\)&lt;/span&gt; be the time of &lt;strong&gt;nth&lt;/strong&gt; arrival, that is, &lt;span class=&#34;math inline&#34;&gt;\(T_n = \sum_{i=1}^n X_i\)&lt;/span&gt; , where &lt;span class=&#34;math inline&#34;&gt;\(X_i \overset{\text{iid}}{\sim}Exp(\lambda)\)&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;What is the pdf of &lt;span class=&#34;math inline&#34;&gt;\(T_n\)&lt;/span&gt;? Answer is &lt;strong&gt;Gamma&lt;/strong&gt;!&lt;/p&gt;
&lt;div class=&#34;proposition&#34;&gt;
&lt;p&gt;&lt;span id=&#34;prp:unnamed-chunk-4&#34; class=&#34;proposition&#34;&gt;&lt;strong&gt;Proposition 1  &lt;/strong&gt;&lt;/span&gt;Gamma is the sum of iid Exponentials.&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Proof:&lt;/p&gt;
&lt;p&gt;Since &lt;span class=&#34;math inline&#34;&gt;\(M_X(t) = \frac{\lambda}{\lambda-t}\)&lt;/span&gt;, where &lt;span class=&#34;math inline&#34;&gt;\(t &amp;lt; \lambda\)&lt;/span&gt;,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
M_{\sum_{i=1}^n X_i}(t) = (M_X(t))^n = (\frac{\lambda}{\lambda-t})^n
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;It is enough to show the MGF of Gamma equals the above value.&lt;/p&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(Y \sim Gamma(n, \lambda)\)&lt;/span&gt;,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{align}
M_Y(t) = E(e^{ty}) &amp;amp;= \int_0^\infty e^{ty} \frac{1}{\Gamma(n)} \lambda^{n}e^{-\lambda y}y^{n-1} dy \\
&amp;amp;= \frac{\lambda^n}{\Gamma(n)} \int_0^\infty e^{-(\lambda - t)y}y^{n-1}dy, \text{ let } u =  (\lambda - t)y \\
&amp;amp;= \frac{\lambda^n}{\Gamma(n)} (\frac{1}{\lambda-t})^n \int_0^\infty e^{-u}u^{n-1}du \\
&amp;amp;= \frac{\lambda^n}{\Gamma(n)} (\frac{1}{\lambda-t})^n\Gamma(n) \\
&amp;amp;= (\frac{\lambda}{\lambda-t})^n
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Proved!&lt;/p&gt;
&lt;div class=&#34;remark&#34;&gt;
&lt;p&gt;&lt;span id=&#34;unlabeled-div-1&#34; class=&#34;remark&#34;&gt;&lt;em&gt;Remark&lt;/em&gt;. &lt;/span&gt;The &lt;strong&gt;Exponential&lt;/strong&gt; is the continuous analog of the &lt;strong&gt;Geometric&lt;/strong&gt;. Similarly, the &lt;strong&gt;Gamma&lt;/strong&gt; is the continuous analog of the &lt;strong&gt;Negative Binomial&lt;/strong&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Probability Theory Study Notes</title>
      <link>https://chenxing.space/blog/probability-theory-study-notes/</link>
      <pubDate>Thu, 17 Dec 2020 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/probability-theory-study-notes/</guid>
      <description>&lt;h2 id=&#34;online-courses&#34;&gt;Online Courses&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://mathweb.ucsd.edu/~tkemp/ProbabilityTube/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Probability Theory and Stochastic Processes (video lecture)&lt;/a&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Description: These lectures encompass a full-year course in probability theory and stochastic processes, as taught at the University of California, San Diego (as Math 280). These lectures were produced in the 2020/2021 school year, in the midst of the Covid-19 pandemic.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/channel/UCeKkMyeKBnec9Y3I_eI2BNQ&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;YouTube Video List here.&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://stars.bilkent.edu.tr/syllabus/view/IE/523/20161?section=1&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Graduate-level probability course by Bilkent University (in Turkey) (video lecture)&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Description: As explained by the professor in the 1st lecture, it uses the measure-theoretic approach to introduce probability concepts.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/playlist?list=PL5B3KLQNAC5jT6yjV1199ji1zUy1YUp6P&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;YouTube Video List here.&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>Notes for LP Inequality</title>
      <link>https://chenxing.space/blog/notes-for-lp-inequality/</link>
      <pubDate>Wed, 24 May 2017 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-for-lp-inequality/</guid>
      <description>&lt;p&gt;Here, I summarize some useful LP Inequalities with proofs.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://chenxing.space/uploads/LP-inequality.pdf&#34;&gt;Download PDF file here.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303241006587.png&#34; alt=&#34;page 1&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303241007316.png&#34; alt=&#34;page 2&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303241006615.png&#34; alt=&#34;page 3&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303241006610.png&#34; alt=&#34;page 4&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303241006622.png&#34; alt=&#34;page 5&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Deriving the conditional distributions of a multivariate normal distribution</title>
      <link>https://chenxing.space/blog/deriving-the-conditional-distributions-of-a-multivariate-normal-distribution/</link>
      <pubDate>Fri, 31 Mar 2017 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/deriving-the-conditional-distributions-of-a-multivariate-normal-distribution/</guid>
      <description>&lt;p&gt;Overall, the intuition behind the conditional distribution of a bivariate normal is that even though the two variables are correlated in the joint distribution, they can still be treated as independent when you&amp;rsquo;re looking at the distribution of one variable, given a fixed value of the other variable. This allows you to use the normal distribution, which is a well-understood and widely-used probability distribution, to model the conditional distribution of a bivariate normal.&lt;/p&gt;
&lt;p&gt;More generally,&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303311241754.png&#34; alt=&#34;image-20230331124124725&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Assume we know there&amp;rsquo;s a theorem that says all conditional distributions of a multivariate normal distribution are normal.&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303311238287.png&#34; alt=&#34;image-20230331123847757&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
