Appendix A — List of Figures, Tables, and Apps

A consolidated index of every numbered figure and table in the book, plus the interactive Shiny apps. Click any entry to jump to it in context.

A.1 List of Shinylive Apps

The book’s interactive Shinylive apps are collected in a standalone archive: Shinylive Apps. Each app runs in the browser and carries its R source and the notes that accompany it.

A.2 List of Figures

A.2.1 Introduction to Statistical Computation

  • Figure 1.1 — Illustrates the two components of estimator quality, bias and variance, by contrasting the sampling distributions of two estimators of the same parameter.
  • Figure 1.2 — Shows how the error rates and the \(p\)-value of a hypothesis test all arise from the null and alternative distributions of a single test statistic.
  • Figure 1.3 — Previews Bayesian updating, in which the posterior combines the prior with the likelihood, here in a case where the data dominate the prior.
  • Figure 1.4 — Demonstrates Monte Carlo simulation as a way of approximating a distribution, using an example whose exact answer is known so that the approximation can be checked.
  • Figure 1.5 — A one-dimensional warm-up for numerical optimization, showing what it means to maximize a function before moving to problems that cannot be plotted.
  • Figure 1.6 — Introduces numerical quadrature, the deterministic alternative to Monte Carlo for computing integrals, by turning the area under a density into a sum of rectangle areas.

A.2.2 Introduction to R and Quarto

  • Figure 2.1 — Demonstrates how to summarize a single numeric variable in R with a histogram, which shows the shape, centre, and spread of the class scores.
  • Figure 2.2 — Demonstrates how to compare a numeric variable across groups in R, using boxplots with jittered points so that both the group summaries and the individual scores are visible.
  • Figure 2.3 — Demonstrates how to visualize a relationship separately within groups in R, by fitting one regression line per program to a scatterplot.

A.2.3 Computer Arithmetic

  • Figure 3.1 — Illustrates how floating-point numbers are spaced on the real line in a toy number system: the gaps between representable numbers grow with magnitude, so precision is relative rather than absolute.
  • Figure 3.2 — Compares implementations of log(1 - exp(-u)) to show which formulas remain accurate for very small and very large u.

A.2.4 Random Numbers and Monte Carlo Methods

  • Figure 4.1 — Checks the quality of a pseudo-random number generator written from scratch by comparing it with R’s built-in generator on three diagnostics: behaviour over time, autocorrelation, and uniformity.
  • Figure 4.2 — Shows the simplest transformation of uniform random numbers, a change of location and scale.
  • Figure 4.3 — Shows geometrically how the inverse-CDF method turns uniform draws into draws from a target distribution.
  • Figure 4.4 — Verifies the inverse-CDF method on a distribution whose inverse CDF has a closed form.
  • Figure 4.5 — Shows the two facts behind the Box-Muller transform: the angle of a pair of independent standard normals is uniform, and its squared radius is exponential. \(3000\) pairs \((Z_1, Z_2)\) of iid \(N(0,1)\) draws (left), with red circles enclosing 50%, 90% and 99% of the points; the squared radius \(R^2\) follows \(\mathrm{Exp}(1/2)\) (top right) and the angle \(\Theta\) is uniform on \((0, 2\pi)\) (bottom right).
  • Figure 4.6 — Shows how uniform random numbers can be transformed into normal ones, and verifies the Box-Muller construction against R’s rnorm().
  • Figure 4.7 — Illustrates the two theorems behind Monte Carlo methods: the Law of Large Numbers, under which averages converge to the mean, and the Central Limit Theorem, under which their fluctuations are approximately normal.
  • Figure 4.8 — Shows the Central Limit Theorem as a statement about distributions: the standardized mean of skewed data becomes approximately normal as the sample size grows.
  • Figure 4.9 — Illustrates Monte Carlo integration with a geometric example, estimating \(\pi\) as the proportion of random points that fall inside a circle. 1,000 uniform points (‘darts’) thrown at the \([-1,1]^2\) square; the fraction landing inside the inscribed circle gives a Monte Carlo estimate of \(\pi\).
  • Figure 4.10 — Shows the Monte Carlo error bound at work along a single simulation run: the running estimate of \(\pi\) stays within a band that narrows at rate \(1/\sqrt{n}\).
  • Figure 4.11 — Uses simulation to compare two estimators of a binomial proportion by their mean squared error, showing when a small bias pays off through lower variance.
  • Figure 4.12 — Shows that the p-value is a tail area of the null distribution of the test statistic, and equivalently the null survival function evaluated at the observed statistic.
  • Figure 4.13 — Illustrates the basic simulation check of a test: p-values are uniform under the null hypothesis and concentrate near 0 under the alternative. p-values of the one-sample \(t\)-test for \(n = 20\) normal observations under \(H_0: \mu = 0\) (left) and \(H_1: \mu = 0.5\) (right); red bars are the rejections at \(\alpha = 0.05\).
  • Figure 4.14 — Uses simulation to check the size of the t-test: under the null hypothesis a valid test gives uniform p-values, and here the data are nearly normal.
  • Figure 4.15 — Repeats the size check of the t-test for heavy-tailed data to show how violating the normality assumption distorts the p-values.
  • Figure 4.16 — Pushes the size check to the extreme case of Cauchy data, which have no mean, to show where the t-test breaks down completely.
  • Figure 4.17 — Shows that a larger sample size can rescue the t-test for heavy-tailed data, because the Central Limit Theorem makes the sample mean nearly normal.
  • Figure 4.18 — Uses simulation to estimate the power of the t-test, its ability to detect a true mean of 0.5, when the data are moderately heavy-tailed. \(t\)-test p-values at \(\nu = 5\) with true mean \(\mu = 0.5\), for \(n = 30\) and \(n = 60\); the red rejection rate is the power.
  • Figure 4.19 — Repeats the power study for heavier-tailed data to show how tail weight affects the ability of the t-test to detect a real effect. \(t\)-test p-values at \(\nu = 3\) with true mean \(\mu = 0.5\), for \(n = 30\) and \(n = 60\); the red rejection rate is the power.
  • Figure 4.20 — Verifies by simulation that the permutation test for correlation has the correct size without any normality assumption, and estimates its power.
  • Figure 4.21 — Demonstrates that choosing the best of many predictors using the data and then testing it as if it had been chosen in advance gives an invalid p-value. p-values of test_lm_sel under \(H_0\) (left) and \(H_1: \beta = 0.5\) (right); under \(H_0\) they are far from uniform, with a rejection rate far above 0.05.
  • Figure 4.22 — Shows why the selection step must be repeated inside the permutation loop: the selected predictor depends on the response, so every permuted response selects its own predictor.
  • Figure 4.23 — Shows that a permutation test which repeats the whole selection step on each shuffled data set restores valid p-values while keeping good power.
  • Figure 4.24 — Shows how plugging estimated parameters into a textbook null distribution biases a goodness-of-fit test, and how a parametric bootstrap that re-estimates the parameters on every simulated data set removes the bias. p-values of the KS normality test with estimated mean and variance for \(n = 50\): the naive test (top) against the Monte Carlo test (bottom), for normal data (\(H_0\)), \(t_3\) data and exponential data; red bars are the rejections at \(\alpha = 0.05\).

A.2.5 Optimization for Maximum Likelihood Estimation

  • Figure 5.1 — Introduces the likelihood and the negative log-likelihood with the simplest possible example, showing that the parameter value making the data most probable is the one that minimizes the function every optimizer in this chapter works with.
  • Figure 5.2 — Relates the three central objects of maximum likelihood, the negative log-likelihood, its gradient, and its curvature, and shows how each changes with sample size.
  • Figure 5.3 — Shows how the Cauchy location negative log-likelihood, which is harder to minimize than the normal one, concentrates around the true value as the sample size grows.
  • Figure 5.4 — Shows the sampling variability of the negative log-likelihood itself: different data sets give different curves, and the minima of those curves, the MLEs, become less variable as \(n\) grows. 50 repeated Cauchy negative log-likelihood curves at \(n = 20\) versus \(n = 200\): the curves cluster more tightly, and their minima concentrate, as \(n\) grows.
  • Figure 5.5 — Sets up the running example of a zero-truncated Poisson model by showing the negative log-likelihood that every method in this section minimizes, together with its gradient, whose root is the minimizer.
  • Figure 5.6 — Illustrates fixed-point iteration, the simplest scheme for solving the stationarity equation, and how its iterates converge to the solution.
  • Figure 5.7 — Illustrates bisection, a slow but robust root-finding scheme that only needs an interval known to contain the root.
  • Figure 5.8 — Illustrates Newton-Raphson, which replaces the gradient of the negative log-likelihood by its tangent line at each step, and contrasts it with Fisher scoring, which uses the expected rather than the observed curvature.
  • Figure 5.9 — Explains why Newton-Raphson can fail on the Cauchy likelihood by showing the shape of the contribution of a single observation.
  • Figure 5.10 — Illustrates golden section search, a derivative-free method for optimizing a one-dimensional function, and why it needs only one new function evaluation per step.
  • Figure 5.11 — Tests golden section search on a smooth likelihood and on a non-differentiable function to show that the method needs no derivatives.
  • Figure 5.12 — Compares the convergence speed of three iterative algorithms on the same problem by showing how quickly each reduces its error.
  • Figure 5.13 — Shows how the curvature of the negative log-likelihood at the MLE yields a standard error and a Wald confidence interval, by approximating it with a parabola.
  • Figure 5.14 — Uses simulation to show why the MLE is worth computing: ignoring the truncation gives a biased estimator, whereas the MLE is centred on the truth.
  • Figure 5.15 — Gives a geometric picture of the gradient and the Hessian in two dimensions, explaining why the direction of steepest descent generally does not point at the minimum.
  • Figure 5.17 — Shows the simulated logistic-regression data used as a running example for multivariate optimization.
  • Figure 5.18 — Shows what a hidden layer buys in a Poisson regression, by displaying the three scaled hidden units and the non-monotone mean curve they combine into, together with counts simulated from it.
  • Figure 5.20 — Illustrates the main weakness of gradient descent: on an elongated objective it makes slow zig-zag progress.
  • Figure 5.21 — Shows how conjugate gradient repairs the zig-zag behaviour of gradient descent by choosing search directions that do not undo earlier progress.
  • Figure 5.22 — Compares stochastic gradient descent, which is cheap per step but noisy, with Newton-Raphson, which is expensive per step but takes large steps, on the same logistic regression problem.

A.2.6 EM Algorithms

  • Figure 6.1 — Illustrates the \(Q\) function of the EM algorithm as the expectation of the complete-data log-likelihood over the unobserved data, approximated here by averaging simulated completions.
  • Figure 6.2 — Illustrates why EM increases the likelihood: each iteration maximizes a lower bound that touches the observed log-likelihood at the current estimate.
  • Figure 6.3 — Introduces the two-component normal mixture model by showing that its density is the weighted sum of the two component densities.
  • Figure 6.4 — Shows that a mixture density can look unimodal or bimodal depending on its parameters, which affects how easily those parameters can be estimated.
  • Figure 6.5 — Illustrates the E-step of EM for a mixture: the responsibility of a component is the posterior probability that an observation came from it.
  • Figure 6.6 — Shows the simulated data used to demonstrate EM for a mixture; the colours reveal the component labels, which are unobserved in practice and must be inferred.
  • Figure 6.7 — Illustrates the conditional distributions needed in the E-step when an observation is only known to lie on one side of a cut point.
  • Figure 6.8 — Illustrates the E-step for censored Poisson counts: the missing count follows the Poisson distribution restricted to the values consistent with the censoring information.
  • Figure 6.9 — Illustrates model-based clustering, the multivariate extension of the normal mixture model, in which EM assigns observations to clusters of different shapes.
  • Figure 6.10 — Illustrates a hidden Markov model, whose unobserved state sequence generates the data, as a further example to which EM applies.

A.2.7 Introduction to Bayesian Inference

  • Figure 7.1 — Shows how Bayes’ rule combines a prior with a likelihood into a posterior, and how a credible interval is read off from the posterior.
  • Figure 7.2 — Shows how the posterior predictive distribution is built by averaging the sampling density over the posterior, so that it accounts for uncertainty about the parameter.
  • Figure 7.3 — Shows the range of prior beliefs about a probability that the Beta family can express, to guide the choice of a prior.
  • Figure 7.4 — Shows how the influence of the prior fades as data accumulate, so that the posterior concentrates around the MLE.
  • Figure 7.5 — Shows why the posterior predictive distribution is more dispersed than a plug-in prediction: it also reflects the remaining uncertainty about the Poisson rate.
  • Figure 7.6 — Shows why the posterior predictive distribution is wider than a plug-in prediction: it also reflects the remaining uncertainty about the mean.

A.2.8 Numerical Quadrature

  • Figure 8.1 — Shows how a change of variable maps an integral over the whole real line onto a finite interval, so that a rule with equally spaced points can be applied.
  • Figure 8.2 — Shows how the error of each quadrature rule shrinks as the number of panels grows, which reveals the order of the rule.
  • Figure 8.3 — Illustrates that the marginal likelihood \(P(y)\) is the integral of the likelihood times the prior, and that it changes with the prior even when the likelihood is fixed.
  • Figure 8.4 — Illustrates how a two-dimensional posterior integral is computed by transforming it to the unit square and applying a product midpoint grid.
  • Figure 8.5 — Illustrates the curse of dimensionality for quadrature: a grid wastes most of its points on negligible regions, and the number of evaluations explodes with the dimension.

A.2.9 Laplace Approximation

  • Figure 9.1 — Shows how the Laplace approximation replaces a posterior by a normal density centred at the mode, and how well it fits at small and large sample sizes.

A.2.10 Rejection Sampling

  • Figure 10.1 — Illustrates the rejection sampling algorithm: candidate points are drawn uniformly under an envelope and kept when they fall under the target density.
  • Figure 10.2 — Verifies that the Gamma rejection sampler produces draws from the correct distribution.
  • Figure 10.3 — Verifies that adaptive rejection sampling produces draws from the correct truncated normal distribution, even far out in the tail.

A.2.11 Importance Sampling

  • Figure 11.1 — Demonstrates a silent failure of importance sampling: when the proposal misses a mode of the target, the weights can look well behaved while the estimate is wrong.
  • Figure 11.2 — Compares a naive bounding-box proposal with a proposal matched to the target’s support for estimating the area of an irregular region. 1,500 darts thrown at the crescent \(A = D_1 \setminus D_2\) (outlined in black) from a generous bounding box (left) and from \(D_1\) itself (right); blue points land in \(A\), red points miss.

A.2.12 Markov Chain Monte Carlo

  • Figure 12.1 — Illustrates the long-run behaviour of Markov chains by contrasting two random walks on ten states with different transition rules.
  • Figure 12.2 — Demonstrates the Gibbs sampler on a normal model with unknown mean and standard deviation, using trace plots to check convergence and densities to summarize the posterior.
  • Figure 12.3 — Demonstrates random-walk Metropolis-Hastings on a difficult target, a truncated normal distribution, and shows how the step size affects the result.
  • Figure 12.4 — Illustrates the trade-off in choosing the step size of a random-walk Metropolis-Hastings sampler, using trace plots and autocorrelation.
  • Figure 12.5 — Illustrates how to tune the step size in practice, by comparing short preliminary runs at several scales.
  • Figure 12.6 — Compares Hamiltonian Monte Carlo with random-walk Metropolis on a highly correlated target, where the random walk struggles.
  • Figure 12.7 — Shows a stochastic-volatility model fitted with Stan: the simulated data and the posteriors of the two parameters of interest.

A.3 List of Tables

A.3.1 Introduction to R and Quarto

  • Table 2.1 — Shows the kind of tidy data frame used to practise data handling in R, with one row per student and one column per assessment category.

A.3.2 Computer Arithmetic

  • Table 3.1 — Compares a naive Taylor-series implementation of the exponential function with the built-in exp() to show that the series is accurate for positive arguments but fails for negative ones.
  • Table 3.2 — Exposes the internals of the failing computation of exp(-50), showing that very large terms of alternating sign cancel and leave only roundoff error.
  • Table 3.3 — Shows that reformulating the computation as 1/exp(|x|) removes the cancellation and restores accuracy for negative arguments.
  • Table 3.4 — Demonstrates catastrophic cancellation in the one-pass variance formula, which breaks down when the data have a large mean, while the two-pass formula stays accurate.

A.3.3 Random Numbers and Monte Carlo Methods

  • Table 4.1 — Shows how the accuracy of the Monte Carlo estimate of \(\pi\), as measured by its margin of error, improves with the number of points.
  • Table 4.2 — Compares three estimators of the centre of a symmetric distribution under normal and heavy-tailed data, to show that the best estimator depends on the data distribution.

A.3.4 Introduction to Bayesian Inference

  • Table 7.1 — Collects every quantity reported in a conjugate Beta-Binomial analysis, from the prior and MLE to the posterior, credible interval, and predictive distribution.
  • Table 7.2 — Collects every quantity reported in a conjugate Poisson-Gamma analysis, from the prior and MLE to the posterior, credible interval, and predictive distribution.
  • Table 7.3 — Collects every quantity reported in a conjugate Normal-Normal analysis, from the prior and MLE to the posterior, credible interval, and predictive distribution.

A.3.5 Numerical Quadrature

  • Table 8.2 — Compares the accuracy of the midpoint, trapezoidal, and Simpson rules on the same integral with the same small number of panels.
  • Table 8.3 — Compares quadrature with naive Monte Carlo for computing the log marginal likelihood, to show that a deterministic grid is much more accurate in low dimensions.
  • Table 8.4 — Shows how the quadrature estimate converges as the grid is refined, which indicates how many points are needed.
  • Table 8.5 — Shows how the log marginal likelihood, the quantity used for Bayesian model comparison, changes with the prior mean and prior standard deviation.
  • Table 8.6 — Shows how the marginal likelihood can be used to compare models, here t models with different degrees of freedom fitted to Gaussian and to heavy-tailed data.
  • Table 8.7 — Shows a failure mode of grid quadrature and naive Monte Carlo: when the posterior mass lies far from where the grid or the prior draws are placed, both estimates are poor.

A.3.6 Laplace Approximation

  • Table 9.1 — Assesses the accuracy of the Laplace approximation by comparing it with the exact marginal likelihood in a conjugate example where the exact answer is known.
  • Table 9.2 — Compares interval estimates based on the Laplace approximation with the exact posterior intervals, to see how much accuracy is lost.
  • Table 9.3 — Compares the Laplace approximation with grid quadrature and naive Monte Carlo against the true value as the posterior moves away from the prior mean, to show where each method succeeds or fails.

A.3.7 Rejection Sampling

  • Table 10.1 — Verifies the rejection sampler for the normal distribution by comparing its acceptance rate and sample moments with theory.
  • Table 10.2 — Checks that a proposed envelope really lies above the target density, which rejection sampling requires, for several parameter values.
  • Table 10.3 — Shows how the efficiency of the Gamma rejection sampler, measured by its acceptance rate, changes with the shape parameter.
  • Table 10.4 — Compares the cost of adaptive rejection sampling with that of naive rejection sampling for truncated normal distributions.

A.3.8 Importance Sampling

  • Table 11.1 — Shows how the choice of bounding constants affects importance sampling of a Student-\(t\) interval, including a case where the weights have infinite variance.
  • Table 11.2 — Compares importance sampling proposals for estimating a marginal likelihood, showing that a Laplace-shaped proposal is more reliable than the prior as \(n\) grows.
  • Table 11.3 — Shows that both proposals are unbiased for the true area, and that only the effective sample size distinguishes them: the disk proposal wastes draws only on \(D_2\), while the box proposal also wastes draws outside \(D_1\) entirely.

A.3.9 Markov Chain Monte Carlo

  • Table 12.1 — Summarizes the posterior distribution obtained from the Gibbs sampler.
  • Table 12.2 — Shows that Gibbs sampling extends to rounded (censored) data by treating the exact values as latent variables, and that the parameters are still recovered.
  • Table 12.3 — Checks the Metropolis-Hastings sampler against Gibbs sampling on the same problem.
  • Table 12.4 — Validates Stan’s sampler against the samplers written by hand earlier in the chapter.
  • Table 12.5 — Reports the convergence diagnostics and posterior summaries of the stochastic-volatility fit.