Skip to contents

Observation models \(p(y_i \mid f_i)\) that link the latent Gaussian process \(f\) to the responses. Like kernel specifications, they are plain data with a class, a print method, and constrained parameters.

Usage

gaussian_likelihood(variance = 1)

bernoulli_likelihood(link = c("logit", "probit"))

poisson_likelihood(link = "log")

Arguments

variance

Positive noise variance of the Gaussian likelihood.

Link function: "logit" or "probit" for Bernoulli responses, and "log" for Poisson counts.

Value

An object of class gaussianprocesses_likelihood.

Details

gaussian_likelihood()

\(y_i \sim N(f_i, \sigma^2)\), the observation model of fit_gp() with noise_variance \(\sigma^2\).

bernoulli_likelihood()

\(\Pr(y_i = 1) = \sigma(f_i)\) with the logistic function, \(\log p = y f - \log(1 + e^f)\), or \(\Phi(f_i)\) with the probit link. Responses may be 0/1 numbers, logical values (TRUE is 1), or a two-level factor whose second level is 1, as in stats::glm(). Multi-class (softmax) likelihoods are not supported.

poisson_likelihood()

\(y_i \sim \mathrm{Poisson}(E_i e^{f_i})\) with exposure \(E_i\), so \(\log p = y (f + \log E) - E e^f - \log y!\). The exposure is data, supplied where the likelihood is used, not part of the specification. Counts more variable than a Poisson process with a smooth rate allows need negative-binomial or zero-inflated likelihoods, which are not yet supported; gp_count_scores() checks for such overdispersion.

Every likelihood provides, vectorized over observations, the log density and its first three derivatives in \(f_i\) (evaluate_likelihood()), derivatives in its parameters, the predictive distribution of a new response when \(f_*\) is Gaussian (likelihood_predictive()), and whether the log density is concave in \(f\). All three are log-concave, which makes Laplace approximations well behaved. Gaussian expectations without a closed form use Gauss-Hermite quadrature: nodes from the Golub-Welsch eigenvalue method, weights from Christoffel sums, cached by order.

The probit derivatives use the inverse Mills ratio \(r = \varphi(z) / \Phi(z)\) computed in log space, with \(z = (2y - 1) f\). For \(z < -5\), where \(r\) grows like \(-z\) and \(z + r\) cancels, they use the continued fraction of the Mills ratio instead. The logistic derivatives compute \(\pi\) and \(1 - \pi\) separately. All derivatives stay finite and accurate for large \(|f|\).

Latent models with any of these likelihoods are fitted with fit_latent_gp() and optimize_latent_gp(); fit_gp() remains the exact Gaussian model.

Stability

Experimental: this interface may change in a minor release, and every change is listed in NEWS. See gaussianprocesses-package for the policy.

References

Golub, G. H. and Welsch, J. H. (1969). Calculation of Gauss quadrature rules. Mathematics of Computation, 23(106), 221–230.

Rasmussen, C. E. and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press. Chapter 3.

Examples

gaussian_likelihood(variance = 0.1)
#> GaussianLikelihood(variance=0.1)
bernoulli_likelihood("probit")
#> BernoulliLikelihood(link=probit)
poisson_likelihood()
#> PoissonLikelihood(link=log)

evaluate_likelihood(bernoulli_likelihood(), c(TRUE, FALSE), f = c(1, 1))
#>   y f log_density      first     second      third
#> 1 1 1  -0.3132617  0.2689414 -0.1966119 0.09085775
#> 2 0 1  -1.3132617 -0.7310586 -0.1966119 0.09085775