Skip to contents

Fits a sparse approximation with fixed inducing locations: the Fully Independent Training Conditional (FITC) approximation or the collapsed variational free energy (VFE) of Titsias (2009). Exact GP regression remains the reference model; these provide a lower-cost path when the number of inducing points is much smaller than the number of observations.

Usage

fit_sparse_gp(
  x,
  y,
  kernel,
  noise_variance = 1e-06,
  mean = zero_mean(),
  n_inducing = NULL,
  inducing_points = NULL,
  selection = c("farthest", "quantile", "variance"),
  method = c("fitc", "vfe"),
  fitc_diagonal_floor = 1e-10,
  initial_jitter = 0,
  fallback_jitter = 1e-10,
  jitter_multiplier = 10,
  max_attempts = 8L,
  symmetry_tolerance = sqrt(.Machine$double.eps)
)

Arguments

x

Numeric vector or matrix of training inputs.

y

Numeric response vector.

kernel

Gaussian-process kernel specification.

noise_variance

Non-negative scalar or one non-negative value per training observation.

mean

Gaussian-process mean specification with fixed coefficients.

n_inducing

Optional number of inducing points selected from x.

inducing_points

Optional explicit inducing locations. If supplied, n_inducing is ignored.

selection

Inducing-point selection method used when explicit points are not supplied: "farthest", "quantile", or "variance" (greedy conditional-variance selection under kernel; see select_inducing_points()).

method

The approximation: "fitc" (the default) or "vfe". VFE needs positive noise variances.

fitc_diagonal_floor

Strictly positive numerical floor applied to the FITC diagonal term before inversion.

initial_jitter

Non-negative jitter tried first in the inducing-system Cholesky factorizations, relative to the scale of each factorized matrix (the mean of its diagonal).

fallback_jitter

Positive relative jitter tried after the first failed factorization.

jitter_multiplier

Multiplicative jitter escalation factor.

max_attempts

Maximum Cholesky attempts.

symmetry_tolerance

Relative tolerance for checking that covariance matrices are symmetric.

Value

An object of class gaussianprocesses_sparse_model.

Details

Let \(m\) be the number of inducing points and \(n\) the number of training observations. FITC replaces the full training covariance with a low-rank inducing covariance plus a diagonal correction. The dominant training cost is \(O(n m^2 + m^3)\) rather than the exact GP's \(O(n^3)\), with \(m \ll n\). The implementation stores an \(m \times n\) cross covariance rather than an \(n \times n\) training covariance.

FITC preserves each training point's prior marginal variance through the diagonal correction, unlike simpler deterministic training conditional approximations. It is still an approximation: inducing locations and their number can affect posterior means, variances, and the approximate marginal likelihood.

If the inducing covariance needs numerical jitter, the jittered matrix is used consistently in the low-rank term, the posterior, and the approximate log marginal likelihood. The inducing posterior is factorized in whitened form, whose eigenvalues are bounded below by one.

VFE

With \(Q_{ff} = K_{fu} K_{uu}^{-1} K_{uf}\) and noise \(\Sigma = \mathrm{diag}(\sigma_i^2)\), the collapsed variational bound of Titsias (2009) is $$\mathcal F = \log N(y \mid m(X), Q_{ff} + \Sigma) - \frac12 \sum_i \frac{k(x_i, x_i) - [Q_{ff}]_{ii}}{\sigma_i^2} \le \log p(y).$$ It is a lower bound on the exact log marginal likelihood, reported as log_marginal_likelihood, so maximizing it over the hyperparameters cannot overfit the way FITC can: FITC's diagonal correction \(\mathrm{diag}(K_{ff} - Q_{ff})\) can absorb noise, and FITC tends to underestimate the noise variance (Bauer, van der Wilk, and Rasmussen, 2016). VFE predicts with the deterministic training conditional, which has the same form as the FITC prediction with \(\Lambda = \Sigma\).

Both models report the trace term \(t = \sum_i (k(x_i, x_i) - [Q_{ff}]_{ii}) / \sigma_i^2\) in trace_term: the prior variance that the inducing points do not explain, in units of noise variance. A trace term much larger than 1 means that the inducing points do not cover the data well: VFE's bound is loose by at least \(t/2\) and FITC leans on its diagonal correction. More or better placed inducing points reduce it.

Stability

Experimental: this interface may change in a minor release, and every change is listed in NEWS. See gaussianprocesses-package for the policy.

References

Titsias, M. (2009). Variational learning of inducing variables in sparse Gaussian processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics, 567–574.

Bauer, M., van der Wilk, M., and Rasmussen, C. E. (2016). Understanding probabilistic sparse Gaussian process approximations. Advances in Neural Information Processing Systems, 29.

See also

gaussianprocesses_model for the structure, versioning, and persistence of fitted models.

Examples

x <- seq(-3, 3, length.out = 200)
y <- sin(x) + 0.1 * cos(7 * x)

model <- fit_sparse_gp(
  x,
  y,
  kernel = rbf_kernel(length_scale = 0.8),
  noise_variance = 0.01,
  n_inducing = 15
)
model
#> Sparse Gaussian-process regression model (FITC)
#>   observations: 200
#>   inducing points: 15
#>   inducing selection: farthest
#>   FITC floor applications: 0
#>   trace term tr(K - Q) / noise: 0.8491
#>   approximate log marginal likelihood: 195.5116