Fits a sparse approximation with fixed inducing locations: the Fully Independent Training Conditional (FITC) approximation or the collapsed variational free energy (VFE) of Titsias (2009). Exact GP regression remains the reference model; these provide a lower-cost path when the number of inducing points is much smaller than the number of observations.
Usage
fit_sparse_gp(
x,
y,
kernel,
noise_variance = 1e-06,
mean = zero_mean(),
n_inducing = NULL,
inducing_points = NULL,
selection = c("farthest", "quantile", "variance"),
method = c("fitc", "vfe"),
fitc_diagonal_floor = 1e-10,
initial_jitter = 0,
fallback_jitter = 1e-10,
jitter_multiplier = 10,
max_attempts = 8L,
symmetry_tolerance = sqrt(.Machine$double.eps)
)Arguments
- x
Numeric vector or matrix of training inputs.
- y
Numeric response vector.
- kernel
Gaussian-process kernel specification.
- noise_variance
Non-negative scalar or one non-negative value per training observation.
- mean
Gaussian-process mean specification with fixed coefficients.
- n_inducing
Optional number of inducing points selected from
x.- inducing_points
Optional explicit inducing locations. If supplied,
n_inducingis ignored.- selection
Inducing-point selection method used when explicit points are not supplied:
"farthest","quantile", or"variance"(greedy conditional-variance selection underkernel; seeselect_inducing_points()).- method
The approximation:
"fitc"(the default) or"vfe". VFE needs positive noise variances.- fitc_diagonal_floor
Strictly positive numerical floor applied to the FITC diagonal term before inversion.
- initial_jitter
Non-negative jitter tried first in the inducing-system Cholesky factorizations, relative to the scale of each factorized matrix (the mean of its diagonal).
- fallback_jitter
Positive relative jitter tried after the first failed factorization.
- jitter_multiplier
Multiplicative jitter escalation factor.
- max_attempts
Maximum Cholesky attempts.
- symmetry_tolerance
Relative tolerance for checking that covariance matrices are symmetric.
Details
Let \(m\) be the number of inducing points and \(n\) the number of training observations. FITC replaces the full training covariance with a low-rank inducing covariance plus a diagonal correction. The dominant training cost is \(O(n m^2 + m^3)\) rather than the exact GP's \(O(n^3)\), with \(m \ll n\). The implementation stores an \(m \times n\) cross covariance rather than an \(n \times n\) training covariance.
FITC preserves each training point's prior marginal variance through the diagonal correction, unlike simpler deterministic training conditional approximations. It is still an approximation: inducing locations and their number can affect posterior means, variances, and the approximate marginal likelihood.
If the inducing covariance needs numerical jitter, the jittered matrix is used consistently in the low-rank term, the posterior, and the approximate log marginal likelihood. The inducing posterior is factorized in whitened form, whose eigenvalues are bounded below by one.
VFE
With \(Q_{ff} = K_{fu} K_{uu}^{-1} K_{uf}\) and noise
\(\Sigma = \mathrm{diag}(\sigma_i^2)\), the collapsed variational
bound of Titsias (2009) is
$$\mathcal F = \log N(y \mid m(X), Q_{ff} + \Sigma) -
\frac12 \sum_i \frac{k(x_i, x_i) - [Q_{ff}]_{ii}}{\sigma_i^2}
\le \log p(y).$$
It is a lower bound on the exact log marginal likelihood, reported as
log_marginal_likelihood, so maximizing it over the hyperparameters
cannot overfit the way FITC can: FITC's diagonal correction
\(\mathrm{diag}(K_{ff} - Q_{ff})\) can absorb noise, and FITC tends to
underestimate the noise variance (Bauer, van der Wilk, and Rasmussen,
2016). VFE predicts with the deterministic training conditional, which
has the same form as the FITC prediction with \(\Lambda = \Sigma\).
Both models report the trace term
\(t = \sum_i (k(x_i, x_i) - [Q_{ff}]_{ii}) / \sigma_i^2\) in
trace_term: the prior variance that the inducing points do not explain,
in units of noise variance. A trace term much larger than 1 means that
the inducing points do not cover the data well: VFE's bound is loose by at
least \(t/2\) and FITC leans on its diagonal correction. More or better
placed inducing points reduce it.
Stability
Experimental: this interface may change in a minor release, and every change is listed in NEWS. See gaussianprocesses-package for the policy.
References
Titsias, M. (2009). Variational learning of inducing variables in sparse Gaussian processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics, 567–574.
Bauer, M., van der Wilk, M., and Rasmussen, C. E. (2016). Understanding probabilistic sparse Gaussian process approximations. Advances in Neural Information Processing Systems, 29.
See also
gaussianprocesses_model for the structure, versioning, and persistence of fitted models.
Examples
x <- seq(-3, 3, length.out = 200)
y <- sin(x) + 0.1 * cos(7 * x)
model <- fit_sparse_gp(
x,
y,
kernel = rbf_kernel(length_scale = 0.8),
noise_variance = 0.01,
n_inducing = 15
)
model
#> Sparse Gaussian-process regression model (FITC)
#> observations: 200
#> inducing points: 15
#> inducing selection: farthest
#> FITC floor applications: 0
#> trace term tr(K - Q) / noise: 0.8491
#> approximate log marginal likelihood: 195.5116