Derivatives of a covariance matrix with respect to log hyperparameters
Source:R/kernel-gradients.R
kernel_gradient.RdComputes analytical derivatives of the covariance matrix \(K(x, y; \theta)\) with respect to the logarithm of every kernel parameter, \(\partial K / \partial \log \theta_j\). These are the building blocks for gradient-based marginal-likelihood optimization, which works in log-parameter space.
Value
A named list with one numeric matrix per flattened kernel
parameter. Each matrix has nrow(x) rows and nrow(y) columns (or
nrow(x) columns when y is NULL).
Details
Every kernel parameter is strictly positive, so derivatives are taken with respect to \(\eta_j = \log \theta_j\). The derivative with respect to the parameter on its natural scale is \(\partial K / \partial \theta_j = \theta_j^{-1} \partial K / \partial \eta_j\).
Elements are named and ordered exactly like
kernel_parameters(kernel, flatten = TRUE), including indexed ARD paths
such as kernel1.length_scale[2], so they can be matched to optimizer
parameters by name. A scalar length scale shared by several input
dimensions has a single derivative.
Leaf derivatives are closed-form expressions in the covariance matrix and
the scaled input differences. Sums differentiate term by term, products use
the product rule (without dividing by any factor, so exact zeros are safe),
and scaled kernels use the chain rule. The derivations are in
vignette("v07-kernel-derivatives", package = "gaussianprocesses").
Stability
Stable: from version 1.0.0 this interface changes incompatibly only in a major release, after a deprecation period. Results and options that concern an experimental model class, kernel, or argument follow that interface's tier. See gaussianprocesses-package for the policy.
See also
kernel_input_gradient() for derivatives with respect to the
inputs.
Examples
kernel <- sum_kernel(
rbf_kernel(variance = 1.5, length_scale = c(0.5, 2)),
white_noise_kernel(variance = 0.1)
)
x <- cbind(c(0, 0.5, 1.5), c(1, 0, 2))
gradient <- kernel_gradient(kernel, x)
names(gradient)
#> [1] "kernel1.variance" "kernel1.length_scale[1]"
#> [3] "kernel1.length_scale[2]" "kernel2.variance"
# Natural-scale derivative with respect to the first ARD length scale.
gradient[["kernel1.length_scale[1]"]] / 0.5
#> [,1] [,2] [,3]
#> [1,] 0.0000000 1.605784 0.2646987
#> [2,] 1.6057843 0.000000 0.9850200
#> [3,] 0.2646987 0.985020 0.0000000