Skip to contents

Computes analytical derivatives of the covariance matrix \(K(x, y; \theta)\) with respect to the logarithm of every kernel parameter, \(\partial K / \partial \log \theta_j\). These are the building blocks for gradient-based marginal-likelihood optimization, which works in log-parameter space.

Usage

kernel_gradient(kernel, x, y = NULL)

Arguments

kernel

A Gaussian-process kernel specification.

x

Numeric vector or matrix of observations.

y

Optional numeric vector or matrix. If NULL, derivatives of the covariance of x with itself are returned.

Value

A named list with one numeric matrix per flattened kernel parameter. Each matrix has nrow(x) rows and nrow(y) columns (or nrow(x) columns when y is NULL).

Details

Every kernel parameter is strictly positive, so derivatives are taken with respect to \(\eta_j = \log \theta_j\). The derivative with respect to the parameter on its natural scale is \(\partial K / \partial \theta_j = \theta_j^{-1} \partial K / \partial \eta_j\).

Elements are named and ordered exactly like kernel_parameters(kernel, flatten = TRUE), including indexed ARD paths such as kernel1.length_scale[2], so they can be matched to optimizer parameters by name. A scalar length scale shared by several input dimensions has a single derivative.

Leaf derivatives are closed-form expressions in the covariance matrix and the scaled input differences. Sums differentiate term by term, products use the product rule (without dividing by any factor, so exact zeros are safe), and scaled kernels use the chain rule. The derivations are in vignette("v07-kernel-derivatives", package = "gaussianprocesses").

Stability

Stable: from version 1.0.0 this interface changes incompatibly only in a major release, after a deprecation period. Results and options that concern an experimental model class, kernel, or argument follow that interface's tier. See gaussianprocesses-package for the policy.

See also

kernel_input_gradient() for derivatives with respect to the inputs.

Examples

kernel <- sum_kernel(
  rbf_kernel(variance = 1.5, length_scale = c(0.5, 2)),
  white_noise_kernel(variance = 0.1)
)
x <- cbind(c(0, 0.5, 1.5), c(1, 0, 2))

gradient <- kernel_gradient(kernel, x)
names(gradient)
#> [1] "kernel1.variance"        "kernel1.length_scale[1]"
#> [3] "kernel1.length_scale[2]" "kernel2.variance"       

# Natural-scale derivative with respect to the first ARD length scale.
gradient[["kernel1.length_scale[1]"]] / 0.5
#>           [,1]     [,2]      [,3]
#> [1,] 0.0000000 1.605784 0.2646987
#> [2,] 1.6057843 0.000000 0.9850200
#> [3,] 0.2646987 0.985020 0.0000000