Skip to contents

Scores Gaussian predictive distributions against observations with proper scoring rules, interval scores and coverage across a grid of levels, and the probability integral transform. gp_holdout_scores() scores a fitted model on held-out data and gp_loo_scores() on its exact leave-one-out predictions.

Usage

gp_scores(
  observed,
  mean,
  sd,
  interval_levels = c(0.5, 0.8, 0.9, 0.95),
  reference_mean = NULL,
  reference_variance = NULL
)

gp_holdout_scores(
  model,
  x,
  y,
  interval_levels = c(0.5, 0.8, 0.9, 0.95),
  n_bins = 10L,
  seed = NULL,
  ...
)

gp_loo_scores(model, interval_levels = c(0.5, 0.8, 0.9, 0.95))

Arguments

observed

Numeric vector of observations.

mean, sd

Predictive means and standard deviations, one per observation; the standard deviations must be positive.

interval_levels

Levels of the central prediction intervals whose coverage and interval score are reported: the calibration curve.

reference_mean, reference_variance

Mean and variance of the reference Gaussian for the MSLL, usually those of the training responses. Without them the MSLL is NA.

model

A fitted model of any class.

x

Held-out inputs, as accepted by the model's stats::predict() method: time values for time-series models.

y

Held-out observations, one per row of x.

n_bins

Number of probability bins for the reliability of classification scores (gp_classification_scores()).

seed

Optional seed for the randomized PIT values of count scores (gp_count_scores()).

...

Further arguments to stats::predict(), such as observation_noise_variance for models fitted with observation-specific noise.

Value

An object of class gaussianprocesses_scores with

summary

Named numeric vector: n, log_score (mean log predictive density), crps (mean), smse, msll, pit_mean, and pit_variance (1/12 for uniform PIT values).

calibration

Data frame with one row per interval level: the level, empirical coverage, mean width, and mean interval_score.

pointwise

Data frame with one row per observation: observed, mean, sd, standardized_error, pit, log_score, and crps.

Details

With \(z_i = (y_i - \mu_i) / \sigma_i\), the scores are

log score

\(\log \varphi(z_i) - \log \sigma_i\), the log predictive density. Higher is better.

CRPS

\(\sigma_i [z_i (2 \Phi(z_i) - 1) + 2 \varphi(z_i) - 1 / \sqrt{\pi}]\), the continuous ranked probability score of Gneiting and Raftery (2007), in closed form. Lower is better; it is in the units of \(y\).

interval score

\((u - l) + \frac{2}{a}(l - y) \mathbf{1}\{y < l\} + \frac{2}{a}(y - u) \mathbf{1}\{y > u\}\) for the central interval \([l, u]\) of level \(1 - a\). Lower is better.

PIT

\(\Phi(z_i)\), uniform on \([0, 1]\) when the forecasts are calibrated.

SMSE

The mean squared error divided by the variance of the observations (Rasmussen and Williams, 2006, section 2.5).

MSLL

The mean of \(-\log p(y_i)\) minus the same under the reference Gaussian; negative values beat that reference.

The log score, CRPS, and interval score are proper: in expectation they are optimized by the true predictive distribution.

gp_holdout_scores() predicts with stats::predict(), so it accepts every model class, and scores the observation predictive distribution, including noise. Its MSLL reference is the mean and variance of the training responses. gp_loo_scores() scores the exact leave-one-out predictions of loo_gp() against the same reference.

For latent models (fit_latent_gp()), gp_holdout_scores() scores the predictive distribution of new responses: with a Gaussian likelihood as above, with a Bernoulli likelihood by gp_classification_scores(), whose log loss uses the predictive log probabilities of likelihood_predictive(), and with a Poisson likelihood by gp_count_scores(), with the exposure of the predictions.

Stability

Stable: from version 1.0.0 this interface changes incompatibly only in a major release, after a deprecation period. Results and options that concern an experimental model class, kernel, or argument follow that interface's tier. See gaussianprocesses-package for the policy.

References

Gneiting, T. and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.

Rasmussen, C. E. and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press. Section 2.5.

Examples

x <- seq(0, 6, length.out = 40)
y <- sin(x) + rnorm(40, sd = 0.1)
training <- seq(1, 40, by = 2)

model <- fit_gp(x[training], y[training], rbf_kernel(),
  noise_variance = 0.01)
scores <- gp_holdout_scores(model, x[-training], y[-training])
scores
#> Predictive scores for 20 observations
#>   log score (mean): 0.7386
#>   CRPS (mean): 0.06505
#>   SMSE: 0.02398
#>   MSLL: -1.821
#>   PIT mean and variance: 0.4906, 0.08746 (0.5 and 0.0833 if calibrated)
#>   coverage: 50% 60%, 80% 80%, 90% 85%, 95% 95%
scores$calibration
#>   level coverage     width interval_score
#> 1  0.50     0.60 0.1619853      0.2834063
#> 2  0.80     0.80 0.3077772      0.4082007
#> 3  0.90     0.85 0.3950278      0.4677165
#> 4  0.95     0.95 0.4707046      0.4877265

gp_loo_scores(model)$summary
#>            n    log_score         crps         smse         msll     pit_mean 
#>  20.00000000   0.66559017   0.07287271   0.02841719  -1.79113791   0.50395187 
#> pit_variance 
#>   0.08889417