Scores Gaussian predictive distributions against observations with proper
scoring rules, interval scores and coverage across a grid of levels, and
the probability integral transform. gp_holdout_scores() scores a fitted
model on held-out data and gp_loo_scores() on its exact leave-one-out
predictions.
Arguments
- observed
Numeric vector of observations.
- mean, sd
Predictive means and standard deviations, one per observation; the standard deviations must be positive.
- interval_levels
Levels of the central prediction intervals whose coverage and interval score are reported: the calibration curve.
- reference_mean, reference_variance
Mean and variance of the reference Gaussian for the MSLL, usually those of the training responses. Without them the MSLL is
NA.- model
A fitted model of any class.
- x
Held-out inputs, as accepted by the model's
stats::predict()method: time values for time-series models.- y
Held-out observations, one per row of
x.- n_bins
Number of probability bins for the reliability of classification scores (
gp_classification_scores()).- seed
Optional seed for the randomized PIT values of count scores (
gp_count_scores()).- ...
Further arguments to
stats::predict(), such asobservation_noise_variancefor models fitted with observation-specific noise.
Value
An object of class gaussianprocesses_scores with
summaryNamed numeric vector:
n,log_score(mean log predictive density),crps(mean),smse,msll,pit_mean, andpit_variance(1/12 for uniform PIT values).calibrationData frame with one row per interval level: the
level, empiricalcoverage, meanwidth, and meaninterval_score.pointwiseData frame with one row per observation:
observed,mean,sd,standardized_error,pit,log_score, andcrps.
Details
With \(z_i = (y_i - \mu_i) / \sigma_i\), the scores are
- log score
\(\log \varphi(z_i) - \log \sigma_i\), the log predictive density. Higher is better.
- CRPS
\(\sigma_i [z_i (2 \Phi(z_i) - 1) + 2 \varphi(z_i) - 1 / \sqrt{\pi}]\), the continuous ranked probability score of Gneiting and Raftery (2007), in closed form. Lower is better; it is in the units of \(y\).
- interval score
\((u - l) + \frac{2}{a}(l - y) \mathbf{1}\{y < l\} + \frac{2}{a}(y - u) \mathbf{1}\{y > u\}\) for the central interval \([l, u]\) of level \(1 - a\). Lower is better.
- PIT
\(\Phi(z_i)\), uniform on \([0, 1]\) when the forecasts are calibrated.
- SMSE
The mean squared error divided by the variance of the observations (Rasmussen and Williams, 2006, section 2.5).
- MSLL
The mean of \(-\log p(y_i)\) minus the same under the reference Gaussian; negative values beat that reference.
The log score, CRPS, and interval score are proper: in expectation they are optimized by the true predictive distribution.
gp_holdout_scores() predicts with stats::predict(), so it accepts
every model class, and scores the observation predictive distribution,
including noise. Its MSLL reference is the mean and variance of the
training responses. gp_loo_scores() scores the exact leave-one-out
predictions of loo_gp() against the same reference.
For latent models (fit_latent_gp()), gp_holdout_scores() scores the
predictive distribution of new responses: with a Gaussian likelihood as
above, with a Bernoulli likelihood by gp_classification_scores(), whose
log loss uses the predictive log probabilities of
likelihood_predictive(), and with a Poisson likelihood by
gp_count_scores(), with the exposure of the predictions.
Stability
Stable: from version 1.0.0 this interface changes incompatibly only in a major release, after a deprecation period. Results and options that concern an experimental model class, kernel, or argument follow that interface's tier. See gaussianprocesses-package for the policy.
References
Gneiting, T. and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.
Rasmussen, C. E. and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press. Section 2.5.
Examples
x <- seq(0, 6, length.out = 40)
y <- sin(x) + rnorm(40, sd = 0.1)
training <- seq(1, 40, by = 2)
model <- fit_gp(x[training], y[training], rbf_kernel(),
noise_variance = 0.01)
scores <- gp_holdout_scores(model, x[-training], y[-training])
scores
#> Predictive scores for 20 observations
#> log score (mean): 0.7386
#> CRPS (mean): 0.06505
#> SMSE: 0.02398
#> MSLL: -1.821
#> PIT mean and variance: 0.4906, 0.08746 (0.5 and 0.0833 if calibrated)
#> coverage: 50% 60%, 80% 80%, 90% 85%, 95% 95%
scores$calibration
#> level coverage width interval_score
#> 1 0.50 0.60 0.1619853 0.2834063
#> 2 0.80 0.80 0.3077772 0.4082007
#> 3 0.90 0.85 0.3950278 0.4677165
#> 4 0.95 0.95 0.4707046 0.4877265
gp_loo_scores(model)$summary
#> n log_score crps smse msll pit_mean
#> 20.00000000 0.66559017 0.07287271 0.02841719 -1.79113791 0.50395187
#> pit_variance
#> 0.08889417