Skip to contents

Scores predicted class probabilities against binary responses with proper scoring rules (log loss and Brier score), discrimination (ROC AUC), and calibration (reliability diagram data and the expected calibration error). gp_holdout_scores() returns these scores for latent models with a Bernoulli likelihood.

Usage

gp_classification_scores(
  observed,
  probability,
  n_bins = 10L,
  log_density = NULL
)

Arguments

observed

Binary responses: 0/1 numbers, logical values, or a two-level factor whose second level is the event coded 1, as in bernoulli_likelihood().

probability

Forecast probabilities \(p_i = \Pr(y_i = 1)\), one per response.

n_bins

Number of equal-width probability bins of the reliability diagram and the expected calibration error.

log_density

Optional log predictive probabilities of the observed responses, for the log loss. They default to \(\log p_i\) or \(\log(1 - p_i)\); gp_holdout_scores() supplies them from the predictive distribution, which keeps their accuracy when \(p_i\) is close to 0 or 1.

Value

An object of class gaussianprocesses_classification_scores with

summary

Named numeric vector: n, log_loss, brier_score, auc, ece, base_rate (the mean response), and mean_probability.

reliability

Data frame with one row per non-empty bin: bin, lower, upper, n, mean_probability, observed_frequency, and standard_error.

pointwise

Data frame with observed (0/1), probability, log_density, and squared_error.

Details

log loss

\(-\frac1n \sum_i \log p(y_i)\), the negative mean log predictive probability. Proper; lower is better; \(\log 2 \approx 0.693\) for \(p_i = 1/2\).

Brier score

\(\frac1n \sum_i (p_i - y_i)^2\). Proper; lower is better; 0.25 for \(p_i = 1/2\).

ROC AUC

The probability that a random event receives a higher forecast than a random non-event, with ties counted as one half: the Mann-Whitney statistic from mid-ranks. 0.5 for uninformative forecasts; NA without both classes. It measures discrimination only and ignores calibration.

reliability

The forecasts are grouped into n_bins equal-width bins of \([0, 1]\). For calibrated forecasts the observed frequency of each bin is close to its mean probability, within about standard_error \(= \sqrt{\bar p (1 - \bar p) / n_b}\).

ECE

The expected calibration error \(\sum_b \frac{n_b}{n} |\bar p_b - \bar y_b|\). It depends on the binning, and even calibrated forecasts have a positive ECE of the order of the bin standard errors.

Stability

Experimental: this interface may change in a minor release, and every change is listed in NEWS. See gaussianprocesses-package for the policy.

References

Gneiting, T. and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.

See also

gp_scores() for Gaussian predictions, predict_latent_gp().

Examples

set.seed(1)
x <- runif(80, 0, 6)
y <- rbinom(80, 1, plogis(2 * sin(x)))
training <- 1:50

model <- fit_latent_gp(x[training], y[training], rbf_kernel(4, 1),
  bernoulli_likelihood())
scores <- gp_holdout_scores(model, x[-training], y[-training], n_bins = 5)
scores
#> Classification scores for 30 responses (base rate 0.6333)
#>   log loss: 0.4022
#>   Brier score: 0.1186
#>   ROC AUC: 0.8995
#>   expected calibration error: 0.09342 (5 occupied bins)
scores$reliability
#>   bin lower upper  n mean_probability observed_frequency standard_error
#> 1   1   0.0   0.2  7        0.1627811          0.1428571      0.1395316
#> 2   2   0.2   0.4  4        0.2186391          0.2500000      0.2066616
#> 3   3   0.4   0.6  4        0.5587810          0.7500000      0.2482664
#> 4   4   0.6   0.8  3        0.7120363          1.0000000      0.2614323
#> 5   5   0.8   1.0 12        0.8409133          0.9166667      0.1055849