Scores predicted class probabilities against binary responses with proper
scoring rules (log loss and Brier score), discrimination (ROC AUC), and
calibration (reliability diagram data and the expected calibration
error). gp_holdout_scores() returns these scores for latent models with
a Bernoulli likelihood.
Arguments
- observed
Binary responses: 0/1 numbers, logical values, or a two-level factor whose second level is the event coded 1, as in
bernoulli_likelihood().- probability
Forecast probabilities \(p_i = \Pr(y_i = 1)\), one per response.
- n_bins
Number of equal-width probability bins of the reliability diagram and the expected calibration error.
- log_density
Optional log predictive probabilities of the observed responses, for the log loss. They default to \(\log p_i\) or \(\log(1 - p_i)\);
gp_holdout_scores()supplies them from the predictive distribution, which keeps their accuracy when \(p_i\) is close to 0 or 1.
Value
An object of class gaussianprocesses_classification_scores with
summaryNamed numeric vector:
n,log_loss,brier_score,auc,ece,base_rate(the mean response), andmean_probability.reliabilityData frame with one row per non-empty bin:
bin,lower,upper,n,mean_probability,observed_frequency, andstandard_error.pointwiseData frame with
observed(0/1),probability,log_density, andsquared_error.
Details
- log loss
\(-\frac1n \sum_i \log p(y_i)\), the negative mean log predictive probability. Proper; lower is better; \(\log 2 \approx 0.693\) for \(p_i = 1/2\).
- Brier score
\(\frac1n \sum_i (p_i - y_i)^2\). Proper; lower is better; 0.25 for \(p_i = 1/2\).
- ROC AUC
The probability that a random event receives a higher forecast than a random non-event, with ties counted as one half: the Mann-Whitney statistic from mid-ranks. 0.5 for uninformative forecasts;
NAwithout both classes. It measures discrimination only and ignores calibration.- reliability
The forecasts are grouped into
n_binsequal-width bins of \([0, 1]\). For calibrated forecasts the observed frequency of each bin is close to its mean probability, within aboutstandard_error\(= \sqrt{\bar p (1 - \bar p) / n_b}\).- ECE
The expected calibration error \(\sum_b \frac{n_b}{n} |\bar p_b - \bar y_b|\). It depends on the binning, and even calibrated forecasts have a positive ECE of the order of the bin standard errors.
Stability
Experimental: this interface may change in a minor release, and every change is listed in NEWS. See gaussianprocesses-package for the policy.
References
Gneiting, T. and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.
See also
gp_scores() for Gaussian predictions, predict_latent_gp().
Examples
set.seed(1)
x <- runif(80, 0, 6)
y <- rbinom(80, 1, plogis(2 * sin(x)))
training <- 1:50
model <- fit_latent_gp(x[training], y[training], rbf_kernel(4, 1),
bernoulli_likelihood())
scores <- gp_holdout_scores(model, x[-training], y[-training], n_bins = 5)
scores
#> Classification scores for 30 responses (base rate 0.6333)
#> log loss: 0.4022
#> Brier score: 0.1186
#> ROC AUC: 0.8995
#> expected calibration error: 0.09342 (5 occupied bins)
scores$reliability
#> bin lower upper n mean_probability observed_frequency standard_error
#> 1 1 0.0 0.2 7 0.1627811 0.1428571 0.1395316
#> 2 2 0.2 0.4 4 0.2186391 0.2500000 0.2066616
#> 3 3 0.4 0.6 4 0.5587810 0.7500000 0.2482664
#> 4 4 0.6 0.8 3 0.7120363 1.0000000 0.2614323
#> 5 5 0.8 1.0 12 0.8409133 0.9166667 0.1055849