Simulation and Benchmark Protocol
Source:vignettes/articles/benchmark-protocol.Rmd
benchmark-protocol.RmdThis research note is maintained in
inst/notes/benchmark-protocol.md and is installed with the
package;
system.file("notes", "benchmark-protocol.md", package = "gaussianprocesses")
returns its location.
Goals
The benchmark suite separates four different questions:
- Statistical recovery — how close is the posterior mean to the latent simulated function?
- Approximation accuracy — how close is FITC to the exact GP posterior?
- Numerical stability — was jitter required and how well-conditioned was the stabilized covariance?
- Performance — how much elapsed time and retained model memory were used?
These are deliberately reported separately. A faster model is not labelled more accurate, and a numerically stable fit is not automatically a better statistical model.
Standard scenarios
smooth_1d
A one-dimensional RBF process with moderate homoscedastic noise.
Purpose: - baseline correctness; - interpolation accuracy; - exact-versus-FITC comparison.
near_singular
A one-dimensional RBF process with two almost coincident inputs and extremely small observation noise.
Purpose: - Cholesky stabilization; - condition-number diagnostics; - sensitivity to near-duplicate observations.
Reproducibility
Every public simulation function accepts an explicit integer seed. Seeded simulation restores the caller’s previous global RNG state after drawing, so running a benchmark does not silently alter later random calculations.
The default suite uses deterministic scenario-specific offsets from the base seed.
Timing
Timing uses repeated system.time() calls and reports the
median elapsed time. This reduces sensitivity to one unusually slow run
but is still not a microbenchmarking framework.
The benchmark functions are intended for local scientific comparison. They should not be used as strict CI performance gates because shared-runner load, BLAS implementations, CPU frequency, and operating system scheduling can change timings materially.
Memory
The suite reports retained R object size through
object.size().
This is not peak resident memory. It answers a narrower question: how much R object memory remains associated with the fitted model after fitting.
External implementations
Comparisons with other implementations (DiceKriging, kernlab, nlme,
and gplite) are part of the validation references in
inst/validation/references.R, which the test suite runs and
the validation article tabulates. They fix the hyperparameters where
practical, so that they compare inference rather than optimizers, and
compare optimizers separately by the likelihood they reach.