Contributing
Source:CONTRIBUTING.md
Contributions should preserve the mathematical transparency and numerical reliability of the project.
Workflow
- Create a focused branch from
main. - Keep each merge request scoped to one issue or tightly related set of changes.
- Add or update tests for every behavioral change.
- Document mathematical assumptions and parameter conventions.
- Avoid unnecessary dependencies.
- Run local tests and package checks before opening a merge request.
R style
- Use clear, descriptive names.
- Prefer small functions with explicit contracts.
- Validate public function inputs.
- Document exported functions with
roxygen2. - Avoid explicit matrix inversion in numerical algorithms where a stable solve is available.
- Make tolerances and jitter policies explicit.
Testing
Run:
testthat::test_local()For package-level validation:
devtools::check()The suite runs in about a minute. Keep each test file under 10 seconds: reduce the work, or mark a longer check with skip_on_cran() and say why. Statistical tests use fixed seeds and bounds of a few Monte Carlo standard errors, with n_statistical_draws draws (helper-sampling.R): 20 000 under testthat::test_local(), devtools::test(), and CI, and 2000 under R CMD check, as on CRAN, where the whole check should take about a minute. The optimization smoke tests, the benchmark-suite test, and the validation references are skipped on CRAN.
Line coverage is measured with covr, which is not a package dependency:
The script prints coverage per file and fails if any file in R/ is below 85%.
Numerical changes
Changes affecting linear algebra, optimization, covariance construction, or predictive uncertainty should include tests covering both ordinary and numerically difficult cases.
Validation references
Every inference method is checked against an independent computation: another implementation, a brute-force computation, or a closed form. The references are in inst/validation/references.R; tests/testthat/test-validation-references.R runs them, and the validation article tabulates them.
- A new exported function must be added to
validation_methodsunder the method it belongs to, or tovalidation_exclusionswith the reason it performs no inference; a test fails otherwise. - A new method needs at least one reference: a call to
validation_reference()whoserun()returns the discrepancy, with a tolerance chosen by the rules at the top of the file. A reference that uses another package names it inpackage, which must be in Suggests, and records the version used to set the tolerance. - Keep each reference to a second or less; the article runs all of them when it is built.
Adding a kernel
Every leaf kernel is defined by one entry in the internal kernel registry, R/kernel-registry.R. The entry is the only place that knows the kernel’s type, label, parameters, covariance formula, diagonal, and derivatives. Evaluation, kernel_diagonal(), kernel_gradient(), kernel_input_gradient(), predict_gradient_gp(), parameter paths and updates, printing, sparse FITC models, and optimize_gp() all read it. Sum, product, scale, and changepoint kernels, and kernels restricted to some input columns with select_dimensions(), are structural and are not registry entries.
A new kernel needs:
-
One registry entry, registered with
.register_kernel(). The comment at the top ofR/kernel-registry.Rlists its fields:- parameters, each declared with
.kernel_parameter(), which sets the default, the shape ("scalar", or"ard"for one value per input dimension), and the constraint (.positive_constraint(),.real_constraint(), or.bounded_constraint(lower, upper)); -
evaluate()anddiagonal(); -
gradient(), which returns the derivative with respect to each parameter’s optimizer coordinate: the log for positive parameters, the value itself for real parameters, and the logit for bounded parameters. Multiply a derivative on the natural scale by the constraint’sjacobian()to convert it. -
smoothness, the number of mean-square derivatives of the process:0, a positive integer, orInf; -
input_gradient()andinput_hessian(), the derivatives with respect to the inputs, ∂k/∂x_d and ∂²k/∂x_d∂y_e, as lists over input dimensions (seeR/kernel-input-derivatives.R). They are required whensmoothnessis at least 1 andNULLotherwise; -
input_parameter_gradient(), the derivatives of those two lists with respect to each parameter’s optimizer coordinate, which derivative observations need for the likelihood gradient. It follows the same rules; -
state_space, a function of the parameters that returns the kernel’s stochastic differential equation on one input (see.matern_state_space()inR/gp-state-space.Rfor the fields), orNULLwhen the kernel has no exact finite-dimensional state. Kalman inference (fit_time_series_gp(method = "state_space")) uses it; a new form needs tests of the Lyapunov equation, of the kernel reconstruction k(τ) = H exp(Fτ) P∞ Hᵀ, and against the exact GP, as intests/testthat/test-state-space.R.
- parameters, each declared with
-
A constructor,
name_kernel(), that calls.new_leaf_kernel(). Its argument defaults must equal the registry defaults; a test checks this. A covariance-matrix functionkernel_name()that calls.evaluate_leaf_parameters()is optional. -
A positive-semidefiniteness test: the Gram matrix on inputs with repeated rows has no eigenvalue below about
-1e-10times its scale. -
A diagonal-consistency test:
kernel_diagonal()equalsdiag(evaluate_kernel()). - A finite-difference test of the hyperparameter derivatives, taken in each parameter’s optimizer coordinate, and, for a differentiable kernel, of the input derivatives and their hyperparameter derivatives.
-
Documentation: the constructor’s help page, the formula and its derivatives in the kernel vignettes, a
_pkgdown.ymlreference entry, and aNEWS.mditem.
expect_kernel_contract() in tests/testthat/test-kernel-registry.R performs checks 3–5, using expect_input_derivative_contract() and expect_input_parameter_gradient_contract() from tests/testthat/helper-input-derivatives.R for the input derivatives. Every built-in kernel is checked with it, and the test-only kernel in that file shows a complete entry with positive, ARD, real, and bounded parameters.
Stability and versioning
Every export is stable, experimental, or deprecated. R/stability.R lists the experimental and deprecated ones; every other export is stable. Each export’s help page states its tier in a “Stability” section, from @template stable or @template experimental (the templates are in man-roxygen/), and the reference index in _pkgdown.yml groups the topics by tier. tests/testthat/test-api-stability.R checks that the three agree, so a new export needs a tier, a template, and a place in the right group of the index. The policy users see is on the package help page, ?gaussianprocesses.
Version numbers follow semantic versioning from 1.0.0:
- Major releases may change stable interfaces incompatibly. Removing a deprecated name is such a change.
- Minor releases add features, deprecate stable names, and may change experimental interfaces. Every change to an experimental interface is listed in NEWS.
- Patch releases only fix defects, compatibly.
Before 1.0.0, minor releases may also change stable interfaces, with deprecation where possible.
A stable function, argument, or result field is replaced in four steps:
- Add the replacement and keep the old name working with a warning:
.Deprecated()for a function, a warning of classdeprecatedWarningwhen an old argument name is supplied (and an error when both names are), or a documented copy for a result field. - List the old name in
R/stability.R, on a?gaussianprocesses-deprecatedhelp page, ininst/notes/api-audit.md, and in NEWS. - Release it in at least one minor release with the warning.
- Remove it in the next major release.
Version 1.0.0 has no deprecated names. The names renamed after 0.1.0 were removed in 1.0.0 without a release of warnings, an exception decided in #40 because 1.0.0 is the first release after 0.1.0; R/stability.R lists them, and a test checks that they do not return.
An experimental interface becomes stable in a minor release, with a NEWS entry; a stable one never becomes experimental.
Fitted models. Models record their schema version (R/model-objects.R). Schema 4 is the format of version 1.0:
- Every 1.x release must read schemas 0 to 4 and give the same results. The fixtures in
tests/testthat/fixtureshold models saved by the 0.1.0 release and schema-4 models written before 1.0, and the tests check both. Never regenerate them. - A new schema needs an upgrade path from schema 4 in
.upgrade_model_from_old_schema(), a schema-history entry, and a test that upgrades the schema-4 fixture.
Extension policy: the kernel registry
The kernel registry (R/kernel-registry.R) stays internal for 1.0: there is no exported way to register a user-defined kernel. The decision was made in #39, for these reasons:
-
Optimizer coordinates. An entry’s
gradient()returns derivatives in each parameter’s optimizer coordinate: log, identity, or logit, depending on its constraint. A public entry format would freeze that convention and the constraint system with it. - Internal conventions. Entries also supply input derivatives, which derivative observations and inducing-point optimization need, and the state-space form of #37. Both slots were added after the registry was designed, and further slots may follow; a public format would make every addition a breaking change for extension authors.
- Validation. A user-defined kernel would need validation that the package cannot fully provide: positive semidefiniteness can be checked on test inputs but not proved, and wrong gradients silently degrade optimization rather than fail.
- Composition. Composite kernels already cover most custom covariances: sums, products, scales, column selection, changepoints, and spectral mixtures.
The registry can become public later as an experimental interface, with an exported validator (positive semidefiniteness on test inputs, finite-difference checks of the gradients and input derivatives, and the diagonal against evaluate_kernel()) and an extension vignette. That needs a separate issue. Until then, new kernels are added to the package itself, as described under “Adding a kernel”.
Compiled code
The package stays pure R for 1.0: no compiled code, and no Imports beyond stats. The decision was made in #39, from these measurements:
-
Linear algebra. Most of the work is already compiled. In a Laplace fit of a Bernoulli model with 2000 observations, 78% of the time is in
chol(), which calls LAPACK. A faster BLAS speeds this up; compiled package code would not. -
The Kalman filter is the only hot loop that runs in R. It costs about 17 microseconds per observation, so one log marginal likelihood for 100 000 observations takes 1.7 s (
inst/benchmarks/state-space-scaling.R). That is fast enough for fitting and forecasting. Optimizing hyperparameters on series that long is slow, because finite-difference gradients need several filter passes per step. Analytical sensitivity recursions would remove most of those passes before compiled code is needed. - Other loops iterate over a few dozen steps, such as Newton and heteroscedastic iterations and greedy inducing-point selection, and each step is dominated by matrix operations.
Compiled code should be reconsidered if:
- profiling a supported workflow shows that R-level loops, not BLAS or LAPACK calls, take most of its time;
- and the workflow matters to users at that size.
The first candidate would be the Kalman recursions of R/gp-state-space.R. Adding compiled code would bring a toolchain requirement for source installs, LinkingTo, and compiled-code checks in CI, so it needs its own issue.
Continuous integration
The pipeline is defined in .gitlab-ci.yml, and every job runs a script from ci/. None of it is needed for local development; the commands in Testing are enough.
| Job | Runs when | What it does |
|---|---|---|
r-cmd-check |
package code, tests, man/, or vignettes change |
Runs the test suite with a JUnit report, then R CMD build and R CMD check --as-cran --no-manual --no-tests; fails on any test failure, ERROR, or WARNING |
ci-scripts |
ci/ or the pipeline changes |
ShellCheck and the release guard tests |
coverage |
scheduled pipelines, or started by hand | Line coverage with covr; fails below 85% for any file in R/. Ordinary merge-request pipelines skip it. |
docs |
a merge request changes documentation | Builds the pkgdown site as a downloadable preview |
pages |
the default branch changes documentation or code | Builds the site and deploys it to GitLab Pages |
release |
a manually started pipeline sets RELEASE_TAG
|
Publishes a release; see RELEASING.md |
Documentation-only changes (README, NEWS, research notes, examples, site configuration) only rebuild the site; they never run R CMD check.
The pipeline is designed for a small compute budget:
- One pipeline per change: merge-request pipelines, plus branch pipelines only for branches without an open merge request. Tags do not start pipelines.
- Jobs are skipped when none of the paths they depend on changed.
- Images are small and pinned by digest. R packages are prebuilt binaries from the image’s dated Posit Package Manager snapshot. Only the packages a job needs are installed, including the suggested packages that the validation references compare against. The R library is cached per image snapshot.
- Tests run once:
R CMD checkskips them after the separate test step. - OpenBLAS and OpenMP use one thread. The image’s multithreaded OpenBLAS spreads the small matrices of these jobs over every runner CPU, which made the test suite ten times slower.
- Benchmark suites in
inst/benchmarksare never run. - Jobs are interruptible, so a newer pipeline cancels redundant runs, and each job is capped at 30 minutes.
- Jobs run on a self-hosted Docker runner; shared runners are disabled for the project. Jobs are untagged, so they also run unchanged on shared runners.
Reproducing jobs locally
Each job can be reproduced with Docker from the repository root, using the same image as CI. For r-cmd-check:
docker run --rm -v "$PWD:/work" -w /work \
-e R_LIBS=/work/.cache/R/library \
-e _R_CHECK_FORCE_SUGGESTS_=false -e _R_CHECK_CRAN_INCOMING_=false \
-e _R_CHECK_SYSTEM_CLOCK_=false \
-e OPENBLAS_NUM_THREADS=1 -e OMP_NUM_THREADS=1 \
rocker/r-ver:4.6.1 sh -c '
mkdir -p "$R_LIBS" &&
apt-get update -qq &&
apt-get install -y -qq --no-install-recommends pandoc qpdf &&
Rscript ci/install-r-deps.R xml2 &&
Rscript ci/run-tests.R &&
sh ci/r-cmd-check.sh --no-tests'For the documentation site, install pandoc libwebpmux3, then run Rscript ci/install-r-deps.R pkgdown and sh ci/build-site.sh. The release guards are tested with sh ci/tests/release-test.sh in any image that has git.