Anytime-Valid Sequential Benchmarking of Algorithms

Sequential, anytime-valid comparison of algorithms on paired losses with a practical-equivalence margin. Confidence sequences for the mean paired difference, superiority/equivalence/continue decisions, cost accounting and auditable trajectories.


seqbench seqbench logo

R-CMD-check Lifecycle: experimental

Anytime-valid sequential benchmarking of algorithms.

seqbench answers one question: when can I stop comparing algorithm A with algorithm B without invalidating the inference? It treats a benchmark as a sequential experiment on paired losses (same instance, same seed) and maintains a confidence sequence for the mean paired difference that stays valid under optional stopping, repeated inspection and adaptive budgets.

Three declarations are possible, all relative to a pre-registered margin of practical relevance δ:

Declaration Condition on the current confidence sequence C_t
A is relevantly better upper limit of C_t below −δ
B is relevantly better lower limit of C_t above +δ
Practically equivalent C_t inside [−δ, +δ]
Continue none of the above and budget remains
Inconclusive none of the above and the budget is exhausted (or C_t is empty)

Status

Version 0.1.0, first release; the API is experimental and may change in minor releases. Implemented: five boundaries (betting, empirical Bernstein, Hoeffding, Bernstein with a declared variance bound, and an invalid fixed-sample negative control), instance-level updating with seeds as clusters, the three-way decision rule, budget accounting, and print, summary, tidy, glance and autoplot methods.

library(seqbench)

design <- comparison_design(alpha = 0.05, margin = 0.02, bounds = c(0, 1),
                            boundary = "betting", budget = 500)
state  <- initialize_comparison(design)

# paired losses: one row per (instance, seed); losses within `bounds`
losses <- data.frame(instance = 1:40, seed = 1,
                     loss_a = runif(40, 0.1, 0.4), loss_b = runif(40, 0.3, 0.6))
state  <- update_comparison(state, losses)

stopping_decision(state)     # "A", "B", "equivalent", "continue" or "inconclusive"
comparison_report(state)     # estimand, estimate, C_t, assumptions, cost, seeds, versions
tidy(state)                  # trajectory, one row per instance
ggplot2::autoplot(state)     # C_t against t with the [-margin, margin] band

Instances are the sequential unit; several seeds of the same instance are averaged into one observation. Losses outside bounds, missing values and instances that reappear after their losses were seen are errors, never silently converted.

Installation

install.packages("seqbench")            # CRAN, once released

# development version
# install.packages("pak")
pak::pak("castlaboratory/seqbench")

Related software

seqbench is not a general confidence-sequence library. For that see confseq (Python), safestats and avlm (R). The CRAN package seqcomp compares probabilistic forecasts sequentially; seqbench compares algorithms on paired losses with an equivalence margin and explicit cost accounting.

License

GPL (>= 3). © CAST Lab, Universidade Federal de Pernambuco.

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("seqbench")

0.1.0 by André Leite, 9 hours ago


https://github.com/castlaboratory/seqbench


Report a bug at https://github.com/castlaboratory/seqbench/issues


Browse source code at https://github.com/cran/seqbench


Authors: André Leite [aut, cre] , Raydonal Ospina [aut] , Cristiano Ferraz [aut]


Documentation:   PDF Manual  


GPL (>= 3) license


Imports cli, generics, ggplot2, rlang, stats, tibble, utils

Suggests knitr, rmarkdown, testthat


See at CRAN