Tools for detecting and modeling underdispersion in count data (conditional variance below the conditional mean), the case the Poisson and negative binomial defaults cannot represent. Provides a screening diagnostic that benchmarks at-risk dispersion against a zero-truncated Poisson, regression-adjusted tests of equidispersion, and a dispersion profile that compares the variance-to-mean curves of competing families against the data; the continuous parameter binomial (CPB) and generalized event count (Katz) regressions with zero-truncated, hurdle, and zero-inflated forms and high-dimensional fixed effects with a split-panel jackknife bias correction; matched Poisson, negative binomial, COM-Poisson (rate- and mean-parameterized), generalized Poisson, gamma-count, and double Poisson regressions through the same interface, with frequency weights, offsets, and analytic, robust, and cluster-robust standard errors; bootstrap and profile-likelihood inference; proper scoring rules, rootograms, PIT histograms, and simulation methods; and quantities of interest including predicted distributions, the implied ceiling, rate ratios, and first differences with an extensive/intensive decomposition. The likelihoods are implemented in C++.
Tools for detecting and modeling underdispersion in count data: the case where the conditional variance falls below the conditional mean, so counts cluster more tightly around their expectation than a Poisson allows.
Underdispersion is common in bounded counts (portfolios of statuses that are
filled and vacated over time, events with regular timing, quotas) but poorly
served by standard software: the negative binomial cannot represent a variance
below the mean and collapses onto the Poisson. underdisp provides the missing
pieces, from a screen that tells you whether an outcome is underdispersed to a
family of estimators, tests, and diagnostics that treat every model on the same
footing.
ud_screen() reports within-unit dispersion, a
zero-truncated-Poisson at-risk benchmark (which separates genuine
underdispersion from the artifact of conditioning on positive counts), and an
over-conditioning guard, with calibrated or parametric-bootstrap thresholds.dispersion_test() tests any fitted family's
dispersion parameter against the Poisson on the fitted design (fixed effects
included), with boundary-aware p-values; dispersion_profile() plots the
conditional variance-to-mean ratio against the fitted mean next to the curve
each family implies, so the mechanism (a hard ceiling, regular event timing,
a soft tail) can be read off the data.cpb() fits King's continuous parameter
binomial and its zero-truncated variant, with an interpretable
observation-specific bound; cpb_fe() absorbs high-dimensional unit fixed
effects by a concentrated likelihood that scales to thousands of units.gec() fits the generalized event count
(Katz) model, whose single dispersion parameter is estimated freely and spans
under-, equi-, and overdispersion; gec_fe() adds concentrated fixed effects.hurdle_cpb()/zi_cpb() and hurdle_gec()/zi_gec()
pair the underdispersed intensities with participation or structural-zero
processes; zi_test() runs the boundary-corrected zero-inflation test.count_reg(), hurdle_count(), and zi_count() fit
the Poisson, negative binomial, COM-Poisson (rate- or mean-parameterized),
generalized Poisson, gamma-count, and double Poisson through the same
interface, with fixed effects, offsets, frequency weights, and analytic,
robust, or cluster-robust standard errors, so compare_models() and
compare_dispersion() adjudicate the whole family on one footing
(information criteria plus proper scores, in sample, held out, or by
cross-validation via score()/cv_score()).bias_correct = "jackknife" removes the
1/T incidental-parameters bias of the fixed-effects dispersion estimate
(split-panel jackknife, Dhaene & Jochmans 2015), behind a validity gate that
refuses the correction, with an informative warning, on panels that violate
the method's time-homogeneity requirement.simulate() methods for every model class, so any fit plugs into DHARMa's
simulated-residual diagnostics via DHARMa::createDHARMa().confint() for every
class; parallel bootstraps via cores =; broom, modelsummary, and
texreg support throughout.# from CRAN
install.packages("underdisp")
# development version
# install.packages("remotes")
remotes::install_github("bagozzib/underdisp")
library(underdisp)
# an underdispersed count (var/mean ~ 0.5)
n <- 400; x <- rnorm(n)
N <- pmax(round(exp(1.6 + 0.5 * x) / 0.5), 1)
d <- data.frame(y = rbinom(n, N, 0.5), x = x)
ud_screen(y ~ x, data = d) # screen
fit <- cpb(y ~ x, data = d, truncated = FALSE, se = "none")
summary(fit)
dispersion_test(fit, cores = 2) # boundary null: parametric-bootstrap p-value
implied_ceiling(fit, newdata = data.frame(x = 0))
compare_dispersion(y ~ x, data = d)$table
dispersion_profile(cpb = fit, # which mechanism fits?
gammacount = count_reg(y ~ x, d, family = "gammacount"),
negbin = count_reg(y ~ x, d, family = "negbin"))
The bundled peacekeeping panel demonstrates the package's central move: a
count that looks overdispersed in the pooled margin but is underdispersed
within countries at risk. See vignette("underdisp") for the full walk-through,
including the two-part models, the bias-corrected fixed effects, the wider
family of underdispersed distributions, and the DHARMa workflow. The package
website is at https://bagozzib.github.io/underdisp/.
Dhaene, Geert, and Koen Jochmans. 2015. "Split-Panel Jackknife Estimation of Fixed-Effect Models." The Review of Economic Studies 82(3): 991–1030.
Efron, Bradley. 1986. "Double Exponential Families and Their Use in Generalized Linear Regression." Journal of the American Statistical Association 81(395): 709–721.
Huang, Alan. 2017. "Mean-Parametrized Conway–Maxwell–Poisson Regression Models for Dispersed Counts." Statistical Modelling 17(6): 359–380.
King, Gary. 1989. "Variance Specification in Event Count Models." American Journal of Political Science 33(3): 762–784.
Winkelmann, Rainer. 1995. "Duration Dependence and Dispersion in Count-Data Models." Journal of Business & Economic Statistics 13(4): 467–474.
Winkelmann, Rainer, Curtis S. Signorino, and Gary King. 1995. "A Correction for an Underdispersed Event Count Probability Distribution." Political Analysis 5: 215–228.