Ongoing surveillance of item parameter drift for continuous
testing programs and pre-equated item banks. Estimates item difficulty in
rolling calibration windows, runs sequential cumulative sum (CUSUM)
detection (Page, 1954,
Not "did it drift?" but "when, how, and what should we do?"
Most drift checks compare two calibrations at equating time. Continuous testing programs and pre-equated banks need ongoing surveillance. driftwatch borrows sequential change detection from statistical process control.
library(driftwatch)
sim <- dw_simulate(seed = 1) # or your responses + bank
est <- dw_estimate(sim$responses, sim$bank) # item x window estimates
tu <- dw_tune(est, target = 0.01) # threshold for 1% false alarms/item
mon <- dw_monitor(est, h = tu$h) # CUSUM + change points + type
mon
dw_actions(mon, anchors = my_anchor_ids) # recommendations + audit log
dw_impact(form_ids, sim$bank, current_b, flagged) # score / pass-rate impact
From CRAN (once released):
install.packages("driftwatch")
Development version from GitHub:
install.packages("pak")
pak::pak("edidatasolutions/driftwatch")
z = (b_t - b_bank) / sqrt(SE_t^2 + SE_bank^2).z (reference value k, threshold h).dw_tune() regenerates responses
under no drift for the program's actual items, windows, sample sizes and
examinees, then reruns estimation and CUSUM. This matters: each item's
bank error is shared by every window and accumulates in the CUSUM. The
design-based threshold (h ≈ 9.4) is far above the iid-normal one (≈ 6.9),
and only the former hits the false-alarm target.undetermined rather than guessed.inst/validation/known_truth.R)300 items, 40 windows, ~80 responses per item per window. 10% of items drift gradually (0.02–0.06 logits/window) and 5% jump (0.4–1.0 logits).
| continuous (driftwatch) | two-point (first vs last 5 windows) | |
|---|---|---|
| false alarms, stable items (target 1%) | 1.3% | — |
| abrupt detected | 100%, median 3.4 windows after onset | 74%, at the end |
| gradual detected | 89%, median 12 windows after onset | 78%, at the end |
| abrupt onset error (windows) | 0.6 | — |
Drift-type classification:
| data used | abrupt: typed / accuracy | gradual: typed / accuracy |
|---|---|---|
| alarm + 3 windows | 79% / 98% | 47% / 65% |
| all 40 windows | 93% / 100% | 74% / 88% |
Gradual drift is hard to type soon after an alarm, because a short ramp looks like a step. Reclassify as windows accumulate.
Score impact (60-item form with 12 drifted items, cut at theta = 0.5): keeping banked values mis-states the pass rate by -1.3 points; recalibrating flagged items cuts that to -0.2 points, and removing them to -0.15.
Done: dw_simulate, dw_estimate, dw_tune, dw_monitor, dw_impact,
dw_actions, dw_twopoint. Rasch difficulty only; examinee ability treated
as known. Next: 2PL discrimination drift, ability uncertainty, anchor-set
re-linking after removals, and a per-item run-length (ARL) view.