Predicts Rasch item difficulty from item features (text
embeddings from any model, template family, content metadata) by ridge
regression (Hoerl and Kennard, 1970,
Calibrate AI-generated items before you have the pretest seats to do it the old way.
Automatic item generation and LLM-assisted writing produce items faster than pretesting can calibrate them. coldstart predicts each new item's difficulty from its features, uses the prediction as a robust prior, finds the families where prediction fails, and tells you how many responses each item needs.
library(coldstart)
sim <- cs_simulate(seed = 1) # legacy + new items with features
it <- sim$items; tr <- it$set == "train"
pr <- cs_predictor(it$b_legacy[tr], sim$features[tr, ], it$family[tr])
pred <- predict(pr, sim$features[!tr, ], it$family[!tr]) # mean + honest SD
cs_plan(pred, target_sd = 0.3) # responses needed per item
resp <- cs_responses(sim, n_per_item = 25) # or your pretest data
cal <- cs_calibrate(resp, pred) # t-prior Bayes vs baseline
chk <- cs_check(cal, setNames(it$family[!tr], it$item[!tr]))
cal <- cs_calibrate(resp, cs_distrust(pred, chk)) # drop priors that failed
From CRAN (once released):
install.packages("coldstart")
Development version from GitHub:
install.packages("pak")
pak::pak("edidatasolutions/coldstart")
Features can be anything numeric: embeddings from any text model, cognitive attribute codes, content metadata. The package does not call a model itself.
cs_distrust() withdraws priors from failing families.inst/validation/known_truth.R)Predictive SDs are honest: stated 0.54 vs actual RMSE 0.53 (seen families), 0.70 vs 0.64 (unseen family); 90% intervals cover 90.8% and 90.4%.
RMSE of difficulty, seen families:
| responses per item | baseline | predicted prior |
|---|---|---|
| 15 | 0.72 | 0.40 |
| 25 | 0.53 | 0.35 |
| 50 | 0.36 | 0.30 |
| 100 | 0.27 | 0.23 |
With the prior, 25 responses do what 50 do without it.
Drifted ("rogue") template family (+1.2 logits vs its history): the prior
alone hurts (0.66 vs 0.61 at n = 25). cs_check flags the family in 80% of
replications at n = 25 and 100% at n ≥ 50, with 0.2 false family flags per
replication. After cs_distrust(), RMSE is back at baseline (0.57 at n = 25).
Planner: a target posterior SD of 0.30 needs a median of 40 responses per item with the prior vs 57 without; achieved SD 0.302, RMSE 0.300.
Done: cs_simulate, cs_responses, cs_predictor (+predict),
cs_calibrate, cs_plan, cs_check, cs_distrust. Next: 2PL
(discrimination priors), sequential updating as responses arrive, ability
uncertainty for pretest examinees (currently treated as known from
operational scoring), and non-linear predictors.