Score Comparability for Translated and Adapted Exams
An adaptation-comparability workflow for small, lower-scoring
and unbalanced language groups where standard differential item
functioning (DIF) tools (Magis, Beland, Tuerlinckx and De Boeck, 2010,
) break down. Calibrates the Rasch model in
each language group, links the groups robustly through the densest
cluster of items rather than assuming DIF cancels out (on anchor
selection see Kopf, Zeileis and Strobl, 2015,
), detects small-sample DIF with an
empirical-Bayes spike-and-slab model and local false discovery rates
(Efron, 2004, ), quantifies whether
item-level DIF accumulates into different pass rates, explains DIF by
item features to give translators actionable guidance, and drafts a
comparability report.