Abstract

Although K‐fold cross‐validation (CV) is widely used for model evaluation and selection, there has been limited understanding of how to perform CV for non‐iid data, including those from sampling designs with unequal selection probabilities. We introduce CV methodology that is appropriate for design‐based inference from complex survey sampling designs. For such data, we claim that we will tend to make better inferences when we choose the folds and compute the test errors in ways that account for the survey design features such as stratification and clustering. Our mathematical arguments are supported with simulations, and our methods are illustrated on real survey data.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call