About this Event
2000 University Drive, Boise, ID 83725
https://www.boisestate.edu/computing/ #University Events CalendarDissertation Information
Title: Prediction-Focused Cross-Validation for Block-Clustered Data
Program: Computing Ph.D. Data Science emphasis
Advisor: Dr. Grady Wright
Committee Members: Dr. Edoardo Serra, Dr. Michael Perlmutter
Abstract:
Cross-validation remains the default tool for estimating prediction error and selecting models for predicting observations in a data set. Block-clustered data, that is, observations nested within larger groups such as patients in clinics or students in schools, are common, and violate the independence that standard k-fold cross-validation assumes. When the goal is prediction to new clusters, k-fold cross-validation is optimistic due to the leakage of within-cluster information into the folds, from the within-cluster correlation of observations and the cluster structure being mistaken for signal. Theoretical work has demonstrated that prediction to observations within existing clusters does not optimistically bias the k-fold cross-validation prediction error, and that leave-one-cluster-out (LOCO) cross-validation targets the true generalization error.
The proposed thesis addresses open questions regarding when the LOCO – 10-fold gap is large enough to affect decisions, and when it might practically be ignored (e.g. if the within-cluster correlation is low, or if the overall sample size is sufficiently large). Proposed simulation studies address these questions for cluster-agnostic and cluster-aware machine learning methods for continuous and binary outcomes, covering parametric models with known corrections to non-linear ensemble methods without these corrections.
For the latter simulations, the generalization error will be built as ground truth from unseen clusters. How precisely the LOCO estimate itself can be reported in the parametric case is also studied. Controlled data-generating mechanisms spanning a range of intraclass correlations, cluster sizes, and total sample sizes are used, with methods judged on the estimated gap size and wall time. It is expected that these studies will provide practical guidance on when cluster-aware validation is necessary, relative to standard k-fold calculations, specific to the data in-hand and the intended prediction approaches.