This Is Auburn

Show simple item record

Statistical Inference and Prediction Across Sparse Model Sets


Metadata FieldValueLanguage
dc.contributor.advisorMolinari, Roberto
dc.contributor.authorYavuz Ozdemir, Yagmur
dc.date.accessioned2026-08-06T22:13:41Z
dc.date.available2026-08-06T22:13:41Z
dc.date.issued2026-08-06
dc.identifier.urihttps://etd.auburn.edu/handle/10415/10619
dc.description.abstractHigh-dimensional scientific data often support many sparse models with similar predictive performance. In such settings, selecting one final model can hide substantial model uncertainty, especially when predictors are correlated, signals are weak, or several variable combinations provide comparable explanations. This dissertation develops inferential and predictive tools for working with sparse model sets produced by the Sparse Wrapper Algorithm (SWAG). The recovered SWAG library is interpreted as an empirical approximation to a Rashomon Variable Set: a collection of competitive variable subsets rather than a single selected support. The dissertation makes three contributions. First, it develops inference for sparse model sets. A global permutation test assesses whether the recovered library has more concentrated net-work structure than expected under a null relationship between the response and predictors. Conditional on global evidence of signal, model-specific p-values are aggregated across the library to summarize variable-level evidence. Simulations show approximately nominal type I error, increasing power with signal strength, and more precise variable recovery from George p-value aggregation than from LASSO in the settings considered. Second, the dissertation studies prediction over sparse model sets by stacking the predictions of SWAG library members. Simulations and public-data benchmarks show that stacked sparse model sets can improve held-out prediction while preserving an interpretable representation. Third, the framework is applied to biomedical data, including ADHD classification from resting-state functional connectivity and DNA methylation analysis of childhood sleep timing. Across these examples, sparse model sets provide a practical way to summarize prediction, stability, co-selection, direction of association, and biological follow-up in high-dimensional scientific problems.en_US
dc.subjectMathematics and Statisticsen_US
dc.titleStatistical Inference and Prediction Across Sparse Model Setsen_US
dc.typePhD Dissertationen_US
dc.embargo.statusNOT_EMBARGOEDen_US
dc.embargo.enddate2026-08-06en_US

Files in this item

Show simple item record