Speech-Based Alzheimer's Detection Under Controlled Acoustic Perturbations: A Baseline Study of Class Imbalance and Robustness
Authors: Aditya Amarjeet Singh
Affiliation: Allen High School; Allen, United States
Publication date: 2026-08-23
Publication pathway: Journal Publication
Collection: NSRI Student Research Journal
NSRI Student Research Journal
Online ISSN: 3143-5653
Volume: 1 Issue: 1 Pages/article: Article 0092
PDF: Open PDF/manuscript
Abstract
Background: Alzheimer's disease (AD) affects over 55 million people worldwide, and early detection through speech analysis offers a low-cost, non-invasive alternative to neuroimaging. However, clinical deployment of speech-based classifiers faces underexplored challenges: domain shifts across recording environments, limited labeled data at new sites, and a need for calibrated uncertainty estimates. Methods: We evaluated Random Forest and XGBoost classifiers on DementiaBank (295 Cookie Theft recordings from 292 unique speakers; 45 control, 250 AD). We extracted 1,589 features combining 53 hand-crafted acoustic descriptors (librosa, Praat) with wav2vec 2.0 and HuBERT embeddings, reduced to 128 dimensions via PCA. Robustness was assessed using 8 controlled acoustic simulations (microphone variation, background noise at 10-25 dB SNR, and reverberation at RT60 0.3-0.8 s) calibrated to published clinical measurements. Results: Both models achieved 81.4% raw accuracy (48/59; macro F1: 0.594; balanced accuracy: 0.581) on clean data, below the majority-class raw baseline of 84.7% but above it in balanced accuracy (0.500) and macro F1 (0.459). Under 7 of 8 controlled acoustic simulations, raw accuracy converged to 84.7%, matching the majority-class rate. A supplementary confusion-matrix analysis using a simplified verification model confirmed that this pattern reflects majority-class defaulting: control recall dropped to 0.000 under microphone and noise simulations. Random Forest achieved low calibration error (ECE = 0.042) on the 59-sample test set. In a single-draw few-shot experiment, XGBoost yielded 67.8% with K=10 examples per class. Conclusion: On clean data, acoustic features combined with self-supervised embeddings produce classifiers with meaningful balanced performance despite severe class imbalance (85% AD). The controlled acoustic simulation framework provides a reproducible evaluation methodology, though our results reveal that apparent robustness may reflect majority-class defaulting rather than genuine invariance, an important methodological caution for future work.
Keywords
Alzheimer's disease, speech analysis, domain robustness, acoustic features, calibration, machine learning, class imbalance
Citation
Publication Details
ISSN: Online ISSN: 3143-5653
License: Author-retained; open access display by NSRI unless a separate article license states otherwise.
Peer review status: NSRI uses editorial and scholarly review. When appropriate, manuscripts may undergo blinded review by reviewers with relevant subject knowledge.
AI disclosure: No AI disclosure is attached to this public record unless stated in the manuscript.
Conflict of interest statement: No conflict of interest statement is attached to this public record unless stated in the manuscript.
References
References are available in the manuscript PDF when provided.