Back to papers
arxiv6.0 / 10

Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening

Wensi Zhang, Tomas Teijeiro, Jérôme Thevenot, David Atienza

Abstract

Cough acoustics are promising for non-invasive tuberculosis (TB) screening, yet whether machine learning (ML) models capture disease-related acoustics or artifacts of data collection remains unresolved. We evaluated the cross-dataset generalizability of classical ML and deep learning (DL) cough-based TB classifiers across three independent datasets. Despite moderate within-dataset performance (ROC-AUC up to $0.755 \pm 0.056$), both pipelines fail to generalize, with external performance frequently below 0.6, indicating a possible limitation of the data. We further observed audio representations are organized by recording device and dataset rather than TB status, predicted TB probability tracks country-level prevalence in CODA, and device mismatch degrades transfer while device-diverse training improves it. Additionally, a clinical-variable baseline generalizes more consistently (ROC-AUC $0.655 - 0.711$), indicating acquisition-specific variability is a stronger driver of poor generalizability than population shift. High within-dataset performance is not enough. External validation is essential before cough-based TB models are clinically ready.

Research area

evaluationgeneralization sciencerobustness to domain shifts
Published
26 Aug 2026
Source
arxiv
Org
Ecole Polytechnique Federale de Lausanne
View paper
Sign in to read and join the discussion.