← All publications
2026 arXiv:2602.03957

Temporal Validation Changes the Apparent Public-Health Utility of Under-Five Mortality Prediction in Bangladesh: A Four-Round DHS Machine-Learning Study

Md Muhtasim Munif Fahim , M. Monimul Huq , M. Sabiruzzaman , Md. Rezaul Karim

BMC Public Health

Abstract

Background: Bangladesh has reduced under-five mortality substantially, but preventable deaths remain unevenly distributed across households and divisions. Prediction models based on Demographic and Health Survey data could help planners prioritise follow-up, referral, and resource allocation, but only if their reported performance reflects future public-health use. We quantified how validation design changes apparent under-five mortality prediction performance in Bangladesh. Methods: We analysed four Bangladesh Demographic and Health Survey rounds (2011, 2014, 2017, and 2022; 33 962 children; 1 290 under-five deaths). The same 26-feature preprocessing pipeline and three model classes were evaluated under four validation regimes: pooled random 80/20, matched-size pooled random 80/20, 2022-only random 80/20, and cross-survey temporal validation trained on 2011+2014, validated and calibrated on 2017, and tested on held-out 2022. The neural model was a 32-unit ELU multilayer perceptron selected by genetic-algorithm neural architecture search. AUROC was estimated with 2 000 bootstrap resamples; screening utility used sensitivity, positive predictive value, and number needed to screen (NNS) at fixed capacity. Results: Validation regime changed the public-health interpretation of model performance more than model class. For the NAS-derived multilayer perceptron, AUROC ranged from 0.669 under 2022-only random validation to 0.775 under pooled random validation, with a temporal estimate of 0.730. At the top-10 percent temporal screening threshold, NAS predictions identified 152 of 355 observed deaths in 2022 (sensitivity 42.8 percent, positive predictive value 13.2 percent, NNS 7.6). Across validation designs, the same model implied NNS values from 5.6 to 11.0, changing the expected follow-up workload and deaths identified. Conclusions: In this four-round Bangladesh benchmark, validation-regime choice changed the screening workload and apparent policy value of under-five mortality prediction more than architecture choice. Cross-round temporal validation gives planners a more defensible basis for estimating community-health-worker follow-up, referral demand, and budget scenarios than random-split AUROC alone. DHS-based child-mortality prediction studies should therefore report capacity-based metrics such as sensitivity, positive predictive value, and NNS before such models are used for programme planning or public-policy decisions.

Machine LearningPublic HealthBangladeshUnder-Five MortalityTemporal ValidationBDHSNeural Architecture SearchFairness

BibTeX

@article{fahim2026temporal,
  title   = {Temporal Validation Changes the Apparent Public-Health Utility of Under-Five Mortality Prediction in Bangladesh: A Four-Round DHS Machine-Learning Study},
  author  = {Md Muhtasim Munif Fahim and M. Monimul Huq and M. Sabiruzzaman and Md. Rezaul Karim},
  year    = {2026},
  journal = {BMC Public Health},
  eprint  = {2602.03957},
  archivePrefix = {arXiv},
  url     = {https://arxiv.org/abs/2602.03957},
}