Prediction and interpretation of factors associated with low birth weight using machine learning with MINSA data, Peru
DOI:
https://doi.org/10.51252/rcsi.v6i2.1388Keywords:
machine learning, low birth weight, CatBoost, explainability, health records, SHAPAbstract
Low birth weight is a significant public health issue in Peru. The objective was to develop, evaluate, and interpret machine learning models to predict it using national records from the Ministry of Health. A retrospective observational study was conducted involving 4,873,146 births registered between 2015 and 2025. Sixteen predictors—covering maternal, obstetric, neonatal, territorial, and healthcare-related factors—were used. CatBoost, LightGBM, and XGBoost were compared using Average Precision and ten stratified partitions from 2023. The selected model was trained on data from 2015–2022, calibrated using 2023 data, and temporally evaluated on data from 2024 and 2025. CatBoost achieved the best performance, with an Average Precision of 0.680 and an ROC-AUC of 0.910 in 2023; in 2024 and 2025, it achieved ROC-AUC values of 0.913 and 0.915, respectively. SHAP analysis identified gestational duration as the most influential predictor; its removal substantially reduced performance. The results demonstrate that MINSA records enable the development of interpretable and temporally stable models to support perinatal surveillance.
Downloads
References
Blencowe, H., Krasevec, J., de Onis, M., Black, R. E., An, X., Stevens, G. A., Borghi, E., Hayashi, C., Estevez, D., Cegolon, L., Shiekh, S., Ponce Hardy, V., Lawn, J. E., & Cousens, S. (2019). National, regional, and worldwide estimates of low birthweight in 2015, with trends from 2000: a systematic analysis. The Lancet Global Health, 7(7), e849–e860. https://doi.org/10.1016/S2214-109X(18)30565-5
Cajachagua-Torres, K. N., Quezada-Pinedo, H. G., Guzman-Vilca, W. C., Tarazona-Meza, C., Carrillo-Larco, R. M., & Huicho, L. (2024). Vulnerable newborn phenotypes in Peru: a population-based study of 3,841,531 births at national and subnational levels from 2012 to 2021. The Lancet Regional Health - Americas, 31, 100695. https://doi.org/10.1016/j.lana.2024.100695
Carrillo-Larco, R. M., Cajachagua-Torres, K. N., Guzman-Vilca, W. C., Quezada-Pinedo, H. G., Tarazona-Meza, C., & Huicho, L. (2021). National and subnational trends of birthweight in Peru: Pooled analysis of 2,927,761 births between 2012 and 2019 from the national birth registry. The Lancet Regional Health - Americas, 1, 100017. https://doi.org/10.1016/j.lana.2021.100017
Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. https://doi.org/10.1145/2939672.2939785
Demšar, J. (2006). Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research, 7(1), 1–30. http://jmlr.org/papers/v7/demsar06a.html
Dominguez Miranda, S. A., & Rodriguez Aguilar, R. (2024). Machine learning models in health prevention and promotion and labor productivity: A co-word analysis. Iberoamerican Journal of Science Measurement and Communication, 4(1), 1–16. https://doi.org/10.47909/ijsmc.85
García, S., Fernández, A., Luengo, J., & Herrera, F. (2010). Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power. Information Sciences, 180(10), 2044–2064. https://doi.org/10.1016/j.ins.2009.12.010
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Conference: Advances in Neural Information Processing Systems 30 (NIPS 2017). https://papers.nips.cc/paper/2017/hash/6449f44a102fde848669bdd9eb6b76fa-Abstract.html
Lundberg, S., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. http://arxiv.org/abs/1705.07874
Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. Proceedings of the 22nd International Conference on Machine Learning - ICML ’05, 625–632. https://doi.org/10.1145/1102351.1102430
Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., & Gulin, A. (2019). CatBoost: unbiased boosting with categorical features. http://arxiv.org/abs/1706.09516
Ranjbar, A., Montazeri, F., Farashah, M. V., Mehrnoush, V., Darsareh, F., & Roozbeh, N. (2023). Machine learning-based approach for predicting low birth weight. BMC Pregnancy and Childbirth, 23(1), 803. https://doi.org/10.1186/s12884-023-06128-w
Saito, T., & Rehmsmeier, M. (2015). The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432
Sultana, N., Afia, Z., Sifat, I. K., Zoha, S., Jisa, T. A., & Kibria, M. K. (2025). Machine learning based prediction of low birth weight and its associated risk factors: Insights from the Bangladesh Demographic and Health Survey 2022. PLOS Global Public Health, 5(9), e0005187. https://doi.org/10.1371/journal.pgph.0005187
Valles-Coral, M., Pinedo, L., Navarro-Cabrera, J. R., Valverde-Iparraguirre, J., Injante, R., Saavedra, S., Quintanilla-Morales, L. K., & Almeida-Espinosa, A. (2025). Skills in searching for and using scientific information among physicians in the context of evidence-based practice: A descriptive study. Iberoamerican Journal of Science Measurement and Communication, 5(1), 1–9. https://doi.org/10.47909/ijsmc.181
Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W. (2019). Calibration: the Achilles heel of predictive analytics. BMC Medicine, 17(1), 230. https://doi.org/10.1186/s12916-019-1466-7
Victor, A., Almeida, F., Xavier, S. P., & Rondó, P. H. C. (2025). Predicting low birth weight risks in pregnant women in Brazil using machine learning algorithms: data from the Araraquara cohort study. BMC Pregnancy and Childbirth, 25(1), 320. https://doi.org/10.1186/s12884-025-07351-3
WHO, & UNICEF. (2023). Nutrition and Food Safety. Word Health Organization. https://www.who.int/teams/nutrition-and-food-safety/monitoring-nutritional-status-and-food-safety-and-events/joint-low-birthweight-estimates
Wu, X., Zhao, Q., Gao, Y., Zhang, Y., Xu, L., Cong, X., Sun, N., Shi, F., & Wang, S. (2025). Interpretable machine learning model for predicting low birth weight in singleton pregnancies: a retrospective cohort study. BMC Pregnancy and Childbirth, 25(1), 1159. https://doi.org/10.1186/s12884-025-08318-0
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Daniel Andrade-Girón, Elsa Oscuvicla-Tapia, Abrahan Neri-Ayala, Américo Peña, Miguel Aguilar-Luna-Victoria, Edgardo Cuevas-Huari

This work is licensed under a Creative Commons Attribution 4.0 International License.
The authors retain their rights:
a. The authors retain their trademark and patent rights, as well as any process or procedure described in the article.
b. The authors retain the right to share, copy, distribute, execute and publicly communicate the article published in the Revista Científica de Sistemas e Informática (RCSI) (for example, place it in an institutional repository or publish it in a book), with an acknowledgment of its initial publication in the RCSI.
c. Authors retain the right to make a subsequent publication of their work, to use the article or any part of it (for example: a compilation of their works, notes for conferences, thesis, or for a book), provided that they indicate the source of publication (authors of the work, journal, volume, number and date).





