Predicción e interpretación de factores asociados al bajo peso al nacer mediante aprendizaje automático con datos del MINSA, Perú
DOI:
https://doi.org/10.51252/rcsi.v6i2.1388Palabras clave:
aprendizaje automático, bajo peso al nacer, CatBoost, explicabilidad, registros de salud, SHAPResumen
El bajo peso al nacer constituye un problema relevante de salud pública en el Perú. El objetivo fue desarrollar, evaluar e interpretar modelos de aprendizaje automático para predecirlo mediante registros nacionales del Ministerio de Salud. Se realizó un estudio observacional y retrospectivo con 4 873 146 nacimientos registrados entre 2015 y 2025. Se utilizaron 16 predictores maternos, obstétricos, neonatales, territoriales y asistenciales. CatBoost, LightGBM y XGBoost fueron comparados mediante Average Precision y diez particiones estratificadas de 2023. El modelo seleccionado fue entrenado con datos de 2015-2022, calibrado con 2023 y evaluado temporalmente en 2024 y 2025. CatBoost obtuvo el mejor desempeño, con Average Precision de 0,680 y ROC-AUC de 0,910 en 2023; en 2024 y 2025 alcanzó ROC-AUC de 0,913 y 0,915, respectivamente. SHAP identificó la duración gestacional como el predictor más influyente. Su eliminación redujo sustancialmente el rendimiento. Los resultados muestran que los registros del MINSA permiten desarrollar modelos interpretables y temporalmente estables para apoyar la vigilancia perinatal.
Descargas
Citas
Blencowe, H., Krasevec, J., de Onis, M., Black, R. E., An, X., Stevens, G. A., Borghi, E., Hayashi, C., Estevez, D., Cegolon, L., Shiekh, S., Ponce Hardy, V., Lawn, J. E., & Cousens, S. (2019). National, regional, and worldwide estimates of low birthweight in 2015, with trends from 2000: a systematic analysis. The Lancet Global Health, 7(7), e849–e860. https://doi.org/10.1016/S2214-109X(18)30565-5
Cajachagua-Torres, K. N., Quezada-Pinedo, H. G., Guzman-Vilca, W. C., Tarazona-Meza, C., Carrillo-Larco, R. M., & Huicho, L. (2024). Vulnerable newborn phenotypes in Peru: a population-based study of 3,841,531 births at national and subnational levels from 2012 to 2021. The Lancet Regional Health - Americas, 31, 100695. https://doi.org/10.1016/j.lana.2024.100695
Carrillo-Larco, R. M., Cajachagua-Torres, K. N., Guzman-Vilca, W. C., Quezada-Pinedo, H. G., Tarazona-Meza, C., & Huicho, L. (2021). National and subnational trends of birthweight in Peru: Pooled analysis of 2,927,761 births between 2012 and 2019 from the national birth registry. The Lancet Regional Health - Americas, 1, 100017. https://doi.org/10.1016/j.lana.2021.100017
Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. https://doi.org/10.1145/2939672.2939785
Demšar, J. (2006). Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research, 7(1), 1–30. http://jmlr.org/papers/v7/demsar06a.html
Dominguez Miranda, S. A., & Rodriguez Aguilar, R. (2024). Machine learning models in health prevention and promotion and labor productivity: A co-word analysis. Iberoamerican Journal of Science Measurement and Communication, 4(1), 1–16. https://doi.org/10.47909/ijsmc.85
García, S., Fernández, A., Luengo, J., & Herrera, F. (2010). Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power. Information Sciences, 180(10), 2044–2064. https://doi.org/10.1016/j.ins.2009.12.010
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Conference: Advances in Neural Information Processing Systems 30 (NIPS 2017). https://papers.nips.cc/paper/2017/hash/6449f44a102fde848669bdd9eb6b76fa-Abstract.html
Lundberg, S., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. http://arxiv.org/abs/1705.07874
Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. Proceedings of the 22nd International Conference on Machine Learning - ICML ’05, 625–632. https://doi.org/10.1145/1102351.1102430
Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., & Gulin, A. (2019). CatBoost: unbiased boosting with categorical features. http://arxiv.org/abs/1706.09516
Ranjbar, A., Montazeri, F., Farashah, M. V., Mehrnoush, V., Darsareh, F., & Roozbeh, N. (2023). Machine learning-based approach for predicting low birth weight. BMC Pregnancy and Childbirth, 23(1), 803. https://doi.org/10.1186/s12884-023-06128-w
Saito, T., & Rehmsmeier, M. (2015). The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432
Sultana, N., Afia, Z., Sifat, I. K., Zoha, S., Jisa, T. A., & Kibria, M. K. (2025). Machine learning based prediction of low birth weight and its associated risk factors: Insights from the Bangladesh Demographic and Health Survey 2022. PLOS Global Public Health, 5(9), e0005187. https://doi.org/10.1371/journal.pgph.0005187
Valles-Coral, M., Pinedo, L., Navarro-Cabrera, J. R., Valverde-Iparraguirre, J., Injante, R., Saavedra, S., Quintanilla-Morales, L. K., & Almeida-Espinosa, A. (2025). Skills in searching for and using scientific information among physicians in the context of evidence-based practice: A descriptive study. Iberoamerican Journal of Science Measurement and Communication, 5(1), 1–9. https://doi.org/10.47909/ijsmc.181
Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W. (2019). Calibration: the Achilles heel of predictive analytics. BMC Medicine, 17(1), 230. https://doi.org/10.1186/s12916-019-1466-7
Victor, A., Almeida, F., Xavier, S. P., & Rondó, P. H. C. (2025). Predicting low birth weight risks in pregnant women in Brazil using machine learning algorithms: data from the Araraquara cohort study. BMC Pregnancy and Childbirth, 25(1), 320. https://doi.org/10.1186/s12884-025-07351-3
WHO, & UNICEF. (2023). Nutrition and Food Safety. Word Health Organization. https://www.who.int/teams/nutrition-and-food-safety/monitoring-nutritional-status-and-food-safety-and-events/joint-low-birthweight-estimates
Wu, X., Zhao, Q., Gao, Y., Zhang, Y., Xu, L., Cong, X., Sun, N., Shi, F., & Wang, S. (2025). Interpretable machine learning model for predicting low birth weight in singleton pregnancies: a retrospective cohort study. BMC Pregnancy and Childbirth, 25(1), 1159. https://doi.org/10.1186/s12884-025-08318-0
Descargas
Publicado
Cómo citar
Número
Sección
Licencia
Derechos de autor 2026 Daniel Andrade-Girón, Elsa Oscuvicla-Tapia, Abrahan Neri-Ayala, Américo Peña, Miguel Aguilar-Luna-Victoria, Edgardo Cuevas-Huari

Esta obra está bajo una licencia internacional Creative Commons Atribución 4.0.
Los autores retienen sus derechos:
a. Los autores retienen sus derechos de marca y patente, y tambien sobre cualquier proceso o procedimiento descrito en el artículo.
b. Los autores retienen el derecho de compartir, copiar, distribuir, ejecutar y comunicar públicamente el articulo publicado en la Revista Científica de Sistemas e Informática (RCSI) (por ejemplo, colocarlo en un repositorio institucional o publicarlo en un libro), con un reconocimiento de su publicación inicial en la RCSI.
c. Los autores retienen el derecho a hacer una posterior publicación de su trabajo, de utilizar el artículo o cualquier parte de aquel (por ejemplo: una compilación de sus trabajos, notas para conferencias, tesis, o para un libro), siempre que indiquen la fuente de publicación (autores del trabajo, revista, volumen, número y fecha).





