Explainable model of supervised and unsupervised learning to characterize the coverage of social programs in vulnerable territories of Peru
DOI:
https://doi.org/10.51252/rcsi.v6i2.1266Keywords:
clustering, social targeting, artificial intelligence, historical poverty, logistic regressionAbstract
This study developed and evaluated an explainable supervised and unsupervised learning approach to characterize social program coverage and identify historically vulnerable territories in Peru. A total of 35,952 district-month observations from 1,893 districts and 19 periods were integrated using MIDIS administrative data and poverty maps. Based on 13 coverage indicators, four clustering algorithms and three classifiers were compared through cross-validation and an independent test set. K-means with two clusters was selected as the main solution, achieving a silhouette score of 0.540 and identifying a profile associated with greater coverage and historical vulnerability. For classification, Elastic Net using only coverage indicators achieved an F1-score of 0.655 and, in the independent test set, a sensitivity of 0.829, an F1-score of 0.647, and a ROC-AUC of 0.871, with no significant differences compared with the other models. Explainability analysis highlighted Juntos coverage, estimated FONCODES users, and the number of programs present. The findings show that administrative records can support territorial prioritization when complemented with updated information and institutional review.
Downloads
References
Aiken, E., Bellue, S., Karlan, D., Udry, C., & Blumenstock, J. E. (2022). Machine learning and phone data can improve targeting of humanitarian aid. Nature, 603(7903), 864–870. https://doi.org/10.1038/s41586-022-04484-9
Altındağ, O., O’Connell, S. D., Şaşmaz, A., Balcıoğlu, Z., Cadoni, P., Jerneck, M., & Foong, A. K. (2021). Targeting humanitarian aid using administrative data: Model design and validation. Journal of Development Economics, 148, 102564. https://doi.org/10.1016/j.jdeveco.2020.102564
Amarante, V., Rossel, C., & Sánchez Laguardia, G. (2026). Social Protection Coverage in Latin America: Aggregate and Individual-level Transformations. Journal of Human Development and Capabilities, 1–27. https://doi.org/10.1080/19452829.2025.2610195
Amarasinghe, K., Rodolfa, K. T., Lamba, H., & Ghani, R. (2023). Explainable machine learning for public policy: Use cases, gaps, and research directions. Data & Policy, 5, e5. https://doi.org/10.1017/dap.2023.2
Arbelaitz, O., Gurrutxaga, I., Muguerza, J., Pérez, J. M., & Perona, I. (2013). An extensive comparative study of cluster validity indices. Pattern Recognition, 46(1), 243–256. https://doi.org/10.1016/j.patcog.2012.07.021
Asri, V., Michaelowa, K., Panda, S., & Paul, S. B. (2022). The pursuit of simplicity: Can simplifying eligibility criteria improve social pension targeting? Journal of Economic Behavior & Organization, 200, 820–846. https://doi.org/10.1016/j.jebo.2022.06.003
Behrens, J. T. (1997). Principles and procedures of exploratory data analysis. Psychological Methods, 2(2), 131–160. https://doi.org/10.1037/1082-989X.2.2.131
Borga, L. G., & D’Ambrosio, C. (2021). Social protection and multidimensional poverty: Lessons from Ethiopia, India and Peru. World Development, 147, 105634. https://doi.org/10.1016/j.worlddev.2021.105634
Ferreira, L. Z., Utazi, C. E., Huicho, L., Nilsen, K., Hartwig, F. P., Tatem, A. J., & Barros, A. J. D. (2022). Geographic inequalities in health intervention coverage – mapping the composite coverage index in Peru using geospatial modelling. BMC Public Health, 22(1), 2104. https://doi.org/10.1186/s12889-022-14371-7
García, S., Fernández, A., Luengo, J., & Herrera, F. (2010). Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power. Information Sciences, 180(10), 2044–2064. https://doi.org/10.1016/j.ins.2009.12.010
García, V. H. B., Martínez, F. de M. G., Llamas Félix, B. I., & Esparza, R. M. V. (2024). Medición de la pobreza multidimensional en México mediante un análisis bibliométrico y de ecuaciones estructurales. Iberoamerican Journal of Science Measurement and Communication, 4(2), 1–23. https://doi.org/10.47909/ijsmc.1354
Hall, O., Ohlsson, M., & Rögnvaldsson, T. (2022). A review of explainable AI in the satellite data, deep machine learning, and human poverty domain. Patterns, 3(10), 100600. https://doi.org/10.1016/j.patter.2022.100600
Harron, K., Dibben, C., Boyd, J., Hjern, A., Azimaee, M., Barreto, M. L., & Goldstein, H. (2017). Challenges in administrative data linkage for research. Big Data & Society, 4(2), 205395171774567. https://doi.org/10.1177/2053951717745678
Henderson, H., & Follett, L. (2022). Targeting social safety net programs on human capabilities. World Development, 151, 105741. https://doi.org/10.1016/j.worlddev.2021.105741
Hinojosa Pérez, J. A., Avalos, H. R. B., Salazar, I. Y. V., & Carrasco Mamani, S. C. (2024). Social Programs and Socioeconomic Variables: Their Impact on Peruvian Regional Poverty (2013–2022). Economies, 12(8), 197. https://doi.org/10.3390/economies12080197
INEI. (2015). Mapa de Pobreza Provincial y Distrital 2013. Gob.pe. https://www.gob.pe/institucion/inei/informes-publicaciones/3204925-mapa-de-pobreza-provincial-y-distrital-2013
INEI. (2020). Mapa de Pobreza Provincial y Distrital 2018. Gob.pe. https://www.gob.pe/institucion/inei/informes-publicaciones/3204872-mapa-de-pobreza-provincial-y-distrital-2018
INEI. (2026). INEI: Pobreza monetaria alcanzó al 25,7% de la población en el Perú durante el año 2025. Gob.pe. https://www.gob.pe/institucion/inei/noticias/1387366-inei-pobreza-monetaria-alcanzo-al-25-7-de-la-poblacion-en-el-peru-durante-el-ano-2025
Jain, A. K. (2010). Data clustering: 50 years beyond K-means. Pattern Recognition Letters, 31(8), 651–666. https://doi.org/10.1016/j.patrec.2009.09.011
Jung, W., Benotsmane, R., Stoeffler, Q., Kim, A. H., Ghadimi, S., Hosseini, M., Ntarlagiannis, D., Ammari, T., Lu, Y., & Steiner, J. (2026). Contextualized poverty targeting with multimodal spatial data and machine learning in Brazzaville, Congo. Cities, 170, 106429. https://doi.org/10.1016/j.cities.2025.106429
Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804
Li, Q., Yu, S., Échevin, D., & Fan, M. (2022). Is poverty predictable with machine learning? A study of DHS data from Kyrgyzstan. Socio-Economic Planning Sciences, 81, 101195. https://doi.org/10.1016/j.seps.2021.101195
Lisboa, P. J. G., Saralajew, S., Vellido, A., Fernández-Domenech, R., & Villmann, T. (2023). The coming of age of interpretable and explainable machine learning models. Neurocomputing, 535, 25–39. https://doi.org/10.1016/j.neucom.2023.02.040
Maco, V. (2025). Lock-in from within: Challenges to expanding cash transfer programs in Peru. Journal of International and Comparative Social Policy, 41(3), 335–349. https://doi.org/10.1017/ics.2025.10066
Marcinkevičs, R., & Vogt, J. E. (2023). Interpretable and explainable machine learning: A methods‐centric overview with concrete examples. WIREs Data Mining and Knowledge Discovery, 13(3). https://doi.org/10.1002/widm.1493
MIDIS. (2026). Cobertura de los Programas Sociales adscritos al MIDIS- [Ministerio de Desarrollo e Inclusión Social - MIDIS]. Gob.pe. https://www.datosabiertos.gob.pe/dataset/cobertura-de-los-programas-sociales-adscritos-al-midis-ministerio-de-desarrollo-e-inclusión
Papadakis, T., Christou, I. T., Ipektsidis, C., Soldatos, J., & Amicone, A. (2024). Explainable and transparent artificial intelligence for public policymaking. Data & Policy, 6, e10. https://doi.org/10.1017/dap.2024.3
Paucara-Charca, R., Almanza Huamán, L. M., Szczcpansky Grobas, D., Landa Almanza, A. I., & Muñiz Laura, K. A. (2024). La eficacia de los efectos políticos en el desarrollo alcance y resultados de las políticas regionales orientadas para mitigar el impacto de la explotación laboral en menores de edad en especial en el sector minero de La Pampa Madre de Dios 2018. Revista Amazónica De Ciencias Sociales, 2(2), 12–21. https://doi.org/10.55873/racs.v2i2.262
Poulin, C., Trimmer, J., Press-Williams, J., Yachori, B., Khush, R., Peletz, R., & Delaire, C. (2022). Performance of a novel machine learning-based proxy means test in comparison to other methods for targeting pro-poor water subsidies in Ghana. Development Engineering, 7, 100098. https://doi.org/10.1016/j.deveng.2022.100098
Schober, P., Boer, C., & Schwarte, L. A. (2018). Correlation Coefficients: Appropriate Use and Interpretation. Anesthesia & Analgesia, 126(5), 1763–1768. https://doi.org/10.1213/ANE.0000000000002864
Smythe, I. S., & Blumenstock, J. E. (2022). Geographic microtargeting of social assistance with high-resolution poverty maps. Proceedings of the National Academy of Sciences, 119(32). https://doi.org/10.1073/pnas.2120025119
Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W. (2019). Calibration: the Achilles heel of predictive analytics. BMC Medicine, 17(1), 230. https://doi.org/10.1186/s12916-019-1466-7
Van den Broeck, J., Argeseanu Cunningham, S., Eeckels, R., & Herbst, K. (2005). Data Cleaning: Detecting, Diagnosing, and Editing Data Abnormalities. PLoS Medicine, 2(10), e267. https://doi.org/10.1371/journal.pmed.0020267
Varma, S., & Simon, R. (2006). Bias in error estimation when using cross-validation for model selection. BMC Bioinformatics, 7(1), 91. https://doi.org/10.1186/1471-2105-7-91
Zou, H., & Hastie, T. (2005). Regularization and Variable Selection Via the Elastic Net. Journal of the Royal Statistical Society Series B: Statistical Methodology, 67(2), 301–320. https://doi.org/10.1111/j.1467-9868.2005.00503.x
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 William J. Marin-Rodriguez, Viviana I. Vellón-Flores, José L. Ausejo-Sánchez, Timoteo Solano-Armas, Yolanda E. Maguiña-Poma, Sixto Samanamud-Malca

This work is licensed under a Creative Commons Attribution 4.0 International License.
The authors retain their rights:
a. The authors retain their trademark and patent rights, as well as any process or procedure described in the article.
b. The authors retain the right to share, copy, distribute, execute and publicly communicate the article published in the Revista Científica de Sistemas e Informática (RCSI) (for example, place it in an institutional repository or publish it in a book), with an acknowledgment of its initial publication in the RCSI.
c. Authors retain the right to make a subsequent publication of their work, to use the article or any part of it (for example: a compilation of their works, notes for conferences, thesis, or for a book), provided that they indicate the source of publication (authors of the work, journal, volume, number and date).





