Prediction and calibration of payment default risk using machine learning in residential water service customers

Authors

DOI:

https://doi.org/10.51252/rcsi.v6i2.1265

Keywords:

billing, collection, payment behavior, stratification, temporal validation

Abstract

Water service payment delinquency limits the availability of resources required to sustain operations and hinders the prioritization of collection activities. This study aimed to evaluate machine learning models for estimating the payment default risk of residential customer bills at EMAPAB S.A. A total of 264,382 bills from 6,358 customers were analyzed using 51 predictors derived from historical billing and payment records. Random Forest, CatBoost, and a feedforward neural network were compared using Optuna optimization, stratified and grouped ten-fold cross-validation, and an independent temporal test. Random Forest achieved the best performance on the temporal test, with an accuracy of 0.9037, an F1-score of 0.6333, a ROC-AUC of 0.8681, and a PR-AUC of 0.6892. Its probabilities were subsequently calibrated using Platt Scaling and stratified into low-, medium-, and high-risk levels, accounting for 74.61%, 13.96%, and 11.44% of the bills, respectively. The results show that the proposed approach can identify bills with higher default risk and support the prioritization of the collection portfolio, although its operational impact requires prospective validation.

Downloads

Download data is not yet available.

References

Abuzaid, A., & Alkronz, E. (2024). A comparative study on univariate outlier winsorization methods in data science context. Italian Journal of Applied Statistics, 36(1). https://doi.org/ 10.26398/IJAS.0036-004

Ayala esquivel, B. D., & Cabrera Tapia, C. F. (2021). La importancia de la economía del agua. RD-ICUAP, 78–91. https://doi.org/10.32399/icuap.rdic.2448-5829.2021.21.630

Bashar, M. A., Nayak, R., Astin-Walmsley, K., & Heath, K. (2023). Machine learning for predicting propensity-to-pay energy bills. Intelligent Systems with Applications, 17, 200176. https://doi.org/10.1016/j.iswa.2023.200176

Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

García Hernández, J. M., & Torres Moreno, W. N. (2023). Predicción de riesgo de impago en institución financiera usando modelos de Machine Learning [Universidad Tecnológica Centroamericana UNITEC]. https://repositorio.unitec.edu/xmlui/handle/123456789/12906

García, S., Fernández, A., Luengo, J., & Herrera, F. (2010). Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power. Information Sciences, 180(10), 2044–2064. https://doi.org/10.1016/j.ins.2009.12.010

Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning. Springer New York. https://doi.org/10.1007/978-0-387-84858-7

Hornik, K., Stinchcombe, M., & White, H. (1989). Multilayer feedforward networks are universal approximators. Neural Networks, 2(5), 359–366. https://doi.org/10.1016/0893-6080(89)90020-8

Lara López, F., Manríquez García, N., & Quintero Rodríguez, J. O. (2023). Comportamiento de la demanda del consumo de agua potable por zonas en Mazatlán, Sinaloa. INTER DISCIPLINA, 11(31), 317–337. https://doi.org/10.22201/ceiich.24485705e.2023.31.86085

Lescano-Delgado, M. (2024). Avances en el uso de inteligencia artificial para la mejora del control y la detección de fraudes en organizaciones. Revista Científica de Sistemas e Informática, 4(2), e671. https://doi.org/10.51252/rcsi.v4i2.671

Li, J., Cheng, K., Wang, S., Morstatter, F., Trevino, R. P., Tang, J., & Liu, H. (2018). Feature Selection. ACM Computing Surveys, 50(6), 1–45. https://doi.org/10.1145/3136625

Mervin, L., H., Afzal, A. M., Engkvist, O., & Bender, A. (2020). Comparison of Scaling Methods to Obtain Calibrated Probabilities of Activity for Protein–Ligand Predictions. Journal of Chemical Information and Modeling, 60(10), 4546–4559. https://doi.org/10.1021/acs.jcim.0c00476

Montenegro-Chasquibol, B., Labajos-Portocarrero, H., Terrones-Suarez, O., Minga-Sarmiento, R., & Fasanando-Garcia, S. (2024). Gestión de cobranzas y su influencia en la liquidez de una empresa inmobiliaria. Revista Amazónica de Ciencias Económicas, 3(2), e741. https://doi.org/10.51252/race.v3i2.741

Odesola, P. A., Adegoke, A. A., & Babalola, I. (2025). Model uncertainty quantification: A post hoc calibration approach for heart disease prediction. https://doi.org/10.1101/2025.09.28.25336834

Olmedo, A. T., Castro, J. R., Reátegui, R., & Castillo, T. (2023). Aplicación de algoritmos de Machine Learning para la segmentación del consumo de agua potable. Caso de estudio en Catamayo, Ecuador. 2023 18th Iberian Conference on Information Systems and Technologies (CISTI), 1–6. https://doi.org/10.23919/CISTI58278.2023.10211900

Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., & Gulin, A. (2018). CatBoost: unbiased boosting with categorical features. Advances in Neural Information Processing Systems 31.

Ramirez Villacorta, J. M. (2024). Modelo predictivo basado en redes neuronales artificiales para pronosticar el consumo de agua potable en la ciudad de Iquitos [Universidad Nacional Federico Villarreal]. https://hdl.handle.net/20.500.13084/8532

Roberts, D. R., Bahn, V., Ciuti, S., Boyce, M. S., Elith, J., Guillera‐Arroita, G., Hauenstein, S., Lahoz‐Monfort, J. J., Schröder, B., Thuiller, W., Warton, D. I., Wintle, B. A., Hartig, F., & Dormann, C. F. (2017). Cross‐validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913–929. https://doi.org/10.1111/ecog.02881

Soudachanh, S., Langergraber, G., & Salhofer, S. (2022). Sanitation planning for resettlement sites in Laos. Journal of Water, Sanitation and Hygiene for Development, 12(3), 248–257. https://doi.org/10.2166/washdev.2022.178

Tharwat, A. (2021). Classification assessment methods. Applied Computing and Informatics, 17(1), 168–192. https://doi.org/10.1016/j.aci.2018.08.003

Yajure Ramírez, C. A. (2022). Uso de algoritmos de Machine Learning para analizar los datos de energía eléctrica facturada en la Ciudad de Buenos Aires durante el período 2010–2021. Ciencia, Ingenierías y Aplicaciones, 5(2), 7–37. https://doi.org/10.22206/cyap.2022.v5i2.pp7-37

Downloads

Published

2026-07-20

How to Cite

Valles-Coral, M. A., Cuello-Sangama, R., Rengifo-Amasifen, R., Vidaurre-Rojas, P., Prieto-Luna, J. C., Holgado-Apaza, L. A., & Injante, R. (2026). Prediction and calibration of payment default risk using machine learning in residential water service customers. Revista Científica De Sistemas E Informática, 6(2), e1265. https://doi.org/10.51252/rcsi.v6i2.1265