Analysis of unpaid care work patterns using K-Means: A Machine Learning application on ENUT 2021 data
Keywords:
care economy, time use, unpaid work, machine learning, k-means, intersectionalityAbstract
This study applies machine learning and cluster analysis to identify sociodemographic profiles of women in Colombia, based on time dedicated to unpaid care work. Using data from the 2021 National Time Use Survey (ENUT), synthetic variables were constructed for domestic work, unpaid care, self-care, leisure, and a traditionalism index. The k-means algorithm, combined with principal component analysis (PCA) and validation criteria such as the elbow method and silhouette index, identified six profiles with contrasting care organization patterns. Results confirm the persistence of the sexual division of labor, with key implications for public policy and the care economy. They also highlight the importance of an intersectional approach to expose structural inequalities in the distribution of unpaid work burdens.
References
Arthur, D., & Vassilvitskii, S. (2007). k-means++: The advantages of careful seeding. In Proceedings of the eighteenth annual ACM-SIAM symposium on Discrete algorithms (pp. 1027–1035). Society for Industrial and Applied Mathematics.
Benería, L. (2003). Gender, development, and globalization: Economics as if all people mat-tered. Routledge.
Carrasco, C. (2006). Invisibles y fundamentales: El trabajo de cuidados en Europa. Icaria. Crenshaw, K. (1991). Mapping the margins: Intersectionality, identity politics, and violence against women of color. Stanford Law Review,
(6), 1241-1299.
Filmer, D., & Pritchett, L. H. (2001). Estimating wealth effects without expenditure data —or tears: An application to educational enrollments in states of India. Demography, 38(1), 115–132. https://doi.org/10.1353/dem.2001.0003
Fraser, N. (2011). Fortunas del feminismo: Del capitalismo gestionado por el Estado a la crisis neoliberal. Traficantes de Sueños.
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow (2nd ed.). O’Reilly Media.
Hill Collins, P. (2000). Black feminist thought: Knowledge, consciousness, and the politics of empowerment. Routledge.
Hunter, J. D. (2007). Matplotlib: A 2D Graphics Environment. Computing in Science & Engineering, 9(3), 90-95.
Jolliffe, I. T. (2002). Principal Component Analysis (2nd ed.). Springer.
Jolliffe, I. T., & Cadima, J. (2016). Principal component analysis: a review and recent developments. Philosophical Transactions of the Royal Society: A: Mathematical, Physical and Engineering Sciences, 374(2065), 20150202.
Kolenikov, S., & Angeles, G. (2009). Socioeconomic status measurement with discrete proxy variables: Is principal component analysis a reliable answer? Review of Income and Wealth, 55(1), 128–165. https://doi.org/10.1111/j.1475-4991.2008.00309.x
MacQueen, J. (1967a). Some Methods for Classification and Analysis of Multiva- riate Observations. Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, 1, 281-297.
MacQueen, J. (1967b). Some methods for classification and analysis of multivariate observations. Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, 1, 281-297.
McKinney, W. (2010). Data Structures for Statistical Computing in Python. Proceedings of the 9th Python in Science Conference, 51-56.
Moreno Salamanca, E. N. (2017). Organización social del cuidado y medición del trabajo no remunerado en Colombia. Colombia Internacional, (91), 23-50.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Courna- peau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12, 2825-2830.
Pérez Orozco, A. (2012). Subversión feminista de la economía: Aportes para un debate sobre el conflicto capital-vida. Traficantes de Sueños.
Rousseeuw, P. J. (1987). Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53-65. https://doi.org/10.1016/0377-0427(87)90125-7
Shlens, J. (2014). A tutorial on principal component analysis [arXiv:1404.1100]. https://arxiv.org/abs/1404.1100
Thorndike, R. L. (1953). Who belongs in the family? Psychometrika, 18(4), 267-276. https://doi.org/10.1007/BF02289263
Vyas, S., & Kumaranayake, L. (2006). Constructing socio-economic status indices: How to use principal components analysis. Health Policy and Planning, 21(6), 459–468. https://doi.org/10.1093/heapol/czl029
Waskom, M. (2021). Seaborn: Statistical Data Visualization. Journal of Open Source Software, 6(60), 3021.
Published
How to Cite
Issue
Section
Copyright (c) 2026 Intercambio

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.