Detección de enfermedades cardiometabólicas pediátricas usando técnicas de machine learning
Fecha
Autores
Título de la revista
ISSN de la revista
Título del volumen
Editor
Resumen
Childhood obesity and its cardiometabolic consequences represent a growing public health challenge in Mexico, where over a third of school-aged children are affected by excess weight. Despite this burden, validated and pediatric-specific predictive tools for early detection of cardiometabolic risk remain scarce. This study develops and evaluates machine learning models for the simultaneous prediction of three cardiometabolic outcomes: nutritional status (Body Mass Index classification), glucose alteration, and hypertension staging, in a pediatric population from Jalisco, Mexico. The dataset comprises 3,229 children and teenagers aged 5 to 16 years, drawn from the statewide Family Wellbeing Modules (Módulos de Bienestar Familiar) screening program of the State Health system. After preprocessing, three supervised learning approaches were implemented and compared: logistic regression, support vector machines (SVM) with a Radial Basis Function (RBF) kernel, and a cluster-then-classify hybrid strategy. Class imbalance was addressed through target-specific synthetic minority over-sampling technique (SMOTE) variants, and feature selection was performed via Wald-based statistical significance testing. Decision thresholds were optimized using the receiver operating characteristic (ROC) curve corner method on a dedicated validation set.
Logistic regression achieved test Area Under the Receiver Operating Characteristic Curve (AUC-ROC) values of 0.897 , 0.635 , and 0.628 for BMI classification, glucose alteration, and hypertension staging, respectively, while maintaining clinical interpretability through odds ratios and significant predictor rankings.
The SMOTE-augmented SVM models reached AUC-ROC values above 0.947 across all three outcomes, serving as a performance upper bound. Key predictors were consistent with established clinical knowledge: waist circumference dominated BMI classification, random glucose drove glucose alteration prediction, and diastolic and systolic pressure were the strongest predictors of hypertension staging.
These findings demonstrate that interpretable machine learning models, trained on field-collected anthropometric and behavioral data, can effectively support early cardiometabolic risk detection in pediatric populations within resource-limited healthcare settings.