Artículo
A labeled medical records corpus for the timely detection of rare diseases using machine learning approaches
Fecha de publicación:
02/2025
Editorial:
Nature
Revista:
Scientific Reports
ISSN:
2045-2322
Idioma:
Inglés
Tipo de recurso:
Artículo publicado
Clasificación temática:
Resumen
Rare diseases (RDs) are a group of pathologies that individually affect less than 1 in 2000 people but collectively impact around 7% of the world’s population. Most of them affect children, are chronic and progressive, and have no specific treatment. RD patients face diagnostic challenges, with an average diagnosis time of 5 years, multiple specialist visits, and invasive procedures. This ‘diagnostic odyssey’ can be detrimental to their health. Machine learning (ML) has the potential to improve healthcare by providing more personalized and accurate patient management, diagnoses, and in some cases, treatments. Leveraging the MIMIC-III database and additional medical notes from different sources such as in-house data, PubMed and chatGPT, we propose a labeled dataset for early RD detection in hospital settings. Applying various supervised ML methods, including logistic regression, decision trees, support vector machine (SVM), deep learning methods (LSTM and CNN), and Transformers (BERT), we validated the use of the proposed resource, achieving 92.7% F-measure and a 96% AUC using SVM. These findings highlight the potential of ML in redirecting RD patients towards more accurate diagnostic pathways and presents a corpus that can be used for future development and refinements.
Palabras clave:
RARE DISEASES
,
MACHINE LEARNING
,
ARTIFICIAL CORPUS
Archivos asociados
Licencia
Identificadores
Colecciones
Articulos(CCT - SAN LUIS)
Articulos de CTRO.CIENTIFICO TECNOL.CONICET - SAN LUIS
Articulos de CTRO.CIENTIFICO TECNOL.CONICET - SAN LUIS
Citación
Rolando, Matias; Raggio, Victor; Naya, Hugo; Spangenberg, Lucia; Cagnina, Leticia Cecilia; A labeled medical records corpus for the timely detection of rare diseases using machine learning approaches; Nature; Scientific Reports; 15; 1; 2-2025; 1-10
Compartir
Altmétricas