目的 构建假肥大型肌营养不良(pseudohypertrophic muscular dystrophy, PMD)数据库及机器学习模型,探究疾病的基因型-表型规律。 方法 回顾性收集2010年1月—2024年12月就诊于中南大学湘雅医院的PMD患儿数据,以及1987年1月—2024年12月PubMed收录文献中报道的病例,构建基因型-表型数据库,通过微信小程序实现在线查询。基于该数据库,整合随机森林、极端梯度提升和轻量级梯度提升机3种算法,通过软投票集成策略构建针对微小变异所致临床表型的机器学习预测模型,并比较预测模型与阅读框规则的表型预测效能,同时基于Streamlit平台完成模型的在线部署。 结果 数据库共纳入17 053例PMD患儿(本地队列472例,文献数据16 581例)。基于筛选后的微小变异数据集开展建模验证,在内部测试集中,机器学习模型的受试者操作特征曲线的曲线下面积(area under the curve, AUC)为0.924(95%CI:0.881~0.963),优于阅读框模型的0.652(95%CI:0.591~0.717),差异有统计学意义(P<0.001)。在外部测试集中,机器学习模型的AUC为0.854(95%CI:0.736~1.000),阅读框模型的AUC为0.667(95%CI:0.500~1.000),两者差异无统计学意义(P>0.05)。 结论 该研究构建的PMD数据库与机器学习预测模型为PMD的表型预测提供了高效可靠的新工具。
Objective To establish a genotype-phenotype database and machine learning models for pseudohypertrophic muscular dystrophy (PMD), and to explore genotype-phenotype correlations of the disease. Methods Clinical data of children with PMD admitted to Xiangya Hospital, Central South University from January 2010 to December 2024, together with cases retrieved from the PubMed database between January 1987 and December 2024, were retrospectively collected to construct a genotype-phenotype database, with an online query function via a WeChat mini-program. Based on this database, Random Forest, Extreme Gradient Boosting, and Light Gradient Boosting Machine algorithms were integrated using a soft voting ensemble strategy to build a machine learning model predicting clinical phenotypes associated with small variants. The predictive performance of the model was compared with that of the reading-frame rule. The model was deployed online via the Streamlit platform. Results The database included 17 053 PMD cases, comprising 472 patients in the local cohort and 16 581 literature-derived cases. Modeling and validation were performed on a filtered dataset comprising small variants. In the internal test set, the machine learning model achieved an area under the receiver operating characteristic curve (AUC) of 0.924 (95%CI: 0.881-0.963), significantly higher than the reading-frame rule AUC of 0.652 (95%CI: 0.591-0.717) (P<0.001). In the external test set, the machine learning model achieved an AUC of 0.854 (95%CI: 0.736-1.000), compared to 0.667 (95%CI: 0.500-1.000) for the reading-frame rule, with no statistically significant difference (P>0.05). Conclusions The constructed PMD genotype-phenotype database and machine learning prediction model provide an efficient and reliable novel tool for phenotype prediction in PMD.