假肥大型肌营养不良的基因型-表型数据库和机器学习模型的构建

楚亦轩, 张慈柳, 彭镜

中国当代儿科杂志 ›› 2026, Vol. 28 ›› Issue (8) : 991-997.

PDF(1562 KB)
HTML
PDF(1562 KB)
HTML
中国当代儿科杂志 ›› 2026, Vol. 28 ›› Issue (8) : 991-997. DOI: 10.7499/j.issn.1008-8830.2511119
论著·临床研究

假肥大型肌营养不良的基因型-表型数据库和机器学习模型的构建

作者信息 +

Construction of a genotype-phenotype database and machine learning models for pseudohypertrophic muscular dystrophy

Author information +
文章历史 +

摘要

目的 构建假肥大型肌营养不良(pseudohypertrophic muscular dystrophy, PMD)数据库及机器学习模型,探究疾病的基因型-表型规律。 方法 回顾性收集2010年1月—2024年12月就诊于中南大学湘雅医院的PMD患儿数据,以及1987年1月—2024年12月PubMed收录文献中报道的病例,构建基因型-表型数据库,通过微信小程序实现在线查询。基于该数据库,整合随机森林、极端梯度提升和轻量级梯度提升机3种算法,通过软投票集成策略构建针对微小变异所致临床表型的机器学习预测模型,并比较预测模型与阅读框规则的表型预测效能,同时基于Streamlit平台完成模型的在线部署。 结果 数据库共纳入17 053例PMD患儿(本地队列472例,文献数据16 581例)。基于筛选后的微小变异数据集开展建模验证,在内部测试集中,机器学习模型的受试者操作特征曲线的曲线下面积(area under the curve, AUC)为0.924(95%CI:0.881~0.963),优于阅读框模型的0.652(95%CI:0.591~0.717),差异有统计学意义(P<0.001)。在外部测试集中,机器学习模型的AUC为0.854(95%CI:0.736~1.000),阅读框模型的AUC为0.667(95%CI:0.500~1.000),两者差异无统计学意义(P>0.05)。 结论 该研究构建的PMD数据库与机器学习预测模型为PMD的表型预测提供了高效可靠的新工具。

Abstract

Objective To establish a genotype-phenotype database and machine learning models for pseudohypertrophic muscular dystrophy (PMD), and to explore genotype-phenotype correlations of the disease. Methods Clinical data of children with PMD admitted to Xiangya Hospital, Central South University from January 2010 to December 2024, together with cases retrieved from the PubMed database between January 1987 and December 2024, were retrospectively collected to construct a genotype-phenotype database, with an online query function via a WeChat mini-program. Based on this database, Random Forest, Extreme Gradient Boosting, and Light Gradient Boosting Machine algorithms were integrated using a soft voting ensemble strategy to build a machine learning model predicting clinical phenotypes associated with small variants. The predictive performance of the model was compared with that of the reading-frame rule. The model was deployed online via the Streamlit platform. Results The database included 17 053 PMD cases, comprising 472 patients in the local cohort and 16 581 literature-derived cases. Modeling and validation were performed on a filtered dataset comprising small variants. In the internal test set, the machine learning model achieved an area under the receiver operating characteristic curve (AUC) of 0.924 (95%CI: 0.881-0.963), significantly higher than the reading-frame rule AUC of 0.652 (95%CI: 0.591-0.717) (P<0.001). In the external test set, the machine learning model achieved an AUC of 0.854 (95%CI: 0.736-1.000), compared to 0.667 (95%CI: 0.500-1.000) for the reading-frame rule, with no statistically significant difference (P>0.05). Conclusions The constructed PMD genotype-phenotype database and machine learning prediction model provide an efficient and reliable novel tool for phenotype prediction in PMD.

关键词

假肥大型肌营养不良 / 阅读框规则 / 机器学习

Key words

Pseudohypertrophic muscular dystrophy / Reading frame rule / Machine learning

引用本文

导出引用
楚亦轩, 张慈柳, 彭镜. 假肥大型肌营养不良的基因型-表型数据库和机器学习模型的构建[J]. 中国当代儿科杂志. 2026, 28(8): 991-997 https://doi.org/10.7499/j.issn.1008-8830.2511119
Yi-Xuan CHU, Ci-Liu ZHANG, Jing PENG. Construction of a genotype-phenotype database and machine learning models for pseudohypertrophic muscular dystrophy[J]. Chinese Journal of Contemporary Pediatrics. 2026, 28(8): 991-997 https://doi.org/10.7499/j.issn.1008-8830.2511119

参考文献

[1]
Duan D, Goemans N, Takeda S, et al. Duchenne muscular dystrophy[J]. Nat Rev Dis Primers, 2021, 7(1): 13. PMCID: PMC10557455. DOI: 10.1038/s41572-021-00248-3 .
[2]
Muntoni F, Torelli S, Ferlini A. Dystrophin and mutations: one gene, several proteins, multiple phenotypes[J]. Lancet Neurol, 2003, 2(12): 731-740. DOI: 10.1016/s1474-4422(03)00585-4 .
[3]
Bushby K, Finkel R, Birnkrant DJ, et al. Diagnosis and management of Duchenne muscular dystrophy, part 1: diagnosis, and pharmacological and psychosocial management[J]. Lancet Neurol, 2010, 9(1): 77-93. DOI: 10.1016/S1474-4422(09)70271-6 .
[4]
Monaco AP, Bertelson CJ, Liechti-gallati S, et al. An explanation for the phenotypic differences between patients bearing partial deletions of the DMD locus[J]. Genomics, 1988, 2(1): 90-95. DOI: 10.1016/0888-7543(88)90113-9 .
[5]
中华医学会罕见病分会, 北京医学会罕见病分会. 抗肌萎缩蛋白病中国诊断指南[J]. 中华医学杂志, 2024, 104(11): 822-833. DOI: 10.3760/cma.j.cn112137-20231217-01402 .
[6]
Ahsan MM, Luna SA, Siddique Z. Machine-learning-based disease diagnosis: a comprehensive review[J]. Healthcare (Basel), 2022, 10(3): 541. PMCID: PMC8950225. DOI: 10.3390/healthcare10030541 .
[7]
Brauner LE, Yao Y, Grigull L, et al. Patient-oriented questionnaires and machine learning for rare disease diagnosis: a systematic review[J]. J Clin Med, 2024, 13(17): 5132. PMCID: PMC11396573. DOI: 10.3390/jcm13175132 .
[8]
Ahmed AS, Ahmed SS, Mohamed S, et al. Advancements in cholelithiasis diagnosis: a systematic review of machine learning applications in imaging analysis[J]. Cureus, 2024, 16(8): e66453. PMCID: PMC11380526. DOI: 10.7759/cureus.66453 .
[9]
Jan Z, Ai-ansari N, Mousa O, et al. The role of machine learning in diagnosing bipolar disorder: scoping review[J]. J Med Internet Res, 2021, 23(11): e29749. PMCID: PMC8663682. DOI: 10.2196/29749 .
[10]
Den dunnen JT, Dalgleish R, Maglott DR, et al. HGVS recommendations for the description of sequence variants: 2016 update[J]. Hum Mutat, 2016, 37(6): 564-569. DOI: 10.1002/humu.22981 .
[11]
Rentzsch P, Witten D, Cooper GM, et al. CADD: predicting the deleteriousness of variants throughout the human genome[J]. Nucleic Acids Res, 2019, 47(D1): D886-D894. PMCID: PMC6323892. DOI: 10.1093/nar/gky1016 .
[12]
Mclaren W, Gil L, Hunt SE, et al. The ensembl variant effect predictor[J]. Genome Biol, 2016, 17(1): 122. PMCID: PMC4893825. DOI: 10.1186/s13059-016-0974-4 .
[13]
Aartsma-Rus A, Van Deutekom JC, Fokkema IF, et al. Entries in the Leiden Duchenne muscular dystrophy mutation database: an overview of mutation types and paradoxical cases that confirm the reading-frame rule[J]. Muscle Nerve, 2006, 34(2): 135-144. DOI: 10.1002/mus.20586 .
[14]
Tuffery-Giraud S, Béroud C, Leturcq F, et al. Genotype-phenotype analysis in 2,405 patients with a dystrophinopathy using the UMD-DMD database: a model of nationwide knowledgebase[J]. Hum Mutat, 2009, 30(6): 934-945. DOI: 10.1002/humu.20976 .
[15]
Bladen CL, Salgado D, Monges S, et al. The TREAT-NMD DMD global database: analysis of more than 7,000 Duchenne muscular dystrophy mutations[J]. Hum Mutat, 2015, 36(4): 395-402. PMCID: PMC4405042. DOI: 10.1002/humu.22758 .
[16]
Stenson PD, Mort M, Ball EV, et al. The human gene mutation database: building a comprehensive mutation repository for clinical and molecular genetics, diagnostic testing and personalized genomic medicine[J]. Hum Genet, 2014, 133(1): 1-9. PMCID: PMC3898141. DOI: 10.1007/s00439-013-1358-4 .
[17]
Landrum MJ, Lee JM, Riley GR, et al. ClinVar: public archive of relationships among sequence variation and human phenotype[J]. Nucleic Acids Res, 2014, 42(D1): D980-D985. PMCID: PMC3965032. DOI: 10.1093/nar/gkt1113 .
[18]
Luce L, Abelleyro MM, Carcione M, et al. Analysis of complex structural variants in the DMD gene in one family[J]. Neuromuscul Disord, 2021, 31(3): 253-263. DOI: 10.1016/j.nmd.2020.11.015 .
[19]
Riley RD, Ensor J, Snell KIE, et al. External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges[J]. BMJ, 2016, 353: i3140. PMCID: PMC4916924. DOI: 10.1136/bmj.i3140 .
[20]
Bello L, Hoffman EP, Pegoraro E. Is it time for genetic modifiers to predict prognosis in Duchenne muscular dystrophy?[J]. Nat Rev Neurol, 2023, 19(7): 410-423. DOI: 10.1038/s41582-023-00823-0 .

脚注

所有作者均声明无利益冲突。


PDF(1562 KB)
HTML

Accesses

Citation

Detail

段落导航
相关文章

/