Start of funding 01.01.2024

Machine Learning-Driven Analysis for Unraveling Genetic Origins of Rare Diseases

Prof. Dr. Julien Gagneur
Technische Universität München
School of Computation, Information and Technology

Prof. Dr. Stephen B. Montgomery
Stanford University
Department of Pathology



A large proportion of patients with rare diseases remain undiagnosed, resulting in inappropriate treatments and psychological distress. Addressing this challenge, our laboratory has developed the tools OUTRIDER and FRASER 2.0 to detect aberrant gene expression and splicing events in RNA sequencing data. Integrated into the Detection of RNA Outliers Pipeline (DROP), these tools have been applied to a large rare disease dataset of the European Solve-RD project, which led to the development of diagnostic guidelines. This project seeks to apply these methodologies to two large North American rare disease datasets from the Undiagnosed Disease Network and the GREGoR consortium. This initiative has three key objectives: firstly, to offer molecular diagnoses to patients within the UDN and GREGoR consortiums. Secondly, to validate the previously established workflows and guidelines and provide clinicians with a validated framework for the effective use of RNA sequencing data, facilitating its wider adoption and efficiency in clinical practice. Lastly, this project will contribute to the refinement of our diagnostic tools, leading to an increased rate of accurate diagnosis for rare diseases.

Final report:
The Undiagnosed Diseases Network (UDN) is a study funded by the National Institutes of Health with the goal of diagnosing patients with rare diseases and better understanding the underlying disease mechanisms. The participants, primarily children with neurodevelopmental and neuromuscular disorders, undergo a thorough clinical evaluation, including specialist consultations, laboratory testing, and phenotyping using a standardized protocol. DNA and RNA are extracted and sequenced from blood, skin, and muscle samples. For this analysis, we used the data available in the database of Genotypes and Phenotypes (dbGaP), which consists of 1,327 families with 4,248 DNA and 831 RNA samples.

The aim of this project was twofold. First, we wanted to test the performance of the Detection of RNA Outliers Pipeline (DROP)1 developed in our laboratory and validate the RNA-seq analysis guidelines developed in the European Solve-RD project using the UDN dataset. Second, we wanted to use this approach to find candidate variants and ultimately diagnoses for UDN cases that remain unsolved.

The effect of a genetic variant is often uncertain, especially when it occurs in a non-coding region of the DNA, making it difficult to predict its impact on the protein. RNA-seq overcomes this hurdle by allowing direct observation of a variant’s effects on the RNA, such as changes in transcript abundance or splicing defects. For the first time in the UDN, we performed a systematic analysis of all RNA samples using DROP to reveal molecular effects that might otherwise be missed by traditional DNA-centric approaches.

During the quality control step of the DROP analysis, we identified inconsistencies in the dataset, including incorrect tissue annotations and DNA/RNA sample swaps. After correcting these errors, we applied the tools OUTRIDER2 and FRASER 2.03, 4 to detect aberrant gene expression and splicing events in the RNA sequencing data. In the 476 RNA samples from 259 families to which DROP was applied, we detected 3,081 expression and 27,569 splicing outliers. We combined the RNA outliers with a phenotypic similarity score, which measures how closely the participant’s symptoms match those associated with defects in the gene where the aberrant event was found. In addition, we annotated the genetic variants with their predicted effects and deleteriousness scores. Using this combined information, we systematically analyzed each case by confirming the association of the gene with the patient’s phenotype, and then identifying the most plausible causal genetic variant by matching the type of aberrant RNA event (e.g., splicing change with a splice site variant, underexpression outlier with a truncating variant).

Of the 259 families analyzed, we identified 74 cases that were either already solved and described in previous publications5, 6, 7, 8 or in ClinVar. In 17 of these cases, aberrant expression or splicing events were previously described in the RNA-seq samples, 14 of which we could identify with DROP. Among the remaining 57 solved cases, where no such events had been reported, DROP detected an RNA-related abnormality in 15 cases. In addition, we found candidate variants for nine previously unsolved cases, one of which is highlighted below. In a male participant with global developmental delay, motor delay, urinary system, visual system, and dermatologic abnormalities, DROP detects four adjacent underexpressed genes on chromosome 15. The underexpressed genes, NIPA1, NIPA2, TUBGCP5, CYFIP1, are the four genes associated with the 15q11.2 microdeletion syndrome. The phenotype of this UDN participant fits well with the abnormalities described in patients with this rare disease, making this the most likely condition from which this patient suffers.

Looking ahead, with the recent introduction of reimbursement for whole genome sequencing for rare diseases in Germany and the establishment of the European Rare Diseases Research Alliance (ERDERA), the amount of sequencing data is increasing and with it the potential for RNA-based diagnostics. We have demonstrated that our DROP-centric RNA-seq workflow is both effective and adaptable, making it a valuable tool for clinicians seeking to diagnose rare diseases.

My time at the Montgomery Lab at Stanford played an important role in the success of this project. In the context of the UDN project described here, I had access to more detailed participant data, including medical and family history, as well as opportunities to discuss cases with lab members and other people at Stanford’s UDN clinical site. These discussions were crucial for identifying candidate variants and retrieving previously solved cases. Furthermore, my research stay at Stanford marks the starting point for future projects and collaborations. We will follow up on the nine undiagnosed cases with genetic counselors at the UDN clinical sites, potentially leading to long awaited diagnoses. Looking ahead, we foresee further collaborations with Stephen Montgomery’s lab, focusing on developing new computational methods for rare disease diagnostics and analyzing new data from the GREGoR (Genomics Research to Elucidate the Genetics of Rare diseases) consortium and the ERDERA project in Europe.

  1. Y'epez VA, Mertes C, Müller MF, Klaproth-Andrade D, Wachutka L, Frésard L, Gusic M, Scheller IF, Goldberg PF, Prokisch H, Gagneur J. Detection of aberrant gene expression events in RNA sequencing data. Nat Protoc. 2021 Feb;16(2):1276-1296. doi: 10.1038/s41596-020-00462-5. Epub 2021 Jan 18. PMID: 33462443.
  2. Brechtmann F, Mertes C, Matusevičiūtė A, Y'epez VA, Avsec Z, Herzog M, Bader DM, ̌Prokisch H, Gagneur J. OUTRIDER: A Statistical Method for Detecting Aberrantly Expressed Genes in RNA Sequencing Data. Am J Hum Genet. 2018 Dec 6;103(6):907-917. doi: 10.1016/j.ajhg.2018.10.025. Epub 2018 Nov 29. PMID: 30503520; PMCID: PMC6288422.
  3. Scheller IF, Lutz K, Mertes C, Y'epez VA, Gagneur J. Improved detection of aberrant splicing with FRASER 2.0 and the intron Jaccard index. Am J Hum Genet. 2023 Dec 7;110(12):2056-2067. doi: 10.1016/j.ajhg.2023.10.014. Epub 2023 Nov 24. PMID: 38006880; PMCID: PMC10716352.
  4. Mertes C, Scheller IF, Y'epez VA, Cçelik MH, Liang Y, Kremer LS, Gusic M, Prokisch H, Gagneur J. Detection of aberrant splicing events in RNA-seq data using FRASER. Nat Commun. 2021 Jan 22;12(1):529. doi: 10.1038/s41467-020-20573-7. Erratum in: Nat Commun. 2022 Jun 16;13(1):3474. doi: 10.1038/s41467-022-31242-2. PMID: 33483494; PMCID: PMC7822922.
  5. Splinter K, Adams DR, Bacino CA, Bellen HJ, Bernstein JA, Cheatle-Jarvela AM, Eng CM, Esteves C, Gahl WA, Hamid R, Jacob HJ, Kikani B, Koeller DM, Kohane IS, Lee BH, Loscalzo J, Luo X, McCray AT, Metz TO, Mulvihill JJ, Nelson SF, Palmer CGS, Phillips JA 3rd, Pick L, Postlethwait JH, Reuter C, Shashi V, Sweetser DA, Tifft CJ, Walley NM, Wangler MF, Westerfield M, Wheeler MT, Wise AL, Worthey EA, Yamamoto S, Ashley EA, Undiagnosed Diseases Network Effect of Genetic Diagnosis on Patients with Previously Undiagnosed Disease. N Engl J Med. 2018 Nov 29;379(22):2131-2139. PMID: 30304647.
  6. Lee H, Huang AY, Wang LK, Yoon AJ, Renteria G, Eskin A, Signer RH, Dorrani N, Nieves-Rodriguez S, Wan J, Douine ED, Woods JD, Dell’Angelica EC, Fogel BL, Martin MG, Butte MJ, Parker NH, Wang RT, Shieh PB, Wong DA, Gallant N, Singh KE, Tavyev Asher YJ, Sinsheimer JS, Krakow D, Loo SK, Allard P, Papp JC, Undiagnosed Diseases Network, Palmer CGS, Martinez-Agosto JA, Nelson SF Diagnostic utility of transcriptome sequencing for rare Mendelian diseases. Genet Med. 2020 Mar;22(3):490-499. PMID: 31607746.
  7. Frésard L, Smail C, Ferraro NM, Teran NA, Li X, Smith KS, Bonner D, Kernohan KD, Marwaha S, Zappala Z, Balliu B, Davis JR, Liu B, Prybol CJ, Kohler JN, Zastrow DB, Reuter CM, Fisk DG, Grove ME, Davidson JM, Hartley T, Joshi R, Strober BJ, Utiramerur S, Undiagnosed Diseases Network, Care4Rare Canada Consortium, Lind L, Ingelsson E, Battle A, Bejerano G, Bernstein JA, Ashley EA, Boycott KM, Merker JD, Wheeler MT, Montgomery SB Identification of rare-disease genes using blood transcriptome sequencing and large control cohorts. Nat Med. 2019 Jun;25(6):911-919. PMID: 31160820.
  8. Murdock DR, Dai H, Burrage LC, Rosenfeld JA, Ketkar S, Müller MF, Y'epez VA, Gagneur J, Liu P, Chen S, Jain M, Zapata G, Bacino CA, Chao HT, Moretti P, Craigen WJ, Hanchard NA, Undiagnosed Diseases Network, Lee B Transcriptome-directed analysis for Mendelian disease diagnosis overcomes limitations of conventional genomic testing. J Clin Invest. 2021 Jan 04;131(1). PMID: 33001864.