US2022316009A1PendingUtilityA1

Precision medicine portal for human diseases

Assignee: GENOME INT CORPORATIONPriority: Mar 26, 2021Filed: Mar 18, 2022Published: Oct 6, 2022
Est. expiryMar 26, 2041(~14.7 yrs left)· nominal 20-yr term from priority
C12Q 1/6883G16H 50/30G16H 20/10G16B 50/10G16H 10/20G16B 25/10G16H 50/20G16H 15/00G16B 20/50G16B 20/20G16B 45/00C12Q 2600/156G16H 70/40G16B 30/10C12Q 1/6869G16B 30/00C12Q 1/6825G16H 50/70G16B 40/20G16B 20/10
77
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method for genome analysis is provided. The method includes receiving a nucleotide string comprising a plurality of nucleotides from at least a portion of one or more individual patients' genome. The method also includes identifying a plurality of variants in said nucleotide string, assigning each identified variant a score based on a location of a variant and a predicted functional consequence, and determining a strength of a variation responsible for a trait or phenotypic manifestation of the variants. The method also includes identifying at least one phenotype, and displaying, in a graphic unit interface of a client device, said nucleotide string, the identified variants, and the at least one phenotype, in one or more genetic elements for one or more individual patients. A system and a non-transitory, computer-readable medium storing instructions to perform the above method are also provided.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method comprising:
 receiving a nucleotide string comprising a plurality of nucleotides from at least a portion of one or more individual patients genome, wherein the portion of the genome includes at least one genetic element of: a 5′-UTR, a promoter, an enhancer, a silencer, an exon, an intron, a coding sequence, a non-protein coding RNA, a splice acceptor, a splice donor, a branch point site, a 3′-UTR, a Kozak sequence, a poly-A addition site or signal, or a cryptic version thereof, from a known protein coding gene or a regulatory, splicing, or functional element of a non-protein coding RNA gene, and within genes not yet identified in a Dark Matter genome;   identifying a plurality of variants in said nucleotide string by comparing a sequence of said nucleotide string with at least one reference genome, wherein at least one of the plurality of variants is in at least one of: the 5′-UTR, the promoter, the enhancer, the silencer, the exon, the intron, the coding sequence, the non-protein coding RNA, the splice acceptor, the splice donor, the branch point site, the 3′-UTR, the Kozak sequence, the poly-A addition site or signal, or the cryptic version thereof, or the regulatory, splicing, or functional element of the non-protein coding RNA gene;   assigning a similarity score to a genetic element comprising an identified variant of the plurality of variants by executing instructions from an algorithm such as Shapiro-Senapathy algorithm, a MaxEntScan algorithm, and NNSplice algorithm, stored in a memory;   determining a difference in the similarity score of the genetic element with the identified variant from that of the genetic element in the reference sequence;   identifying, for the genetic element with the variant, at least one phenotype such as a disease, a drug response, a therapeutic indication, and a harmful side effect for a medication or substance, based on the pathogenicity or strength alteration of the genetic element with the variant determined using the difference between the similarity score of the genetic element with and without the variant; and,   sending, to a client device, an output comprising said nucleotide string, the identified variant, or the at least one phenotype, in the genetic element for one or more individual patients.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein determining the difference in the strength of the identified variant comprises:
 comparing the strength of the identified variant with a strength of the reference sequence;   determining a deleteriousness or non-deleteriousness of the identified variant; and,   determining a zygosity or mode of inheritance for the identified plurality of variants.   
     
     
         3 . The computer-implemented method of  claim 1 , wherein assigning the similarity score comprises:
 determining the strength of the identified variant based on the similarity score by executing instructions from modifications of an algorithm such as Shapiro-Senapathy algorithm, a MaxEntScan algorithm, and NNSplice algorithm, stored in a memory, based on the length or variability of sequence signals;   
     
     
         4 . The computer-implemented method of  claim 1 , wherein assigning the similarity score comprises:
 determining the similarity score by executing instructions from an algorithm selected from a group consisting of the Shapiro-Senapathy algorithm, the MaxEntScan algorithm, and the NNSplice algorithm, stored in a memory; and   determining a combined score from the group based on their average or differentially weighted scores.   
     
     
         5 . The computer-implemented method of  claim 1 , wherein sending the output comprises displaying, in a graphical user interface of the client device, said nucleotide string, the identified variant, altered strengths or deleteriousness and biological effects or consequences thereof, in the genetic element leading to the at least one phenotype, in sequence view, structure view, or other graphical or tabular representations, in the one or more individual patients. 
     
     
         6 . The computer-implemented method of  claim 1 , further comprising:
 determining a genomic variation such as copy number variants, gene fusion or structural variants affecting one or more genes in one or both alleles, in the plurality of nucleotides;   correlating the affected genes with disease gene panels, therapeutic gene panels, or PGx gene panels, for diagnosing one or more of: disease, therapeutics, immunotherapy, or side effects; and   graphically illustrating the genomic variation or associated mechanism of action.   
     
     
         7 . The computer-implemented method of  claim 1 , further comprising:
 determining that a particular pathogenic coding DNA sequence (CDS) mutation, a regulatory or splicing element mutation, or a genomic structural variation (GSV) mutation in a gene is an indicator of a disease, drug response phenotype, therapeutic indications, or harmful side effects; and,   determining that another pathogenic CDS, regulatory or splicing element or GSV mutation, in the gene, indicates the disease, drug response phenotype, therapeutic indications, or harmful side effects;   
     
     
         8 . The computer-implemented method of  claim 1 , further comprising:
 correlating a number of mismatch repair (MMR) and/or DNA damage repair genes, comprising a deleterious or strength altering mutation, with a tumor mutation burden (TMB) or an microsatellite instability (MSI) status;   predicting a level of the TMB and the MSI status such as low, medium or high, based on the number of MMR and/or DNA damage repair genes with the deleterious or strength altering mutation, and,   indicating or predicting cancers or immunotherapy.   
     
     
         9 . The computer-implemented method of  claim 1 , further comprising:
 identifying a mutation in the plurality of nucleotides in two or more genes, or combinations thereof, specifying cellular features such as cell structure, transcriptional regulation, extracellular matrix, cell proliferation, angiogenesis, or cancer metastasis associated with the disease, as indicative of a gene signature that is predictive of a phenotype;   
     
     
         10 . The computer-implemented method of  claim 1 , further comprising:
 determining a pathogenicity of a mutation in at least one amino acid in a CDS of the known protein coding gene based on various scoring algorithms; and   determining a combined score from scoring algorithms yielding scores above a score threshold based on their average or differentially weighted scores, to obtain a pathogenicity score for the identified variant, indicative of a phenotype.   
     
     
         11 . A computer-implemented method comprising:
 receiving a nucleotide string comprising a plurality of nucleotides from at least a portion of one or more individual patients genome, wherein the portion of the genome includes at least one of: a 5′-UTR, a promoter, an enhancer, a silencer, an exon, an intron, a coding sequence, a splice acceptor, a splice donor, a branch point site, a 3′-UTR, a Kozak sequence, a poly-A addition site or signal, or a cryptic version thereof, from a known protein coding gene or a non-protein coding RNA gene, and within genes not yet identified in a Dark Matter genome;   identifying a variant from the plurality of nucleotides;   determining a pathogenic or strength altering mutation in a gene corresponding to the nucleotide string;   determining, based on the mutation in the gene, various aberrations that occur in gene transcription, splicing, or translation, wherein the various aberrations comprise abolition, increase or decrease of rate or quantity of transcription, exon skipping, partial exon deletion, intron retention, intron skipping, premature termination of translation, polyadenylation, or translation;   determining, for the various aberrations, a mechanism of a causation of an aberration due to the mutation in the gene, gene transcript, or the resulting defective protein, in one or more patients; and   graphically illustrating the various aberrations in a structure or sequence view.   
     
     
         12 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of a mutation in a true acceptor, a true donor, a true branch point site, a true enhancer, or a true silencer in the gene, based on a position and a similarity score of a cryptic acceptor, donor, a branch point site, a enhancer, or a silencer within the exon or intron, as a partial exon deletion, exon skipping, intron retention, or intron skipping.   
     
     
         13 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of a mutation in a cryptic acceptor, donor, branch site, enhancer or silencer in the gene or within the exon or intron, based on a position and a similarity score relative to a true or cryptic acceptor, donor, branch point, enhancer or silencer within the exon or intron, as a partial exon deletion, exon skipping, intron retention, or intron skipping;   
     
     
         14 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of a mutation in a cryptic acceptor, donor, branch site, enhancer or silencer in the gene, as a cryptic exon creation causing intron retention, wherein the cryptic exon score is equal or greater than the score of a true exon bordering it.   
     
     
         15 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of a mutation in a true acceptor, donor, branch point site, enhancer or silencer in a gene, wherein the mutated exon score is lower than the score of the original exon or an adjacent true exon, as an exon skipping.   
     
     
         16 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of a mutation in a true promoter element such as TATA box, CAAT box, GC box or initiator box, a promoter motif, enhancer or silencer within, upstream, or downstream of the gene, leading to pathogenic or strength altering effect, based on a position and a similarity score of a true or cryptic promoter box, motifs, enhancer or silencer, as enhanced or reduced transcription, abolition of transcription, or use of a cryptic element leading to erroneous transcription.   
     
     
         17 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of the mutation in a cryptic promoter element, a cryptic promoter motif, cryptic enhancer or silencer within, upstream, or downstream of the gene, leading to a pathogenic or strength altering effect, based on a position and a similarity score of a true or cryptic promoter box, cryptic motifs, cryptic enhancer or silencer, as enhancement, reduction, or abolition of transcription, or use of a cryptic element leading to erroneous transcription.   
     
     
         18 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of a mutation in a true polyA site or signal, a polyA motif, enhancer or silencer within, upstream, or downstream of the gene, leading to pathogenic or strength altering effect, based on a position and a similarity score of a true or cryptic polyA element, motif, enhancer or silencer, as enhancement, reduction, or abolition of polyadenylation, or use of a cryptic element leading to erroneous polyadenylation.   
     
     
         19 . The computer-implemented method of  claim 11 , wherein determining the mechanism of causation of an aberration comprises:
 determining an effect of a mutation in a cryptic polyA site or signal, a cryptic polyA motif, cryptic enhancer or silencer within, upstream, or downstream of the gene, leading to pathogenic or strength altering effect, based on a position and a similarity score of a true or cryptic polyA element, motif, enhancer or silencer, as enhancement, reduction, or abolition of polyadenylation, or use of a cryptic element leading to erroneous polyadenylation.   
     
     
         20 . The computer-implemented method of  claim 11 , wherein graphically illustrating the different aberrations comprises:
 graphically displaying the mutated gene and showing an effect of splicing aberrations such as exon skipping, partial exon deletion, intron retention, intron skipping, frameshift, premature termination, or amino acid deletion or insertion, in gene or protein structure, sequence views, or in animation;   
     
     
         21 . The computer-implemented method of  claim 11 , wherein graphically illustrating the different aberrations comprises:
 determining that a mutation within an exon, indicative of a coding sequence mutation leading to an amino acid change, is a splicing mutation based on a presence of splicing signals such as donor, acceptor, branch point, enhancers, silencers, or their cryptic versions containing that mutation, their position relative to real or cryptic splicing signals, or a difference between an original similarity score and a mutated similarity score, and, leading to the splicing aberrations; and   graphically displaying a comparison of an effect of a same mutation when considered as an exonic coding region mutation or a splicing mutation.

Join the waitlist — get patent alerts

Track US2022316009A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.