US2022310275A1PendingUtilityA1

Unified portal for regulatory and splicing elements for genome analysis

Assignee: GENOME INT CORPORATIONPriority: Mar 26, 2021Filed: Mar 18, 2022Published: Sep 29, 2022
Est. expiryMar 26, 2041(~14.6 yrs left)· nominal 20-yr term from priority
G16B 20/00A61K 31/7088G16B 30/00G16B 20/20C12N 2320/33G16B 40/20G16H 70/60C12N 15/113G16B 30/10G16B 45/00
72
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method, including identifying, in a nucleotide string, at least two exons, at least one acceptor, at least one donor, and at least one intron between the at least two exons, is provided. The method includes identifying, in the nucleotide string, a cryptic splice site comprising a sequence of nucleotides based on a similarity score with at least one of the acceptor or the donor, and graphically marking, in a display for a user, the nucleotide string at a location indicative of an exon, an intron, a true splice site, and optionally a cryptic splice site when the similarity score is higher than a pre-selected threshold. A system and a non-transitory, computer-readable medium including instructions to cause the system to perform the method are also provided.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method comprising:
 receiving a nucleotide string comprising a plurality of nucleotides from at least a portion of two or more individuals' genome, wherein the portion of the genome includes at least one genetic element of: a 5′-UTR, a promoter, an enhancer, a silencer, an exon, an intron, a coding sequence, a non-protein coding RNA, a splice acceptor, a splice donor, a branch point site, a 3′-UTR, a Kozak sequence, a poly-A addition site or signal, or a cryptic version thereof, from a known protein coding gene or a regulatory, splicing, or functional element of a non-protein coding RNA gene, and within genes not yet identified in a Dark Matter genome;   identifying, in the nucleotide string based on a chromosomal position, a genetic element such as a coding element, exon, intron, 5′-UTR, 3′-UTR, promoter, a splice acceptor, a splice donor, a branch point site, a Kozak sequence, a poly-A addition site or signal, an enhancer, or a silencer, or their cryptic version thereof, from a known protein coding gene or a regulatory, splicing, or functional element of the non-protein coding RNA gene; and,   determining a variable sequence signature or position weight matrix (PWM) for a particular genetic element of a particular gene, based on a multiple sequence alignment of the same element at the same chromosomal or genomic position within a gene or in the genome sequences of one or more individuals from a same species.   
     
     
         2 . The computer-implemented method of  claim 1 , further comprising:
 determining a variable sequence signature or position weight matrix (PWM) based on di, tri, or longer oligonucleotides, for a particular genetic element of a particular gene, based on a multiple sequence alignment of the same element at the same chromosomal or genomic position within a gene or in the genome sequences of one or more individuals from a same species.   
     
     
         3 . The computer-implemented method of  claim 1 , further comprising:
 predicting novel genetic elements such as promoters, recognition sequences, binding sites or regulatory and splicing elements throughout the genome, based on the PWM constructed from a multiple sequence alignment of a nucleotide sequence at a particular chromosomal position in genomes from multiple individuals of a same species or organism, wherein the novel elements show variable nucleotide frequencies that exhibit non-random characteristics indicative of the PWM of genuine structural or functional genetic elements, or statistically distinct characteristics indicative of functional regions, compared to random nucleotide positions.   
     
     
         4 . The computer-implemented method of  claim 1 , further comprising:
 predicting novel genetic elements such as promoters, recognition sequences, binding sites or other regulatory and splicing elements, based on the PWMs constructed from a set length identified from a multiple sequence alignment of a nucleotide sequence at a chromosomal position in the genomes from multiple individuals of a same species or organism, at every consecutive position in a genome, wherein, the PWMs show a variable nucleotide frequencies that exhibit non-random characteristics, typical of PWMs of genuine functional genetic elements, or other statistically distinct characteristics indicative of functional regions, compared to random nucleotide positions.   
     
     
         5 . The computer-implemented method of  claim 1 , further comprising:
 modifying a Shapiro Senapathy algorithm, a MaxEntScan algorithm, a NNSplice algorithm, or any algorithm for identifying the genetic elements based on the PWM or variable sequence signature for the genetic element constructed from mono, di, tri, or longer oligo-nucleotides;   assigning a score to the structural or functional genetic element; and,   identifying deleterious or strength altering mutations in the functional genetic element based on similarity scores calculated from a modified algorithm such as a Shapiro Senapathy algorithm, MaxEntScan algorithm, NNSplice algorithm, or any algorithm based on di, tri, or longer oligo-nucleotides.   
     
     
         6 . The computer-implemented method of  claim 1 , further comprising:
 aligning the plurality of nucleotides from unknown regions such as multiple protein binding promoter sites, polyA sites, splice sites upstream of, downstream of, or within or around the gene from a number of individuals from a same species or organism, so as to create a recognizable pattern of the PWM consisting of invariable or variable nucleotides that are not randomly distributed at a given position as having structural, functional or biological implications of gene regulatory or splicing elements.   
     
     
         7 . A computer-implemented method comprising:
 identifying, in a nucleotide string, based on a chromosomal position, a non-coding RNA gene such as a miRNA, tRNA, rRNA, or snoRNA;   determining, based on a similarity score using a prediction algorithm, a genetic element comprising a regulatory, splicing, or a functional RNA element of the non-coding RNA gene;   identifying, a difference between the first similarity score of a normal genetic element and the second similarity score of a mutated genetic element of the non-coding RNA gene;   determining the causality of a phenotype by a sequence variant based on the difference between the first similarity score and the second similarity score; and,   graphically marking, in a display for a user, the nucleotide string at a location indicative of an exon, an intron, regulatory, splicing or functional RNA element when the first similarity score or the second similarity score is higher or lower than, or equal to, a pre-selected threshold on a gene structure or sequence view.   
     
     
         8 . The computer-implemented method of  claim 7 , further comprising,
 identifying, in the nucleotide string, a positive signature, and a negative signature from an allowable mono, di, tri or longer oligo-nucleotide and a disallowed mono, di, tri or longer oligo-nucleotide;   graphically marking a mutation of the nucleotide string on the positive signature and the negative signature; and,   determining a deleterious effect of the mutation based on whether the mutation occurs within the positive signature or the negative signature.   
     
     
         9 . The computer-implemented method of  claim 7 , further comprising,
 displaying a recognition sequence, regulatory, splicing, or processed functional element on the gene structure;   depicting the processing steps of the non-coding RNA gene into an active element;   elaborating the processing steps in the gene structure or sequence view;   indicating mutations and the processing steps at which a processing error occurs; and,   elaborating on a mechanism of aberrations within the ncRNA gene causing a biological or clinical phenotype.   
     
     
         10 . The computer-implemented method of  claim 7 , further comprising:
 constructing a position weight matrix (PWM) for regulatory, or splicing elements, and recognition sequences for the processing of a non-coding RNA gene; and   constructing the PWM for a processed functional non-coding RNA gene product, for an individual type of the non-coding RNA gene such as the miRNA, tRNA or rRNA; and   constructing the PWM for the non-coding RNA gene, by aligning the nucleotide sequences of a particular non-coding RNA gene from a number of individuals of the same organism at a particular chromosomal position.   
     
     
         11 . The computer-implemented method of  claim 7 , further comprising:
 constructing a variable sequence signature based on a number or frequency of variable mono, di, tri or longer oligonucleotides at each position of the genetic element from aligned sequences; and,   determining a deleteriousness score of a mutation of the non-coding RNA gene, or regulatory, splicing, recognition sequence elements, or a processed functional non-coding RNA product, based on the difference between the first similarity score of a normal genetic element and the second similarity score of the mutated element, calculated from the position weight matrix (PWM).   
     
     
         12 . The computer-implemented method of  claim 7 , further comprising:
 determining the similarity score by executing instructions from an algorithm selected from a group consisting of algorithms such as Shapiro-Senapathy algorithm, a MaxEntScan algorithm, and NNSplice algorithm, stored in a memory;   determining the similarity score by executing instructions from a modified algorithm selected from a group consisting of algorithms such as Shapiro-Senapathy algorithm, a MaxEntScan algorithm, and NNSplice algorithm, stored in a memory, based on characteristics of splicing element sequence signals such as length or variability; and,   determining a combined score of the group of algorithms based on corresponding average or differentially weighted scores.   
     
     
         13 . The computer-implemented method of  claim 7 , further comprising:
 determining and graphically marking variable nucleotides by stacking a non-redundant mono, di, tri or longer oligo-nucleotides at each position of an additional nucleotide string from a multiple sequence alignment of multiple non-coding RNA genes of the same type; and,   applying a position weight matrix (PWM) methodology using these oligo-nucleotides for detecting the multiple non-coding RNA genes and corresponding genetic elements.   
     
     
         14 . A computer-implemented method comprising:
 identifying a first amino acid string corresponding to a functional protein or a protein domain;   aligning said first amino acid string with at least one additional amino acid string that encodes a functional variant of said functional protein;   identifying, at each amino acid position within said additional amino acid string, multiple variable amino acids that appear in the at least one additional amino acid string for each aligned location in the first amino acid string; and   graphically marking, in a display for a user, a variable amino acid as an allowable amino acid at an aligned location in said first amino acid string.   
     
     
         15 . The computer-implemented method of  claim 14 , further comprising:
 identifying an amino acid that is different from an allowable amino acid as a disallowed amino acid at the aligned location;   graphically stacking a non-redundant disallowed amino acid as a variable amino acid at each position of the additional amino acid string in the functional protein; and   graphically distinguishing, in the display for a user, an allowed amino acid and the disallowed amino acid at each aligned location.   
     
     
         16 . The computer-implemented method of  claim 14 , further comprising:
 distinguishing allowed variable amino acids of a protein or a domain as a positive signature, and disallowed variable amino acids of the protein or the domain as a negative signature;   determining a deleterious effect of a mutation based on whether the mutation occurs within the positive signature or negative signature; and   graphically marking a mutation on the positive signature or the negative signature.   
     
     
         17 . The computer-implemented method of  claim 14 , further comprising:
 graphically indicating a hydropathy value of each variable amino acid at each aligned location;—   determining a hydropathy value based on the average of hydropathy values of each of the variable amino acids at each location;   determining a hydropathy value for a region of amino acids based on the average of hydropathy values at each amino acid position in a given amino acid sequence region;   determining a normal hydropathy signature of a protein domain based on the hydropathy value of an allowed amino acid;   determining a mutated hydropathy signature of a sequence portion of a protein domain based on the hydropathy value of a mutated amino acid; and   determining a deleteriousness score for the mutation based on a difference between the normal hydropathy signature and the mutated hydropathy signature, or an average hydropathy index of the plurality of variable amino acids of a mutated position before and after mutation.   
     
     
         18 . The computer-implemented method of  claim 14 , further comprising:
 correlating an invariance or a degree of variance of an amino acid position with a deleteriousness of a mutation;   indicating that the mutation at an invariant amino acid position is deleterious, wherein decreasing deleteriousness is correlated with increasing amino acid variability; and   applying the correlated invariance or degree of variance to determine the deleteriousness of the mutation.   
     
     
         19 . The computer-implemented method of  claim 14 , further comprising:
 constructing an allowable and a non-allowable variable amino acid sequence signature based on variable amino acid strings of a protein or domain sequence at the same chromosomal position from different individuals of a same organism;   determining a frequency of each allowable amino acid at every position that occurs across the different individuals;   defining an algorithm based on the frequencies of different amino acids to assign scores for individual allowable amino acids at each position;   determining the deleteriousness of a variable amino acid at a position based on the frequency of an allowable amino acid, with deleteriousness decreasing with increasing variability score.   
     
     
         20 . The computer-implemented method of  claim 14 , further comprising,
 aligning a genome sequence from an individual of an organism with the genome of another individual of the same organism, at the same chromosomal or genomic position, to construct variable mono, di, tri or oligonucleotides, or mono, di, tri or oligo amino acids; and   predicting, based on aligning the genome sequence of multiple individuals of the same organism, regulatory elements, splicing elements, variable amino acids, domains, proteins, genes, exons, introns, or intergenic regions, throughout the genome.   
     
     
         21 . The computer-implemented method of  claim 14 , further comprising:
 constructing a variable amino acid sequence signature based on variable amino acid strings of the plurality of variable amino acids representing a possible domain or portion of the possible domain, in open reading frames (ORFs), exonic, intronic, intergenic, or genic regions throughout the genome, at the same chromosomal or genomic position, from different individuals of a same organism; and   identifying new genes, exons, introns, coding sequence, regulatory or splicing elements, domains, or proteins, protein coding genes or non coding RNA genes, based on portions of variable amino acid sequence signature by comparing with genes predicted within the genome, employing gene prediction programs using various parameters;   
     
     
         22 . The computer-implemented method of  claim 14 , further comprising:
 constructing a variable amino acid sequence signature based on variable amino acid strings of the plurality of variable amino acids representing a possible domain or portion of the possible domain, in open reading frames (ORFs), exonic, intronic, intergenic, or genic regions of the genome, at the same chromosomal or genomic position, from different individuals of a same organism; and   discovering new domains by the presence of the variable amino acid sequence signature similar to and characteristic of variable sequence signatures of genuine domains, by searching in all six reading frames of a nucleotide sequence throughout the genome from different individuals of the same organism.

Join the waitlist — get patent alerts

Track US2022310275A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.