Unified portal for regulatory and splicing elements for genome analysis
Abstract
A method, including identifying, in a nucleotide string, at least two exons, at least one acceptor, at least one donor, and at least one intron between the at least two exons, is provided. The method includes identifying, in the nucleotide string, a cryptic splice site comprising a sequence of nucleotides based on a similarity score with at least one of the acceptor or the donor, and graphically marking, in a display for a user, the nucleotide string at a location indicative of an exon, an intron, a true splice site, and optionally a cryptic splice site when the similarity score is higher than a pre-selected threshold. A system and a non-transitory, computer-readable medium including instructions to cause the system to perform the method are also provided.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method comprising:
identifying, in a nucleotide string, at least two exons, at least one acceptor, at least one donor, and at least one intron between the at least two exons; identifying, in the nucleotide string, a cryptic splice site comprising a sequence of nucleotides based on a similarity score with the at least one of the acceptor or donor; and, graphically marking, in a display for a user, the nucleotide string at a location indicative of the at least two exons, the at least one intron, a true splice site, or a cryptic splice site when the similarity score is higher or lower than, or equal to, a pre-selected threshold.
2 . The computer-implemented method of claim 1 , wherein identifying the at least two exons, at least one acceptor, at least one donor, and at least one intron comprises:
identifying, in the nucleotide string, a first exon that lacks the at least one acceptor and contains the at least one donor; identifying, in the first exon, an open reading frame between an initiator codon for a gene and a first splice junction within the at least one donor; identifying an open reading frame between a subsequent splice junction within the at least one acceptor and at least one donor, as a middle exon for a gene; identifying, in the nucleotide string, a last exon that contains the at least one acceptor and lacks the at least one donor, with an open reading frame between the splice junction within the at least one acceptor and the terminator codon for the gene; identifying, in the nucleotide string, a branch point within the at least one intron, wherein the branch point is associated with a splicing site of the nucleotide string to combine the at least two exons; identifying, in the nucleotide string, a mutation, wherein the mutation comprises a modification in the at least two exons, the at least one intron, the at least one acceptor or the at least one donor, or a branch point, enhancer or silencer; and graphically marking, in the display for the user, the mutation in the nucleotide string, gene structure or sequence view of the display.
3 . The computer-implemented method of claim 1 , further comprising:
identifying, in the nucleotide string, a cryptic branch point, enhancer, or a silencer within the at least two exons or at least one intron, wherein the cryptic branch point, enhancer, or silencer is associated with the splicing site of the nucleotide string to combine the at least two exons.
4 . The computer-implemented method of claim 1 , wherein the similarity scores are computed from one of:
a. determining the similarity score of an element by executing instructions from an algorithm selected from a group consisting of algorithms such as Shapiro-Senapathy algorithm, a MaxEntScan algorithm, and NNSplice algorithm, stored in a memory; b. determining the similarity score of an element by executing instructions from a modified algorithm selected from a group consisting of algorithms such as Shapiro-Senapathy algorithm, a MaxEntScan algorithm, and NNSplice algorithm, stored in a memory, based on characteristics of splicing element sequence signals such as length or variability; c. determining a combined score of the group of algorithms based on corresponding average or differentially weighted scores.
5 . The computer-implemented method of claim 1 , further comprises:
identifying, in the nucleotide string, a cryptic exon that comprises at least one cryptic acceptor and one cryptic donor or an open reading frame, between the cryptic acceptor and the cryptic donor, when a cryptic splice site score is higher or lower than, or equal to, a pre-selected threshold, and a length of the cryptic exon conforms to a pre-selected threshold; identifying a cryptic branch point upstream of the cryptic exon; identifying a cryptic splice enhancer site within, upstream, or downstream of the cryptic exon; and, identifying a cryptic splice silencer site within, upstream, or downstream of the cryptic exon.
6 . The computer-implemented method of claim 1 , further comprising:
identifying, within the at least two exons or the at least one intron, a splice enhancer site comprising a binding site for a spliceosome enhancer factor that promotes a splicing of the at least two exons of a gene, wherein the gene comprises at least a portion of the at least two exons and the at least one intron; and, identifying, within the at least two exons or the at least one intron, the splice silencer site comprising a binding site for an inhibitor factor that suppresses a splicing of at least two exons of the gene.
7 . The computer-implemented method of claim 1 , further comprising:
determining a deleteriousness score of a mutation of the true or the cryptic splice site, a branch point site, enhancers, or silencers, based on variability in the similarity score, relative to a reference sequence.
8 . The computer-implemented method of claim 1 , further comprising:
aligning a plurality of nucleotides from regular, uncommon or unusual regulatory or splicing elements, or cryptic versions thereof in the gene from a plurality of individuals from a same organism to create a recognizable pattern of a position weight matrix (PWM) including invariable or a variable nucleotides of the plurality of nucleotides at a given sequence position of a particular element in a particular gene; determining a distinct PWM or variable sequence signature for a distinct splicing element, or regulatory element within the gene, based on a multiple sequence alignment within the gene or genome sequences of numerous individuals from a species or a group of organisms consisting of similar species; and creating a database of the distinct PWMs or variable sequence signatures for one or more elements from a genome of an organism.
9 . The computer-implemented method of claim 1 , further comprising:
graphically marking, in the display for the user, a true or a cryptic exon, splicing elements or motifs, when the similarity score is higher or lower than, or equal to, a pre-selected threshold, on the tabular, gene structure or sequence view.
10 . The computer-implemented method of claim 1 , further comprising:
determining an exon score as an average of scores or differentially weighted scores of the at least one acceptor and the at least one donor, branch point site, or splicing enhancers, and subtracting the average of the scores or the differentially weighted scores of splicing silencers.
11 . The computer-implemented method of claim 1 , further comprising:
modifying a Shapiro Senapathy algorithm based on a position weight matrix (PWM) or variable sequence signature for a genetic element constructed from mono, di, tri, or longer oligo-nucleotides.
12 . The computer-implemented method of claim 1 , further comprising:
predicting novel genetic elements such as promoters, recognition sequences, binding sites or regulatory and splicing elements, based on the PWM constructed from a multiple sequence alignment of a nucleotide sequence at a specific chromosomal position in genomes from multiple individuals of a same organism, wherein the novel elements show variable nucleotide frequencies that exhibit non-random characteristics indicative of the PWM of genuine functional genetic elements, or statistically distinct characteristics indicative of functional regions, compared to random nucleotide positions.
13 . A computer-implemented method comprising:
identifying, in a nucleotide string, a true gene transcriptional regulatory element such as a promoter sequence comprising at least one of a TATA box, a CAAT box, a GC box or a transcription initiator box, or a transcription termination site; by calculating a similarity score using a position weight matrix (PWM) based on a sequence length or variability of a gene transcriptional regulatory element; associating a score to the gene transcriptional regulatory element based on the corresponding similarity score; and graphically marking using graphical notations, in a display for a user, the nucleotide string at a location indicative of the gene transcriptional regulatory element when the score is higher or lower than, or equal to, a preselected threshold.
14 . The computer-implemented method of claim 13 , further comprising:
identifying a promoter motif that comprises a combination of a promoter elements such as the TATA box, the CAAT box, the GC box, or the transcription initiator box; identifying a cryptic version of a promoter box, motif or other elements such as enhancer or a silencer; determining and assigning a score to the promoter motif based on combined or variably weighted scores of individual elements of the promoter motif, and determining and assigning scores for various genetic element motifs such as splicing motifs (donor, acceptor, exon motifs), initiator motifs, enhancer or silencer motifs, and termination motifs.
15 . The computer-implemented method of claim 13 , further comprising:
identifying mutations within promoter elements or promoter motifs; determining a deleteriousness of the mutations, or alterations in binding strengths of the mutated promoter elements or motifs, based on score differences between normal elements or motifs, and the mutated elements or motifs; identifying the promoter elements or promoter motifs in a nucleotide string by modifying a Shapiro Senapathy algorithm, MaxEntScan algorithm, or NNSplice algorithm based on di, tri, or longer oligo-nucleotides; and determining the deleteriousness of mutations in various genetic motifs such as splicing motifs, initiator motifs, enhancer or silencer motifs, and termination motifs.
16 . The computer-implemented method of claim 13 , further comprising:
assigning a score to a functional genetic element such as a TATA box, CAAT box, GC box, Initiator box, donor, acceptor or polyA site or signal; determining a score for a promoter complex, as an average of scores or differentially weighted scores of the promoter elements and enhancers, and subtracting the average of the scores or the differentially weighted scores of silencer elements; and determining and assigning scores for various genetic regulatory or splicing regulatory or splicing motifs such as splicing motifs comprising donor, acceptor, exon motifs, initiator motifs, enhancer or silencer motifs, or termination motifs.
17 . A computer-implemented method, comprising:
identifying, in a nucleotide string, a poly-A addition site, a signal, enhancer, silencer, or a Kozak sequence, by calculating a similarity score using a position weight matrix (PWM) based on a sequence length and variability of the poly-A addition site, signal, enhancer, silencer, or the Kozak sequence; and associating a score to the respective element based on the similarity score.
18 . The computer-implemented method of claim 17 , further comprising:
identifying, within the nucleotide string, a cryptic poly-A site, signal, enhancer, silencer, a Kozak sequence, a sequence of nucleotides resembling at least one of the true poly-A site, signal, enhancer, silencer, or the Kozak sequence; identifying, in the nucleotide string, a polyA motif comprising a combination of a polyA site, the signal, the enhancer, and the silencer; identifying, in the nucleotide string, a Kozak motif comprising a combination of surrounding elements such as the enhancer, and the silencer; determining and assigning a score to the polyA motif comprising a combination of the polyA site, signal, enhancer, silencer, or Kozak sequence based on combined or variably weighted scores of individual elements of the polyA motif; and, graphically marking, in the display for the user, the polyA elements or motifs, and Kozak elements or motifs.
19 . The computer-implemented method of claim 17 , further comprising:
identifying a mutation of a true or a cryptic translational regulatory element such as the polyA site, a signal, enhancer, silencer, or Kozak sequence, based on a difference between the similarity score of an original element and a mutated element; determining a deleteriousness of the mutation, or alteration in binding strength of a mutated translational regulatory element comprising an element or motif, based on a range of score difference between a normal element or motif and the mutated element or motif; and graphically marking, in the display for the user, the mutation in the nucleotide string, in gene structure or sequence view.
20 . The computer-implemented method of claim 17 , further comprising:
identifying a polyA or Kozak element or motif in a nucleotide string by modifying a Shapiro Senapathy, MaxEntScan, NNSplice or other algorithms based on di, tri, or longer oligo-nucleotides.
21 . The computer-implemented method of claim 17 , further comprising:
assigning a score to a functional genetic element; determining a score for a polyA motif, as an average of scores, or differentially weighted scores of a plurality of polyA elements such as poly-A signals, sites, enhancers, and subtracting the average of the scores or the differentially weighted scores of silencer elements; and, determining a score for a Kozak motif, as an average of scores, or differentially weighted scores of a plurality of Kozak elements including enhancers, and subtracting the average of the scores or the differentially weighted scores of the silencer elements.Join the waitlist — get patent alerts
Track US2022307026A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.