US2010153400A1PendingUtilityA1
Systems and methods for rational selection of context sequences and sequence templates
Est. expiryAug 21, 2027(~1.1 yrs left)· nominal 20-yr term from priority
Inventors:Yoav Shalom Namir
G16B 40/30G16B 30/00G16B 50/10G16B 20/30G16B 20/20G16B 40/00G06F 16/285G06F 16/20G06F 16/2246G16B 50/00G16B 20/00
59
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Provided are systems and methods for rational selection of context sequences and sequence templates including a computer implemented method for obtaining a repository of attributes sets where the attributes sets are statistically associated with a sequence template representing two or more context sequences.
Claims
exact text as granted — not AI-modified1 . A computer implemented method for obtaining a repository of attributes sets, wherein attributes sets are statistically associated with a sequence template representing two or more context sequences, comprising:
(a) obtaining a dataset of context sequences; (b) transforming each context sequence to a sequence template, thereby obtaining a dataset of sequence templates; (c) clustering said dataset of sequence templates into a plurality of clusters according to a distance formula; wherein at least one cluster is statistically associated with at least one attributes set; and (d) inserting into said repository each of said clusters and said attributes set which is statistically associated with said each of said clusters.
2 . The computer implemented of claim 1 , wherein said dataset of context sequences of step (a) is further subjected to multiple sequence alignment.
3 . A repository obtained by the computer implemented method of claim 1 .
4 . A computer implemented method for identifying a sequence template as statistically associated with an attributes set of interest, comprising:
(a) providing a repository of attributes sets; wherein attributes sets are statistically associated with a sequence template representing two or more context sequences; (b) selecting an attributes set; and (c) retrieving at least one sequence template statistically associated with said attributes set.
5 . The computer implemented method of claim 4 , further comprising the step of merging at least two of said retrieved sequence templates.
6 . The computer implemented method of claim 4 , said attributes are selected from the group consisting of the Gene Ontology Project (GO), Interpro annotation (European Molecular Biology Laboratory, EMBL), SMART (a Simple Modular Architecture Research Tool, found at http://smart.embl.de/), UniProt Knowledgebase (SwissProt), OMIM (by NCBI) PROSITE (by the Swiss Institute of Bioinformatics), Protein Information Resource (PIR), GeneCards, and Kyoto Encyclopedia of Genes and Genomes (KEGG).
7 . A method of preparing a polynucleotide construct, comprising:
(a) identifying a sequence template as statistically associated with an attributes set of interest according to the method of claim 4 ; and (b) preparing a polynucleotide construct having at least one portion operably linked to a context sequence; wherein said context sequence is characterized as having either 80%-85%, 85%-90%, or 90%-100% homology with said sequence template.
8 . The method of claim 7 , wherein the preparing comprises synthesizing said context sequence.
9 . The method of claim 7 , wherein the preparing comprises constructing an expression vector comprising said context sequence.
10 . The method of claim 7 , wherein the preparing comprises constructing a probe comprising said context sequence.
11 . A computer memory system comprising a plurality of tree topologies representing plurality of (k) heaps, wherein the plurality of tree topologies is managed through a common interface; and (k≧1).
12 . The computer memory system of claim 11 , wherein said heaps are min heaps.
13 . The computer memory system of claim 11 , wherein said heaps are max heaps.
14 . The computer memory system of claim 11 , wherein an active subset of heaps is held in Random Access Memory (RAM), while the rest of said heaps are maintained on a secondary storage.
15 . A computer implemented method for clustering a plurality of polynucleotide sequences, comprising:
(a) determining an attributes set for the plurality of polynucleotide sequences; and (b) clustering said polynucleotide sequences into a plurality of clusters according to values of said attributes set.
16 . A computerized system configured for identifying a sequence template as statistically associated with an attributes set of interest, the computerized system comprising: context sequence clustering module, configured to cluster said sequences into a plurality of clusters; and an enrichment analysis module, configured to provide enrichment appraisal, wherein context sequence clustering module being communicatively coupled to the enrichment analysis module.Join the waitlist — get patent alerts
Track US2010153400A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.