Data analysis of dna sequences
Abstract
Systems and methods for data analysis are provided. In one embodiment, a method for analysis is provided, including electronically receiving sequence data; electronically receiving one or more reference data sequences related to at least an expression vector; associating the sequence data with at least one of the reference data sequences to identify a transgene flanking sequence; searching a genome for one or more insertion sites of the transgene flanking sequence; and annotating the genome and the one or more insertion sites within the genome when one or more insertion sites are found in said searching step.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for analysis, comprising:
electronically receiving sequence data; electronically receiving one or more reference data sequences related to at least an expression vector; associating the sequence data with at least one of the reference data sequences to identify a transgene flanking sequence; searching a genome for one or more insertion sites of the transgene flanking sequence; and annotating the genome and the one or more insertion sites within the genome when one or more insertion sites are found in said searching step.
2 . The method of claim 1 , wherein the reference data is further related to at least one of a left cloning vector, a primer, an adapter, and a right cloning vector.
3 . The method of claim 1 , wherein the reference data is further related to a left cloning vector, a primer, an adapter, and a right cloning vector.
4 . The method of claim 1 , further comprising:
searching the sequence data for a first reference data sequence; and searching the sequence data for a second reference data sequence when said first reference data sequence is located.
5 . The method of claim 4 , wherein the first reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector.
6 . The method of claim 5 , wherein the second reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector, the second reference data sequence being selected independently of the first reference data sequence.
7 . The method of claim 4 , wherein the first reference data sequence is an expression vector and the second reference data sequence is an adapter.
8 . The method of claim 4 , wherein the first and second reference data sequences are independently selected from the group consisting of: a primer and an adapter.
9 . The method of claim 1 , further comprising visualizing the transgene flanking sequence and the reference data.
10 . The method of claim 1 , further comprising visualizing the one or more insertion sites within the genome.
11 . The method of claim 1 , further comprising characterizing sequence information of the genome upstream and downstream of the insertion site.
12 . The method of claim 11 , wherein sequence information of the genome 10 kilobase pairs upstream and 10 kilobase pairs downstream of the insertion site are characterized.
13 . The method of claim 1 , further comprising:
aligning the sequence data with one or more of the reference data sequences; and conducting a qualitative analysis of the aligned sequences.
14 . The method of claim 1 , further comprising:
aligning the sequence data with one or more of the reference data sequences; and conducting a quantitative analysis of the aligned sequences.
15 . The method of claim 1 , wherein the genome is at least a portion of a plant genome.
16 . The method of claim 1 , wherein associating the sequence data with at least one of the reference data sequences includes using an algorithm to match at least one of the reference data sequences against the sequence data.
17 . The method of claim 16 , wherein the algorithm is a LASTZ algorithm.
18 . The method of claim 1 , wherein searching a genome for one or more insertion sites of the transgene flanking sequence includes using an algorithm to locate sequences upstream and downstream of the at least one insertion site with the genome.
19 . The method of claim 18 , wherein the algorithm is a BLAST algorithm.
20 . A system for analysis, comprising:
a module for receiving sequence data related to a sequence; a module for receiving one or more reference sequences related to at least an expression vector; and a calculation module operable to:
associate the sequence data with at least one of the reference data sequences to identify a transgene flanking sequence;
search a genome for one or more insertion sites of the transgene flanking sequence; and
annotate the genome and the one or more insertion sites within the genome. when the one or more insertion site is found.
21 . The system of claim 20 , wherein the reference sequences are further related to at least one of a left cloning vector, a primer, an adapter, and a right cloning vector.
22 . The system of claim 20 , wherein the reference sequences are further related to a left cloning vector, a primer, an adapter, and a right cloning vector.
23 . The system of claim 20 , wherein said computation module is further operable to:
search the sequence data for a first reference data sequence; and search the sequence data for a second reference data sequence when said first reference data sequence is located.
24 . The system of claim 23 , wherein the first reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector.
25 . The system of claim 24 , wherein the second reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector, the second reference data sequence being selected independently of the first reference data sequence.
26 . The system of claim 23 , wherein the first reference data sequence is an expression vector and the second reference data sequence is an adapter.
27 . The system of claim 23 , wherein the first and second reference data sequences are independently selected from the group consisting of: a primer and an adapter.
28 . The system of claim 20 , further comprising a module for visualizing the transgene flanking sequence and at least one of the left cloning vector, the expression vector, the primer, the adapter, and the right cloning vector.
29 . The system of claim 20 , further comprising a module for visualizing the one or more insertion sites within the genome.
30 . The system of claim 20 , wherein said computation module is further operable to characterize sequence information of the genome upstream and downstream of the insertion site.
31 . The system of claim 30 , wherein said computation module is operable to characterize sequence information of the genome 10 kilobase pairs upstream and 10 kilobase pairs downstream of the insertion site.
32 . The system of claim 20 , wherein said computation module is operable to:
align the sequence data with one or more of the reference data sequences; and conduct a qualitative analysis of the aligned sequences.
33 . The system of claim 20 , wherein said computation module is operable to:
align the sequence data with one or more of the reference data sequences; and conduct a quantitative analysis of the aligned sequences.
34 . The system of claim 20 , wherein the genome is at least a portion of a plant genome.
35 . The system of claim 20 , wherein associating the sequence data with at least one of the reference data sequences includes using an algorithm to match at least one of the reference data sequences against the sequence data.
36 . The system of claim 35 , wherein the algorithm is a LASTZ algorithm.
37 . The system of claim 20 , wherein searching a genome for one or more insertion sites of the transgene flanking sequence includes using an algorithm to locate sequences upstream and downstream of the at least one insertion site with the genome.
38 . The system of claim 37 , wherein the algorithm is a BLAST algorithm.Join the waitlist — get patent alerts
Track US2013211729A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.