US2013211729A1PendingUtilityA1

Data analysis of dna sequences

Assignee: DOW AGROSCIENCES LLCPriority: Feb 8, 2012Filed: Feb 7, 2013Published: Aug 15, 2013
Est. expiryFeb 8, 2032(~5.5 yrs left)· nominal 20-yr term from priority
G16B 50/10G16B 30/10G16B 50/00G16B 30/00G06F 19/28
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for data analysis are provided. In one embodiment, a method for analysis is provided, including electronically receiving sequence data; electronically receiving one or more reference data sequences related to at least an expression vector; associating the sequence data with at least one of the reference data sequences to identify a transgene flanking sequence; searching a genome for one or more insertion sites of the transgene flanking sequence; and annotating the genome and the one or more insertion sites within the genome when one or more insertion sites are found in said searching step.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for analysis, comprising:
 electronically receiving sequence data;   electronically receiving one or more reference data sequences related to at least an expression vector;   associating the sequence data with at least one of the reference data sequences to identify a transgene flanking sequence;   searching a genome for one or more insertion sites of the transgene flanking sequence; and   annotating the genome and the one or more insertion sites within the genome when one or more insertion sites are found in said searching step.   
     
     
         2 . The method of  claim 1 , wherein the reference data is further related to at least one of a left cloning vector, a primer, an adapter, and a right cloning vector. 
     
     
         3 . The method of  claim 1 , wherein the reference data is further related to a left cloning vector, a primer, an adapter, and a right cloning vector. 
     
     
         4 . The method of  claim 1 , further comprising:
 searching the sequence data for a first reference data sequence; and   searching the sequence data for a second reference data sequence when said first reference data sequence is located.   
     
     
         5 . The method of  claim 4 , wherein the first reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector. 
     
     
         6 . The method of  claim 5 , wherein the second reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector, the second reference data sequence being selected independently of the first reference data sequence. 
     
     
         7 . The method of  claim 4 , wherein the first reference data sequence is an expression vector and the second reference data sequence is an adapter. 
     
     
         8 . The method of  claim 4 , wherein the first and second reference data sequences are independently selected from the group consisting of: a primer and an adapter. 
     
     
         9 . The method of  claim 1 , further comprising visualizing the transgene flanking sequence and the reference data. 
     
     
         10 . The method of  claim 1 , further comprising visualizing the one or more insertion sites within the genome. 
     
     
         11 . The method of  claim 1 , further comprising characterizing sequence information of the genome upstream and downstream of the insertion site. 
     
     
         12 . The method of  claim 11 , wherein sequence information of the genome 10 kilobase pairs upstream and 10 kilobase pairs downstream of the insertion site are characterized. 
     
     
         13 . The method of  claim 1 , further comprising:
 aligning the sequence data with one or more of the reference data sequences; and   conducting a qualitative analysis of the aligned sequences.   
     
     
         14 . The method of  claim 1 , further comprising:
 aligning the sequence data with one or more of the reference data sequences; and   conducting a quantitative analysis of the aligned sequences.   
     
     
         15 . The method of  claim 1 , wherein the genome is at least a portion of a plant genome. 
     
     
         16 . The method of  claim 1 , wherein associating the sequence data with at least one of the reference data sequences includes using an algorithm to match at least one of the reference data sequences against the sequence data. 
     
     
         17 . The method of  claim 16 , wherein the algorithm is a LASTZ algorithm. 
     
     
         18 . The method of  claim 1 , wherein searching a genome for one or more insertion sites of the transgene flanking sequence includes using an algorithm to locate sequences upstream and downstream of the at least one insertion site with the genome. 
     
     
         19 . The method of  claim 18 , wherein the algorithm is a BLAST algorithm. 
     
     
         20 . A system for analysis, comprising:
 a module for receiving sequence data related to a sequence;   a module for receiving one or more reference sequences related to at least an expression vector; and   a calculation module operable to:
 associate the sequence data with at least one of the reference data sequences to identify a transgene flanking sequence; 
 search a genome for one or more insertion sites of the transgene flanking sequence; and 
 annotate the genome and the one or more insertion sites within the genome. when the one or more insertion site is found. 
   
     
     
         21 . The system of  claim 20 , wherein the reference sequences are further related to at least one of a left cloning vector, a primer, an adapter, and a right cloning vector. 
     
     
         22 . The system of  claim 20 , wherein the reference sequences are further related to a left cloning vector, a primer, an adapter, and a right cloning vector. 
     
     
         23 . The system of  claim 20 , wherein said computation module is further operable to:
 search the sequence data for a first reference data sequence; and   search the sequence data for a second reference data sequence when said first reference data sequence is located.   
     
     
         24 . The system of  claim 23 , wherein the first reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector. 
     
     
         25 . The system of  claim 24 , wherein the second reference data sequence is selected from the group consisting of: an expression vector, an adapter, a primer, and a cloning vector, the second reference data sequence being selected independently of the first reference data sequence. 
     
     
         26 . The system of  claim 23 , wherein the first reference data sequence is an expression vector and the second reference data sequence is an adapter. 
     
     
         27 . The system of  claim 23 , wherein the first and second reference data sequences are independently selected from the group consisting of: a primer and an adapter. 
     
     
         28 . The system of  claim 20 , further comprising a module for visualizing the transgene flanking sequence and at least one of the left cloning vector, the expression vector, the primer, the adapter, and the right cloning vector. 
     
     
         29 . The system of  claim 20 , further comprising a module for visualizing the one or more insertion sites within the genome. 
     
     
         30 . The system of  claim 20 , wherein said computation module is further operable to characterize sequence information of the genome upstream and downstream of the insertion site. 
     
     
         31 . The system of  claim 30 , wherein said computation module is operable to characterize sequence information of the genome 10 kilobase pairs upstream and 10 kilobase pairs downstream of the insertion site. 
     
     
         32 . The system of  claim 20 , wherein said computation module is operable to:
 align the sequence data with one or more of the reference data sequences; and   conduct a qualitative analysis of the aligned sequences.   
     
     
         33 . The system of  claim 20 , wherein said computation module is operable to:
 align the sequence data with one or more of the reference data sequences; and   conduct a quantitative analysis of the aligned sequences.   
     
     
         34 . The system of  claim 20 , wherein the genome is at least a portion of a plant genome. 
     
     
         35 . The system of  claim 20 , wherein associating the sequence data with at least one of the reference data sequences includes using an algorithm to match at least one of the reference data sequences against the sequence data. 
     
     
         36 . The system of  claim 35 , wherein the algorithm is a LASTZ algorithm. 
     
     
         37 . The system of  claim 20 , wherein searching a genome for one or more insertion sites of the transgene flanking sequence includes using an algorithm to locate sequences upstream and downstream of the at least one insertion site with the genome. 
     
     
         38 . The system of  claim 37 , wherein the algorithm is a BLAST algorithm.

Join the waitlist — get patent alerts

Track US2013211729A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.