US2015302144A1PendingUtilityA1

Hierarchical genome assembly method using single long insert library

Assignee: PACIFIC BIOSCIENCES CALIFORNIAPriority: Jul 13, 2012Filed: May 19, 2015Published: Oct 22, 2015
Est. expiryJul 13, 2032(~5.9 yrs left)· nominal 20-yr term from priority
G06F 19/18G06F 19/22G16B 20/50G16B 30/20G16B 20/20G16B 30/10G16B 30/00G16B 20/00
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention is generally directed to a hierarchical genome assembly process for producing high-quality de novo genome assemblies. The method utilizes a single, long-insert, shotgun DNA library in conjunction with Single Molecule, Real-Time (SMRT®) DNA sequencing, and obviates the need for additional sample preparation and sequencing data sets required for previously described hybrid assembly strategies. Efficient de novo assembly from genomic DNA to a finished genome sequence is demonstrated for several microorganisms using as little as three SMRT® cells, and for bacterial artificial chromosomes (BACs) using sequencing data from just one SMRT® Cell. Part of this new assembly workflow is a new consensus algorithm which takes advantage of SMRT® sequencing primary quality values, to produce a highly accurate de novo genome sequence, exceeding 99.999% (QV 50) accuracy. The methods are typically performed on a computer and comprise an algorithm that constructs sequence alignment graphs from pairwise alignment of sequence reads to a common reference.

Claims

exact text as granted — not AI-modified
What is claimed: 
     
         1 . A computer-implemented method to determine a consensus sequence from a set of polynucleotide sequence reads without using a previously known reference sequence, the method comprising:
 a) providing a set of polynucleotide sequence reads that comprise errors introduced by a sequencing reaction, wherein said polynucleotide sequence reads in the set comprise overlapping polynucleotide sequences that are alignable to each other;   b) choosing a seed read from the set of polynucleotide sequence reads;   c) performing pairwise alignment of all other polynucleotide sequence reads in the set to the seed read to generate a set of sequence alignments;   d) constructing a multiple sequence alignment from the set of sequence alignments, wherein the errors in the set of polynucleotide sequence reads are present in the resulting multiple sequence alignment and further wherein the multiple sequence alignment is constructed without the use of a previously known reference sequence;   e) performing an error correction step on the multiple sequence alignment by applying a consensus algorithm to the multiple sequence alignment, wherein the consensus algorithm reduces the number of errors in the seed sequence using information within the multiple sequence alignment and generates a consensus sequence for the set of polynucleotide sequence reads, thereby determine a consensus sequence from a set of polynucleotide sequence reads.   
     
     
         2 . The method of  claim 1 , wherein the set of polynucleotide sequence reads comprises raw sequencing data. 
     
     
         3 . The method of  claim 1 , wherein the seed read is greater than 1000 base pairs in length. 
     
     
         4 . The method of  claim 1 , wherein the seed read has a length between 1000 and 10,000 base pairs. 
     
     
         5 . The method of  claim 1 , wherein the seed read is at least 2, 3, 5, 10, 15, or 20 kb in length. 
     
     
         6 . The method of  claim 1 , wherein the seed read has an accuracy of less than 90%. 
     
     
         7 . The method of  claim 1 , further comprising normalizing the set of sequence alignments prior to constructing the multiple sequence alignment. 
     
     
         8 . The method of  claim 1 , wherein the consensus algorithm uses a dynamic programming process to generate the consensus sequence. 
     
     
         9 . The method of  claim 1 , wherein the seed read is generated using a single-molecule sequencing technology. 
     
     
         10 . The method of  claim 1 , wherein each read in the set of polynucleotide sequence reads contains at least a portion of a region of interest. 
     
     
         11 . The method of  claim 1 , wherein the set of sequence alignments is generated by a method comprising a pairwise local alignment algorithm. 
     
     
         12 . The method of  claim 1 , wherein the consensus algorithm is performed iteratively, with each iteration reducing the number of errors in a resulting consensus sequence. 
     
     
         13 . The method of  claim 12 , wherein at the end of each iteration, the resulting consensus sequence is used as the seed read for the subsequent iteration. 
     
     
         14 . The method of  claim 1 , wherein the polynucleotide sequence reads are genomic DNA sequence reads. 
     
     
         15 . The method of  claim 1 , wherein the polynucleotide sequence reads comprise replicate sequence information. 
     
     
         16 . A method of determining a consensus sequences for a region of interest, the method comprising:
 a) providing a mixed population of nucleic acid sequence reads from the region of interest;   b) choosing a sequence read from the mixed population of nucleic acid sequence reads as a seed sequence;   c) aligning the nucleic acid sequence reads to the seed sequence to generate a set of sequence alignments;   d) constructing a multiple sequence alignment using the set of sequence alignments; and   e) based upon the multiple sequence alignment, determining a consensus sequence for the mixed population without the use of a reference sequence, thereby determining a consensus sequence for the region of interest.   
     
     
         17 . The method of  claim 16 , wherein the aligning comprises subjecting the set of sequence alignments to normalization prior to constructing the multiple sequence alignment. 
     
     
         18 . The method of  claim 17 , wherein the normalization comprises at least one of the group consisting of: changing mismatches to indels and moving gaps to right-most equivalent positions. 
     
     
         19 . The method of  claim 16 , wherein said constructing the multiple sequence alignment comprises constructing a multigraph and merging nodes in the multigraph. 
     
     
         20 . The method of  claim 16 , wherein the seed sequence is at least 2, 3, 5, 10, 15, or 20 kb in length.

Join the waitlist — get patent alerts

Track US2015302144A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.