System and method for secondary analysis of nucleotide sequencing data
Abstract
Disclosed herein are systems and methods for performing secondary analyses of nucleotide sequencing data in a time-efficient manner. Some embodiments include performing a secondary analysis iteratively while sequence reads are generated by a sequencing system. Secondary analyses can encompass both alignment of sequence reads to a reference sequence (e.g., the human reference genome sequence) and utilization of this alignment to detect differences between a sample and the reference. Secondary analysis can enable detection of genetic differences, variant detection and genotyping, identification of single nucleotide polymorphisms (SNPs), small insertions and deletion (indels) and structural changes in the DNA, such as copy number variants (CNVs) and chromosomal rearrangements.
Claims
exact text as granted — not AI-modified1 .- 27 . (canceled)
28 . A system for sequencing polynucleotides, comprising:
a sequencing apparatus configured to determine the nucleotide sequence of a polynucleotide; and a processor configured to control the sequencing apparatus and to execute instructions that perform a method comprising:
receiving a first nucleotide subsequence of the polynucleotide;
determining whether the first nucleotide subsequence aligns to a reference sequence at a first plurality of candidate locations beyond a threshold confidence level using a first process;
receiving a second nucleotide subsequence of the polynucleotide from the sequencing apparatus, wherein the second nucleotide subsequence comprises the first nucleotide subsequence plus one or more additional nucleotides;
comparing the one or more additional nucleotides in the second nucleotide subsequence to the reference sequence based in part on the first plurality of candidate locations, if the first nucleotide subsequence is aligned to the reference sequence beyond the threshold confidence level, or
repeating the first process by aligning the entire second nucleotide subsequence to the reference sequence if the first nucleotide subsequence is not aligned to the reference sequence beyond the threshold confidence level.
29 . The system of claim 28 , wherein the threshold confidence level depends on a number of mismatches or a probability of a correct match.
30 . The system of claim 28 , wherein the first nucleotide subsequence is one or more nucleotides in length.
31 . The system of claim 28 , wherein the second nucleotide subsequence is one or more nucleotides in length.
32 . The system of claim 28 , wherein comparing the one or more additional nucleotides in the second nucleotide subsequence to the reference sequence comprises a simple alignment process, the simple alignment process being more computationally efficient than the first process in memory usage or the number of computation operations.
33 . The system of claim 32 , wherein the processor is further configured to determine a simple alignment score based on the simple alignment process.
34 . The system of claim 28 , wherein the processor is further configured to store data corresponding to at least one of the first plurality of candidate locations if the first nucleotide subsequence is aligned to the reference sequence.
35 . The system of claim 28 , wherein the processor is further configured to store data corresponding to at least one of a second plurality of candidate locations resulting from comparing the second nucleotide subsequence to the reference sequence.
36 . The system of claim 28 , wherein comparing the one or more additional nucleotides in the second nucleotide subsequence to the reference sequence comprises comparing the second nucleotide subsequence with corresponding sequences of the second nucleotide subsequence on the reference sequence based on the first plurality of candidate locations.
37 . The system of claim 36 , wherein the processor is further configured to determine a mapping quality (MapQ) score for each of the second plurality of candidate locations.
38 . The system of claim 28 , wherein determining whether the first nucleotide subsequence aligns to the reference sequence is initiated before the sequencing reactions are completed.
39 . The system of claim 28 , wherein the processor is further configured to perform variant calling for the first nucleotide subsequence or the second nucleotide subsequence.
40 . The system of claim 39 , wherein performing the variant calling comprises:
performing variant calling using a first variant calling process or a second, simple, variant calling process, wherein the second variant calling process is more computationally efficient than the first variant calling process in variant calling of the second nucleotide subsequence.
41 . The system of claim 39 , wherein the variant calling is performed using the output of the first process or the process used to compare the one or more additional nucleotides in the second nucleotide subsequence to the reference sequence, based on a variant calling metric.
42 . The system of claim 41 , wherein the variant calling metric is determined based on a number of different base types called at a position of the reference sequence.
43 . The system of claim 28 , wherein comparing the one or more additional nucleotides in the second nucleotide subsequence to the reference sequence is initiated before the sequencing reactions are completed.
44 . The system of claim 28 , wherein the sequencing apparatus implements sequencing-by-synthesis.
45 . A computer-implemented method for efficient sequencing of polynucleotides, comprising:
receiving a first nucleotide subsequence of a read from a sequencing apparatus during a sequencing run of the first nucleotide subsequence; performing a secondary analysis of the first nucleotide subsequence of the read based on a reference sequence using a first process or a second process, wherein the first nucleotide subsequence comprises one or more additional nucleotides compared to a previous iteration, wherein the second process is more computationally efficient than the first process in performing the secondary analysis, wherein the first process aligns the entire first nucleotide subsequence to the reference sequence, wherein the second process aligns the one or more additional nucleotides to the reference sequence based in part on results from the previous iteration, and wherein the secondary analysis comprises: comparing the first nucleotide subsequence to the reference sequence to determine a first subsequence of the reference sequence that has a high degree of similarity to the first nucleotide subsequence; and determining if the sequencing apparatus should generate additional nucleotide reads.
46 . The method of claim 45 , wherein performing the secondary analysis comprises processing the first nucleotide subsequence to determine a first plurality of candidate locations of the read that align to the reference sequence using:
the first process if the read is not aligned to the reference sequence in the previous iteration, the second process if otherwise, wherein the second process is more computationally efficient than the first process to determine the first plurality of candidate locations of the read.
47 . The method of claim 46 , wherein performing the secondary analysis of the first nucleotide subsequence using the second process comprises performing a simple alignment to determine a simple alignment score.
48 . The method of claim 46 , wherein results of the secondary analysis comprises output of the first process, or output of the second process.
49 . The method of claim 45 , wherein performing the secondary analysis comprises performing variant calling of the first nucleotide subsequence, comprising:
performing variant calling on the output of the first process or the second process using a first variant calling process or a second variant calling process, wherein the second variant calling process is more computationally efficient than the first variant calling process in variant calling of the first nucleotide subsequence.
50 . The method of claim 49 , wherein results of the secondary analysis comprises output of the first variant calling process, output of the second variant calling process.
51 . The method of claim 45 , further comprising providing a user with results of the secondary analysis during the sequencing run.
52 . The method of claim 51 , wherein the results of the secondary analysis are provided to the user at fixed intervals.
53 . The method of claim 51 , wherein the results of the secondary analysis are provided to the user at request of the user.
54 . The method of claim 45 , wherein performing the secondary analysis is based on whether the first nucleotide subsequence aligns to the reference sequence beyond a threshold confidence in the previous iteration.
55 . A computer readable recording medium having recorded a program for implementing in a computer the functions of a system according to claim 28 .
56 . A computer readable recording medium having recorded a program that causes a computer to execute a method according to claim 45 .Join the waitlist — get patent alerts
Track US2023410945A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.