US2022115090A1PendingUtilityA1
Systems and methods for nucleic acid sequence assembly
Est. expiryJun 26, 2034(~7.9 yrs left)· nominal 20-yr term from priority
G16B 30/00G16B 30/10G16B 45/00G16B 30/20
69
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods, processes, and particularly computer implemented processes and computer program products are provided for use in the analysis of genetic sequence data. The processes and products are employed in the assembly of shorter nucleic acid sequence data into longer linked and preferably contiguous genetic constructs, including large contigs, chromosomes and whole genomes.
Claims
exact text as granted — not AI-modified1 . A sequencing method of assembling nucleic acid sequence reads comprising:
at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:
obtaining a plurality of sequence reads derived from a larger contiguous nucleic acid, wherein two or more sequence reads derived from a common fragment of the larger contiguous nucleic acid comprise a common barcode sequence,
identifying a first subset of sequence reads in the plurality of sequence reads that comprise both overlapping sequences and a common barcode sequence; and
aligning the first subset of sequence reads to provide a contiguous linear nucleic acid sequence.
2 . The method of claim 1 , further comprising repeating the identifying and aligning steps with a plurality of different subsets of sequence reads to provide a plurality of contiguous linear nucleic acid sequences.
3 . The method of claim 2 , further comprising ordering the plurality of different contiguous linear nucleic acid sequences in a sequence context within the larger contiguous nucleic acid.
4 . The method of claim 3 , wherein the ordering comprises mapping the plurality of different contiguous linear nucleic acid sequences against a reference sequence.
5 . The method of claim 3 , wherein the ordering comprises:
identifying one or more sequence reads that comprise a barcode sequence common to a first contiguous linear nucleic acid sequence, but include overlapping sequences with a second contiguous linear nucleic acid sequence; and identifying the first and second contiguous linear nucleic acid as structurally linked.
6 . A method of assembling nucleic acid sequence reads into larger contiguous sequences, comprising:
at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors: obtaining a plurality of sequence reads derived from a larger contiguous nucleic acid, identifying a first subsequence from a set of overlapping sequence reads in the plurality of sequence reads; extending the first subsequence to one or more adjacent or overlapping sequences based upon the presence of a barcode sequence on the adjacent sequence that is common to the first subsequence; and providing a linear nucleic acid sequence that comprises the first subsequence and the one or more adjacent sequences.
7 . A sequencing method comprising, at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors:
(A) obtaining a plurality of sequence reads, wherein
the plurality of sequence reads comprises a plurality of sets of sequence reads,
each respective sequence read in a set of sequence reads includes (i) a first portion that corresponds to a subset of a larger contiguous nucleic acid and (ii) a common second portion that forms an identifier that is independent of the sequence of the larger contiguous nucleic acid and that identifies a partition, in a plurality of partitions, in which the respective sequence read was formed, and
each respective set of sequence reads in the plurality of sets of sequence reads is formed in a partition in the plurality of partitions and each partition includes one or more fragments of the larger contiguous nucleic acid that is used as the template for each respective sequence read in the partition;
(B) creating a respective set of k-mers for each sequence read in the plurality of sequence reads, wherein
the sets of k-mers collectively comprise a plurality of k-mers,
the identifiers of the sequence reads for each k-mer in the plurality of k-mers is retained,
k is less than the average length of the sequence reads in the plurality of sequence reads, and
each respective set of k-mers includes at least eighty percent of the possible k-mers of length k of the first portion of the corresponding sequence read;
(C) tracking, for each respective k-mer in the plurality of k-mers, an identity of each sequence read in the plurality of sequence reads that contains the respective k-mer and the identifier of the set of sequence reads that contains the sequence read; (D) graphing the plurality k-mers as a graph comprising a plurality of nodes connected by a plurality of directed arcs, wherein
each node comprises an uninterrupted set of k-mers in the plurality of k-mers of length k with k−1 overlap,
each arc connects an origin node to a destination node in the plurality of nodes,
a final k-mer of an origin node has k−1 overlap with an initial k-mer of a destination node, and
a first origin node has a first directed arc with both a first destination node and a second destination node in the plurality of nodes; and
(E) determining whether to merge the origin node with the first destination node or the second destination node in order to derive a contig sequence that is more likely to be representative of a portion of the larger contiguous nucleic acid, wherein the contig sequence comprises (i) the origin node and (ii) one of the first destination node and the second destination node, wherein the determining uses at least the identifiers of the sequence reads for k-mers in the first origin node, the first destination node, and the second destination node.
8 . The sequencing method of claim 7 , wherein,
the first origin node and the first destination node are part of a first path in the graph that includes one or more additional nodes other than the origin node and the first destination node, the first origin node and the second destination node are part of a second path in the graph that includes one or more additional nodes other than the origin node and the second destination node, and the determining (E) comprises determining whether the first path is more likely representative of the larger contiguous nucleic acid than the second path by evaluating a number of identifiers shared between the k-mers of the nodes of a first portion of the first path and the k-mers of the nodes of a second portion of the first path versus a number of identifiers shared between the k-mers of the nodes of a first portion of the second path and the k-mers of the nodes of a second portion of the second path.
9 . The sequencing method of claim 7 , wherein
the first origin node and the first destination node are part of a first path in the graph that includes one or more additional nodes other than the origin node and the first destination node, the first origin node and the second destination node are part of a second path in the graph that includes one or more additional nodes other than the origin node and the second destination node, and the determining (E) up-weights the first path relative to the second path when the first path has higher average coverage than the second path.
10 . The sequencing method of claim 7 , wherein the
the first origin node and the first destination node are part of a first path in the graph that includes one or more additional nodes other than the origin node and the first destination node, the first origin node and the second destination node are part of a second path in the graph that includes one or more additional nodes other than the origin node and the second destination node, and the determining (E) up-weights the first path relative to the second path when the first path represents a longer contiguous portion of the larger contiguous nucleic acid sequence than the second path.
11 . The sequencing method of claim 7 , wherein a first k-mer in the first node is present in a sub-plurality of the plurality of sequence reads and the identity of each sequence read in the sub-plurality of sequence reads is retained for the first k-mer and used by the determining (E) to determine whether the first path is more likely representative of the larger contiguous nucleic acid sequence than the second path.
12 . The sequencing method of claim 7 , wherein
a partition in the plurality of partitions comprises at least 1000 molecules with the common second portion, and each molecule in the at least 1000 molecules further comprises a primer sequence complementary to at least a portion of the larger contiguous nucleic acid.
13 . The sequencing method of claim 7 , wherein
a partition in the plurality of partitions comprises at least 1000 molecules with the common second portion, and each molecule in the at least 1000 molecules further comprises a primer site and a semi-random N-mer priming sequence that is complementary to part of the larger contiguous nucleic acid.
14 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions are greater than 50 kilobases in length.
15 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions are between 20 kilobases and 200 kilobases in length.
16 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions consists of between 1 and 500 different fragments of the larger contiguous nucleic acid.
17 . The sequencing method of claim 7 , wherein the one or more fragments of the larger contiguous nucleic acid in a partition in the plurality of partitions consists of between 5 and 100 fragments of the larger contiguous nucleic acid.
18 . The sequencing method of claim 7 , wherein the plurality of sequence reads is obtained from less than 5 nanograms of nucleic or ribonucleic acid.
19 . The sequencing method of claim 7 , wherein the identifier in the second portion of each respective sequence read in the set of sequence reads encodes a common value selected from the set {1, . . . , 1024}, the set {1, . . . , 4096}, the set {1, . . . , 16384}, the set {1, . . . , 65536}, the set {1, . . . , 262144}, the set {1, . . . , 1048576}, the set {1, . . . , 4194304}, the set {1, . . . , 16777216}, the set {1, . . . , 67108864}, or the set {1, . . . , 1×10 12 }.
20 . The sequencing method of claim 7 , wherein the identifier is an N-mer, and N is an integer selected from the set {4, . . . , 20}.
21 .- 44 . (canceled)Join the waitlist — get patent alerts
Track US2022115090A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.