US2023197197A1PendingUtilityA1
Methods and systems for sequence calling
Est. expiryMar 10, 2039(~12.6 yrs left)· nominal 20-yr term from priority
G16B 40/00G16B 30/10G06N 3/08G16B 40/20G16B 5/00G06N 3/045G06N 3/0455G06N 3/09G06N 3/096G06N 3/0464G16B 30/20
74
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The present disclosure provides methods, systems, and media for accurate and efficient estimation of a genome of a genus.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method for estimating at least a part of a genome, the method comprising:
receiving or generating a plurality of actual sequencing signals that represent at least a part of the genome, wherein the plurality of actual sequencing signals is generated by imaging a substrate comprising a plurality of substrate segments as part of a flow sequencing process that uses non-terminated nucleotides; and estimating at least a part of the genome by applying a first trained algorithm to a first set of actual sequencing signals from the plurality of actual sequencing signals, wherein the first set of actual sequencing signals are associated with a first substrate segment of the plurality of substrate segments, and applying a second trained algorithm that differs from the first trained algorithm to a second set of actual sequencing signals from the plurality of actual sequencing signals, wherein the second set of actual sequencing signals are associated with a second substrate segment of the plurality of substrate segments.
22 . The computer-implemented method of claim 21 , wherein the plurality of substrate segments is determined based on expected or actual differences between illumination of substrate segments in the plurality of substrate segments.
23 . The computer-implemented method of claim 21 , wherein the plurality of substrate segments is determined based on expected or actual differences between a plurality of measurements of radiation from the plurality of substrate segments.
24 . The computer-implemented method of claim 21 , wherein the plurality of substrate segments is determined based on expected or actual distribution of samples or sample sources over the plurality of substrate segments.
25 . The computer-implemented method of claim 24 , wherein the plurality of actual sequencing signals is associated with at least one image of at least a portion of the substrate that is linked to multiple samples or sample sources.
26 . The computer-implemented method of claim 24 , wherein the samples or sample sources comprise a plurality of beads, and wherein each bead comprises a clonal population of amplified products of nucleic acids.
27 . The computer-implemented method of claim 21 , wherein substrate segments in the plurality of substrate segments each comprise a same shape and/or a same size.
28 . The computer-implemented method of claim 21 , further comprising identifying the first substrate segment and the second substrate segment from the plurality of substrate segments.
29 . The computer-implemented method of claim 28 , wherein the identifying is performed subsequent to the imaging.
30 . The computer-implemented method of claim 28 , wherein the identifying is performed during the imaging.
31 . The computer-implemented method of claim 21 , wherein:
the first trained algorithm comprises a first mapping between the first set of actual sequencing signals and a first set of trusted reference sequencing signals, wherein the first set of actual sequencing signals and the first set of trusted reference sequencing signals represent parts of a first reference genome; and the second trained algorithm comprises a second mapping between the second set of actual sequencing signals and a second set of trusted reference sequencing signals, wherein the second set of sequencing signals and the second set of trusted reference sequencing signals represent parts of a second reference genome.
32 . The computer-implemented method of claim 31 , wherein the first reference genome is from a first genus, and the second reference genome is from a second genus.
33 . The computer-implemented method of claim 32 , wherein the second trained algorithm is trained on training data generated based, at least in part, on applying the first mapping to the second set of actual sequencing signals.
34 . The computer-implemented method of claim 32 , wherein the first genus differs from the second genus.
35 . The computer-implemented method of claim 32 , wherein the first genus is the same as the second genus.
36 . The computer-implemented method of claim 32 , wherein the first reference genome is smaller than the second reference genome.
37 . The computer-implemented method of claim 31 , wherein:
the second mapping comprises processing the second set of actual sequencing signals to provide a set of accurate sequencing signals; and estimating at least a part of the genome comprises aligning the set of accurate sequencing signals to the second reference genome.
38 . The computer-implemented method of claim 21 , wherein the genome is the human genome.
39 . The computer-implemented method of claim 38 , wherein the genome is from a subject.
40 . The computer-implemented method of claim 21 , wherein estimating at least a part of the genome comprises calculating a confidence level for at least one estimated nucleotide of the genome.
41 . The computer-implemented method of claim 21 , wherein the plurality of substrate segments is grouped into a plurality of groups of substrate segments, wherein the first substrate segment is in a first group and the second substrate segment is in a second group.
42 . The method of claim 41 , wherein:
the first trained algorithm is applied to signals from the first set of actual sequencing signals, wherein the first set of actual sequencing signals is associated with substrate segments in the first group of substrate segments; and the second trained algorithm is applied to signals from the second set of actual sequencing signals, wherein the second set of actual sequencing signals is associated with substrate segment in the second group of substrate segments.
43 . The method of claim 21 , wherein the second set of actual sequencing signals correspond to sequential incorporation of two or more non-terminated nucleotides that are complementary to a homopolymer sequence, and wherein applying the second trained algorithm to the second set of actual sequencing signals allows accurate quantification of homopolymer length.Join the waitlist — get patent alerts
Track US2023197197A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.