Systems and methods for off-target sequence detection
Abstract
A computer-implemented method, computer system and computer-readable medium for identifying off-target matches from a set of candidate primer sequences on a genome reference sequence can include: receiving onto a data storage unit a plurality of candidate primer sequences; for each candidate primer sequence, calculating, using a processor, a plurality of candidate matches on the genome reference sequence for the candidate primer sequences; calculating, using the processor, verified matches on the genome reference sequence based on the candidate matching locations satisfying a plurality of matching verification rules; performing matching calculations of the verified matches to determine whether the verified matches form a match condition on the genome reference sequence; and generating a location profile on the genome reference sequence based on the match condition from the verified matches that meet a predetermined threshold.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A computer-implemented method of identifying off-target matches from a set of candidate primer sequences on a genome reference sequence, the method comprising:
receiving onto a data storage unit a plurality of candidate primer sequences; decomposing each candidate primer sequence into k-mers comprising substrings or subsequences of length k; for each candidate primer sequence, identifying a plurality of candidate matches on the genome reference sequence for the candidate primer sequences, wherein identifying the plurality of candidate matches comprises searching a k-mer index of the genome reference sequence for occurrences of each k-mer; calculating verified matches on the genome reference sequence based on the candidate matches satisfying a plurality of matching verification rules, wherein calculating the verified matches comprises reusing one or more rule satisfaction calculations for k-mers when present in a cache; performing matching calculations to determine off-target match conditions for one or more of the verified matches with respect to the genome reference sequence; and selecting, in response to determining off-target match conditions, candidate primer sequences that do not form an off-target match on the genome reference sequence as acceptable primer sequences for use in a multiplex PCR assay.
22 . The method of claim 21 , further comprising updating the cache with one or more respective rule satisfaction calculations for respective k-mers based on the respective k-mers satisfying one or more matching verification rules.
23 . The method of claim 22 , wherein the one or more matching verification rules comprise a rule that there are at least k consecutive matching nucleotides.
24 . The method of claim 22 , wherein the one or more matching verification rules comprise a rule that a threshold number of nucleotides mismatches is not exceeded.
25 . The method of claim 24 , wherein the threshold number is a function of the length of the respective candidate primer sequence.
26 . The method of claim 22 , wherein the one or more matching verification rules comprise a rule that a threshold number of nucleotide mismatches on an end of the respective candidate primer sequence is not exceeded.
27 . The method of claim 22 , wherein identifying the plurality of candidate matches comprises reusing one or more rule satisfaction calculations for k-mers when present in the cache.
28 . The method of claim 21 , further comprising generating a location profile on the genome reference sequence based on the off-target match condition from the verified matches that meet a predetermined off-target threshold.
29 . The method of claim 21 , wherein:
the cache is organized based on different sequence lengths.
30 . The method of claim 21 , wherein the candidate matches comprise a candidate matching location of the corresponding candidate primer sequence on the genome reference sequence, and
wherein the verified matches comprise a verified matching location of the verified matches on the genome reference sequence.
31 . The method of claim 21 , wherein the plurality of candidate primer sequences comprises a first candidate primer sequence, the method further comprising:
generating a prediction of a number of matches on the reference genome sequence for the first candidate primer sequence by comparing k-mers derived from the first candidate primer sequence to the k-mer index of the genome reference sequence; responsive to determining that the predicted number of matches exceeds a threshold, discarding the first candidate primer sequence from further evaluation.
32 . The method of claim 21 further comprising:
clustering a plurality of sequences corresponding to the verified matches into sequence proximity groupings; and
using the sequence proximity groupings to identify the off-target match condition.
33 . The method of claim 32 wherein candidate primer sequences that form an off-target match cause amplification of non-target sequences or interference with amplification of target locations.
34 . A computing system for identifying off-target matches from a set of candidate primer sequences on a genome reference sequence, the computing system comprising:
at least one processor; and a memory storing instructions that, when executed by the at least one processor, causes the computing system to perform:
receiving onto a data storage unit a plurality of candidate primer sequences;
decomposing each candidate primer sequence into k-mers comprising sub strings or subsequences of length k;
for each candidate primer sequence, identifying a plurality of candidate matches on the genome reference sequence for the candidate primer sequences, wherein identifying the plurality of candidate matches comprises searching a k-mer index of the genome reference sequence for occurrences of each k-mer;
calculating verified matches on the genome reference sequence based on the candidate matches satisfying a plurality of matching verification rules, wherein calculating the verified matches comprises reusing one or more rule satisfaction calculations for k-mers when present in a cache;
performing matching calculations to determine off-target match conditions for one or more of the verified matches with respect to the genome reference sequence; and
selecting, in response to determining off-target match conditions, candidate primer sequences that do not form an off-target match on the genome reference sequence as acceptable primer sequences for use in a multiplex PCR assay.
35 . The computing system of claim 34 , wherein:
the cache is coupled to the memory, and wherein the cache stores rule satisfaction calculations separately for different parameters of a respective rule.
36 . The computing system of claim 34 wherein the memory further stores instructions that when executed by the at least one processor causes the computer system to perform skipping a rule satisfaction calculation for at least one of the candidate primer sequences based on rule satisfaction calculations stored in the cache.
37 . The computing system of claim 34 wherein:
the cache is organized based on different sequence lengths.
38 . The computing system of claim 34 wherein the plurality of matching verification rules comprise a rule that there are at least k consecutive matching nucleotides.
39 . The computing system of claim 34 wherein the plurality of matching verification rules comprise a rule that a threshold number of nucleotide mismatches is not exceeded.
40 . The computing system of claim 34 , wherein the plurality of matching verification rules comprise a rule that a threshold number of nucleotide mismatches on an end of the respective candidate primer sequence is not exceeded.
41 . A non-transitory computer-readable storage medium for identifying off-target matches from a set of candidate primer sequences on a genome reference sequence comprising computer-executable instructions that when executed cause a computing system to:
receive onto a data storage unit a plurality of candidate primer sequences; decompose each candidate primer sequence into k-mers comprising substrings or sub sequences of length k; for each candidate primer sequence, identify a plurality of candidate matches on the genome reference sequence for the candidate primer sequences, wherein identifying the plurality of candidate matches comprises searching a k-mer index of the genome reference sequence for occurrences of each k-mer; calculate verified matches on the genome reference sequence based on the candidate matches satisfying a plurality of matching verification rules, wherein calculating the verified matches comprises reusing one or more rule satisfaction calculations for k-mers when present in a cache; perform matching calculations to determine off-target match conditions for one or more of the verified matches with respect to the genome reference sequence; and selecting, in response to determining off-target match conditions, candidate primer sequences that do not form an off-target match on the genome reference sequence as acceptable primer sequences for use in a multiplex PCR assay.Join the waitlist — get patent alerts
Track US2021233612A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.