US2013268207A1PendingUtilityA1
Systems and methods for identifying somatic mutations
Est. expiryApr 9, 2032(~5.7 yrs left)· nominal 20-yr term from priority
G16B 30/00G16B 20/20G16B 30/10G06F 19/22
67
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and method for identifying somatic mutations can receive first ans second sequence information, determine if a variant present in the first sequencing information is also present in the second sequence information, and identify variants present in the first sequence information are somatic mutations when the variant is either not present in the second sequence information or the presence of the variant in the second sequence information is likely due to a sequencing error.
Claims
exact text as granted — not AI-modified1 . A method of identifying a somatic mutation, comprising:
identifying a variant in a first sequence information; determining if the variant is present in a second sequence information; calculating, when the variant is present in the second sequence information, a likelihood that the variant is present in the second sequence information above an expected error rate; and identifying the variant as a somatic mutation when the likelihood is below a threshold.
2 . The method of claim 1 , further comprising determining a coverage level within the second sequence information for a position corresponding to the variant when the variant is not present in the second sequence information, and identifying the variant as a somatic mutation when the coverage level not less than a threshold.
3 . The method of claim 2 , further comprising identifying the variant as a non-confident somatic mutation when the coverage level is less than a threshold.
4 . The method of claim 1 , further comprising identifying the variant as a non-confident somatic mutation when a coverage level within the first sequence information for a position corresponding to the variant is less than a threshold.
5 . The method of claim 1 , further comprising calculating a somatic call quality value based on the likelihood that the variant is present in the second sequence information above the expected error rate and a likelihood that the variant is present in the first sequence information above a second expected error rate.
6 . A computer program product, comprising a computer-readable storage medium whose contents include a program with instructions to be executed on a processor, the instructions comprising:
instructions to identify a variant in a first sequence information; instructions to determine if the variant is present in a second sequence information; instructions to calculate, when the variant is present in the second sequence information, a likelihood that the variant is present in the second sequence information above an expected error rate; and instructions to identify the variant as a somatic mutation when the likelihood is below a threshold.
7 . The method of claim 6 , further comprising instructions to determine a coverage level within the second sequence information for a position corresponding to the variant when the variant is not present in the second sequence information, and identifying the variant as a somatic mutation when the coverage level not less than a threshold.
8 . The method of claim 7 , further comprising instructions to identify the variant as a non-confident somatic mutation when the coverage level is less than a threshold.
9 . The method of claim 6 , further comprising instructions to identify the variant as a non-confident somatic mutation when a coverage level within the first sequence information for a position corresponding to the variant is less than a threshold.
10 . The method of claim 6 , further comprising instructions to calculate a somatic call quality value based on the likelihood that the variant is present in the second sequence information above the expected error rate and a likelihood that the variant is present in the first sequence information above a second expected error rate.
11 . A system for identifying a somatic mutation, comprising:
a processor configured to:
receive instructions to perform a paired analysis on a first sequence information and a second sequence information, the first sequence information corresponding to a first sample and the second sequence information corresponding to a second sample; and
provide an indication of a sequence difference between the first sequence information and the second sequence information and a somatic call quality value indicative of a relative confidence that the difference is indicative of a variation between the first sample and the second sample, the variation between the first sample and the second sample being representative of a somatic mutation.
12 . The system of claim 11 , wherein the first sample is a tumor sample and the second sample is a non-tumor sample.
13 . The system of claim 11 , wherein the processor is further configured to receive the first sequence information
14 . The system of claim 11 , wherein the processor is further configured to receive information
15 . The system of claim 11 , wherein the processor is further configured to identify variants in the first sequence information including variants found at a low frequency.
16 . The system of claim 11 , wherein the processor is further configured to identify variants in the second sequence information with a low stringency.
17 . The system of claim 11 , where the processor is further configured to identify a variant in a first sequence information, and determine if the variant is present in a second sequence information.
18 . The system of claim 11 , where the processor is further configured to identify the variant as a somatic mutation when the variant is not present in the second sequence information.
19 . The system of claim 11 , where the processor is further configured to identify the variant as a somatic mutation when the variant is present in the second sequence information at a rate that is likely due to an expected error.
20 . The system of claim 11 , where the processor is further configured to calculate the somatic call quality value based on the likelihood that the variant is present in the second sequence information above the expected error rate and a likelihood that the variant is present in the first sequence information above a second expected error rate.
21 . (canceled)
22 . (canceled)Join the waitlist — get patent alerts
Track US2013268207A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.