Methods for multi-resolution analysis of cell-free nucleic acids
Abstract
The present disclosure provides a method for enriching for multiple genomic regions using a first bait set that selectively hybridizes to a first set of genomic regions of a nucleic acid sample and a second bait set that selectively hybridizes to a second set of genomic regions of the nucleic acid sample. These bait set panels can selectively enrich for one or more nucleosome-associated regions of a genome, said nucleosome-associated regions comprising genomic regions having one or more genomic base positions with differential nucleosomal occupancy, wherein the differential nucleosomal occupancy is characteristic of a cell or tissue type of origin or disease state.
Claims
exact text as granted — not AI-modified1 .- 33 . (canceled)
34 . A computer system for classifying a candidate insertion or deletion (indel) detected in a plurality of sequence reads as a true indel or an indel not in a subject, the plurality of sequence reads generated from cell-free deoxyribonucleic acid (cfDNA) molecules in a bodily sample of the subject, the system comprising:
a processor programmed to:
access the plurality of sequence reads;
map the plurality of sequence reads to a reference genome associated with the cfDNA molecules;
determine a plurality of families of sequence reads based on the mapped plurality of sequence reads;
detect the candidate indel in the plurality of sequence reads;
determine a plurality of model parameters that each affects a classification of whether the candidate indel is a true indel;
generate, via a multi-parameter likelihood function, a first multi-parameter likelihood function value based on the plurality of model parameters and the plurality of families of sequence reads, wherein the first multi-parameter likelihood function value represents a probability that the candidate indel is a true indel;
generate, via the multi-parameter likelihood function, a second multi-parameter likelihood function value based on the plurality of model parameters and the plurality of families of sequence reads, wherein the second multi-parameter likelihood function value represents a probability that the candidate indel is an indel not in a subject; and
classify the candidate indel as a true indel or an indel not in the subject based on the first multi-parameter likelihood function value and the second multi-parameter likelihood function value.
35 . The system of claim 34 , wherein to classify the candidate indel as a true indel or an indel not in the subject, the processor is further programmed to:
determine a log likelihood ratio (LLR) based on the first multi-parameter likelihood function value and the second multi-parameter likelihood function value; and compare the LLR to a predetermined threshold value, wherein the candidate indel is classified based on the comparison.
36 . The system of claim 35 , wherein the processor is further programmed to:
classify the candidate indel (i) as a true indel if the LLR is greater than the predetermined threshold value; or (ii) as an indel not in the subject if the LLR is less than the predetermined threshold value.
37 . The system of claim 34 , wherein the plurality of model parameters comprise:
a first set of the plurality of model parameters for the first multi-parameter likelihood function value; and a second set of the plurality of model parameters for the second multi-parameter likelihood function value.
38 . The system of claim 37 , wherein the first and/or second set of the plurality of model parameters comprises two or more of:
a. a frequency of the variant allele in the plurality of cfDNA molecules, b. a frequency of non-reference alleles other than the variant allele in the plurality of cfDNA molecules, c. a frequency of an indel error affecting all the forward strands of a family of strands, d. a frequency of an indel error affecting all the entire reverse strands of a family of strands, and e. a frequency of an indel error in a sequence read.
39 . The system of claim 37 , wherein the first and/or second set of the plurality of model parameters comprises:
a. a frequency of the variant allele in the plurality of sequence reads, and b. a frequency of an indel error in a sequence read.
40 . The system of claim 39 , wherein the first and/or second set of plurality of model parameters comprises:
a. a frequency of an indel error affecting all the forward strands of a family of strands, and b. a frequency of an indel error affecting all the reverse strands of a family of strands.
41 . The system of claim 34 , wherein a family of sequence reads of the plurality of families of sequence reads comprises sequence reads derived from a single cfDNA molecule.
42 . The system of claim 41 , wherein a sequence read within the family of sequence reads is classified into either one of:
a. the sequence read derived from a forward strand of the cfDNA molecule; and b. the sequence read derived from a reverse strand of the cfDNA molecule.
43 . The system of claim 41 , wherein a plurality of sequence reads within a family is classified into either one of:
a. the plurality of sequence reads derived from a forward strand of the cfDNA molecule; and b. the plurality of sequence reads derived from a reverse strand of the cfDNA molecule.
44 . The system of claim 38 , wherein the frequency of the variant allele in the plurality of cfDNA molecules for the second set of the plurality of model parameters is zero.
45 . The system of claim 34 , wherein for each family of the plurality of families:
generate a first configuration of reads based on a first number of reads in the family having a variant allele that includes the candidate indel; generate a second configuration of reads based on a second number of reads in the family having a non-reference allele other than the variant allele; and generate a third configuration of reads based on a third number of reads in the third family having a reference allele that appears in the genome sequence, wherein the multi-parameter likelihood function value is based on the first configuration of reads of the plurality of families, the second configuration of reads of the plurality of families, and the third configuration of reads of the plurality of families.
46 . The system of claim 45 , wherein the processor is further programmed to:
generate, based on at least one or more of the plurality of model parameters, an overall configuration of the plurality of sequence reads, the overall configuration omitting a value for the frequency of the variant allele in the plurality of sequence reads, wherein the multi-parameter likelihood function value is generated based on the overall configuration.
47 . The system of claim 45 , wherein the processor is further programmed to:
for each family of the plurality of families:
generate forward strand data based on a number of reads of a forward strand of the family having the candidate indel; and
generate reverse strand data based on a number of reads of a reverse strand of the family having the candidate indel,
wherein the multi-parameter likelihood function value is based further on the forward strand data of the plurality of families and the reverse strand data of the plurality of families.
48 . The system of claim 47 , wherein the processor is further programmed to:
generate, based on at least one or more of the plurality of model parameters, an overall configuration of the plurality of sequence reads, the overall configuration omitting a value for the frequency of the variant allele in the plurality of sequence reads, wherein the multi-parameter likelihood function value is generated based on the overall configuration.
49 . The system of claim 45 , wherein the processor is further programmed to:
for each family of the plurality of families:
generate first read data based on a number of sequence reads of the family that supports the variant allele;
generate second read data based on a number of the sequence reads of the family that supports the non-reference allele other than the variant allele; and
generate third read data based on a number of the sequence reads of the family that supports the reference allele that appears in the reference genome,
wherein the multi-parameter likelihood function value is based on the first read data of the plurality of families, the second read data of the plurality of families, and the third read data of the plurality of families.
50 . The system of claim 45 , wherein the set of plurality of model parameters comprises one or more of:
a frequency of the variant allele in the plurality of sequence reads, a frequency of non-reference alleles other than the variant allele in the plurality of sequence reads, a frequency of a frequency of an indel error affecting all the forward strands of a family of strands, a frequency of an indel error affecting all the reverse strands of a family of strands, and a frequency of an indel error in a sequence read.
51 . The system of claim 45 , wherein the multi-parameter likelihood function comprises a Nelder-Mead function.
52 . The system of claim 34 , wherein the processor is further programmed to:
classify the candidate indel as a true indel; and correlate the true indel with a mutation associated with a disease.
53 . The system of claim 52 , wherein the processor is further programmed to:
determine that the mutation is at least a partial cause of the disease.
54 . The system of claim 52 , wherein the processor is further programmed to:
determine that the mutation is reflects a response to treatment of the disease.
55 . The system of claim 34 , wherein the indel not in the subject is caused by an error in sequencing at a plurality of genomic base locations.
56 . The system of claim 34 , wherein the indel not in the subject is caused by an error in amplification at a plurality of genomic base locations.
57 . The system of claim 34 , wherein to map the plurality of sequence reads to the reference genome, the processor is further programmed to:
align the plurality of sequence reads with the reference genome.
58 . The system of claim 34 , wherein to map the plurality of sequence reads to the reference genome, the processor is further programmed to:
map the plurality of sequence reads based on at least one or more respective molecular barcode sequences attached to the plurality of sequence reads.
59 . A computer-implemented method of classifying a candidate insertion or deletion (indel) detected in a plurality of sequence reads as a true indel or an indel not in a subject, the plurality of sequence reads generated from cell-free deoxyribonucleic acid (cfDNA) molecules in a bodily sample of the subject, the method comprising:
accessing, by a processor, the plurality of sequence reads; mapping, by the processor, the plurality of sequence reads to a reference genome associated with the cfDNA molecules; determining, by the processor, a plurality of families of sequence reads based on the mapped plurality of sequence reads; detecting, by the processor, the candidate indel in the plurality of sequence reads; determining, by the processor, a plurality of model parameters that each affects a classification of whether the candidate indel is a true indel; generating, by the processor, via a multi-parameter likelihood function, a multi-parameter likelihood function value based on the plurality of model parameters, wherein the multi-parameter likelihood function value represents a probability that the candidate indel is a true indel; generating, by the processor, via the multi-parameter likelihood function, a second multi-parameter likelihood function value based on the plurality of model parameters and the plurality of families of sequence reads, wherein the second multi-parameter likelihood function value represents a probability that the candidate indel is an indel not in a subject; and classifying, by the processor, the candidate indel as a true indel or an indel not in the subject based on the first multi-parameter likelihood function value and the second multi-parameter likelihood function value.
60 . The method of claim 59 , wherein classifying the candidate indel as a true indel or an indel not in the subject comprises:
determining a log likelihood ratio (LLR) based on the first multi-parameter likelihood function value and the second multi-parameter likelihood function value; and comparing the LLR to a predetermined threshold value, wherein the candidate indel is classified based on the comparison.
61 . The method of claim 60 , further comprising:
classifying the candidate indel (i) as a true indel if the LLR is greater than the predetermined threshold value; or (ii) as an indel not in the subject if the LLR is less than the predetermined threshold value.
62 . The method of claim 59 , wherein the plurality of model parameters comprise:
a first set of the plurality of model parameters for the first multi-parameter likelihood function value; and a second set of the plurality of model parameters for the second multi-parameter likelihood function value.
63 . The method of claim 62 , wherein the first and/or second set of the plurality of model parameters comprises two or more of:
a. a frequency of the variant allele in the plurality of cfDNA molecules, b. a frequency of non-reference alleles other than the variant allele in the plurality of cfDNA molecules, c. a frequency of an indel error affecting all the forward strands of a family of strands, d. a frequency of an indel error affecting all the entire reverse strands of a family of strands, and e. a frequency of an indel error in a sequence read.
64 . The method of claim 62 , wherein the first and/or second set of plurality of model parameters comprises:
a. a frequency of the variant allele in the plurality of sequence reads, and b. a frequency of an indel error in a sequence read.
65 . The method of claim 64 , wherein the first and/or second set of plurality of model parameters comprises:
a. a frequency of an indel error affecting all the forward strands of a family of strands, and b. a frequency of an indel error affecting all the reverse strands of a family of strands.
66 . The method of claim 59 , wherein a family of sequence reads of the plurality of families of sequence reads comprises sequence reads derived from a single cfDNA molecule.
67 . The method of claim 66 , wherein a sequence read within the family of sequence reads is classified into either one of:
a. the sequence read derived from a forward strand of the cfDNA molecule; and b. the sequence read derived from a reverse strand of the cfDNA molecule.
68 . The method of claim 66 , wherein a plurality of sequence reads within a family is classified into either one of:
a. the plurality of sequence reads derived from a forward strand of the cfDNA molecule; and b. the plurality of sequence reads derived from a reverse strand of the cfDNA molecule.
69 . The method of claim 68 , wherein the frequency of the variant allele in the plurality of cfDNA molecules for the second multi-parameter likelihood function value is zero.
70 . The method of claim 59 , wherein generating the multi-parameter likelihood function value comprises:
for each family of the plurality of families:
generating a first configuration of reads based on a first number of reads in the family having a variant allele that includes the candidate indel;
generating a second configuration of reads based on a second number of reads in the family having another non-reference allele other than the variant allele; and
generating a third configuration of reads based on a third number of reads in the family having a reference allele that appears in the genome sequence,
wherein the multi-parameter likelihood function value is based on the first configuration of reads of the plurality of families, the second configuration of reads of the plurality of families, and the third configuration of reads of the plurality of families.
71 . The method of claim 70 , further comprising:
generating, based on at least one or more of the plurality of model parameters, an overall configuration of the plurality of sequence reads, the overall configuration omitting a value for the frequency of the variant allele in the plurality of sequence reads, wherein the multi-parameter likelihood function value is generated based on the overall configuration.
72 . The method of claim 70 , wherein further comprising:
for each family of the plurality of families:
generating forward strand data based on a number of reads of a forward strand of the family having the candidate indel; and
generating reverse strand data based on a number of reads of a reverse strand of the family having the candidate indel,
wherein the multi-parameter likelihood function value is based further on the forward strand data of the plurality of families and the reverse strand data of the plurality of families.
73 . The method of claim 70 , wherein generating the family probability comprises:
for each family of the plurality of families:
generating first read data based on a number sequence reads of the family that supports the variant allele;
generating second read data based on a number of sequence reads of the family that supports another a non-reference allele other than the variant allele; and
generating third read data based on a number of the sequence reads of the family that supports a reference allele that appears in the reference genome,
wherein the multi-parameter likelihood function value is based on the first read data of the plurality of families, the second read data of the plurality of families, and the third read data of the plurality of families.
74 . A computer system for classifying a candidate insertion or deletion (indel) detected in a plurality of sequence reads as a true indel or an indel not in a subject, the plurality of sequence reads generated from deoxyribonucleic acid (DNA) molecules in a bodily sample of the subject, the system comprising:
a processor programmed to:
access the plurality of sequence reads;
map the plurality of sequence reads to a reference genome associated with the DNA molecules;
detect a candidate indel in the plurality of sequence reads;
generate a first probability that the candidate indel is a true indel;
generate a second probability that the candidate indel is an indel not in the subject; and
determine whether the candidate indel is a true indel or an indel not in the subject based on the first probability and the second probability.
75 . The computer system of claim 74 , wherein the processor is further programmed to:
locate the plurality of sequence reads with respect to a reference genome associated with the DNA molecules; determine a plurality of families of sequence reads based on the location; and determine a plurality of model parameters that each affects a classification of whether the candidate indel is a true indel, wherein the first probability and the second probability are each based on the plurality of model parameters and the plurality of families.
76 . The computer system of claim 75 , wherein the processor is further programmed to:
for each family of the plurality of families:
generate a first configuration of reads based on a first number of reads in the family having a variant allele that includes the candidate indel;
generate a second configuration of reads based on a second number of reads in the family having a non-reference allele other than the variant allele; and
generate a third configuration of reads based on a third number of reads in the third family having a reference allele that appears in the genome sequence,
wherein the first probability and the second probability are each based on: the first configuration of reads of the plurality of families, the second configuration of reads of the plurality of families, and the third configuration of reads of the plurality of families.
77 . The computer system of claim 76 , wherein the processor is further programmed to:
for each family of the plurality of families:
generate forward strand data based on a number of reads of a forward strand of the family having the candidate indel; and
generate reverse strand data based on a number of reads of a reverse strand of the family having the candidate indel,
wherein the first value and the second value are each based on the forward strand data of the plurality of families and the reverse strand data of the plurality of families.
78 . The computer system of claim 77 , wherein the processor is further programmed to:
for each family of the plurality of families:
generate first read data based on a number of sequence reads of the family that supports the variant allele;
generate second read data based on a number of the sequence reads of the family that supports the non-reference allele other than the variant allele; and
generate third read data based on a number of the sequence reads of the family that supports the reference allele that appears in the reference genome,
wherein the first probability and the second probability are each based on the first read data of the plurality of families, the second read data of the plurality of families, and the third read data of the plurality of families.
79 . A method of classifying a candidate insertion or deletion (indel) detected in a plurality of sequence reads as a true indel or an indel not in a subject, the plurality of sequence reads generated from deoxyribonucleic acid (DNA) molecules in a bodily sample of the subject, comprising:
accessing, by a processor, the plurality of sequence reads; mapping, by the processor, the plurality of sequence reads to a reference genome associated with the DNA molecules; detecting, by the processor, a candidate indel in the plurality of sequence reads; generating, by the processor, a first probability that the candidate indel is a true indel; generating, by the processor, a second probability that the candidate indel is an indel not in the subject; and determining, by the processor, whether the candidate indel is a true indel or an indel not in the subject based on the first probability and the second probability.
80 . The method of claim 79 , further comprising:
locating the plurality of sequence reads with respect to a reference genome associated with the DNA molecules; determining a plurality of families of sequence reads based on the location; and determining a plurality of model parameters that each affects a classification of whether the candidate indel is a true indel, wherein the first probability and the second probability are each based on the plurality of model parameters and the plurality of families.
81 . The method of claim 80 , further comprising:
for each family of the plurality of families:
generating a first configuration of reads based on a first number of reads in the family having a variant allele that includes the candidate indel;
generating a second configuration of reads based on a second number of reads in the family having a non-reference allele other than the variant allele; and
generating a third configuration of reads based on a third number of reads in the third family having a reference allele that appears in the genome sequence,
wherein the first probability and the second probability are each based on: the first configuration of reads of the plurality of families, the second configuration of reads of the plurality of families, and the third configuration of reads of the plurality of families.
82 . The method of claim 81 , further comprising:
for each family of the plurality of families:
generating forward strand data based on a number of reads of a forward strand of the family having the candidate indel; and
generating reverse strand data based on a number of reads of a reverse strand of the family having the candidate indel,
wherein the first value and the second value are each based on the forward strand data of the plurality of families and the reverse strand data of the plurality of families.
83 . The method of claim 82 , further comprising:
for each family of the plurality of families:
generating first read data based on a number of sequence reads of the family that supports the variant allele;
generating second read data based on a number of the sequence reads of the family that supports the non-reference allele other than the variant allele; and
generating third read data based on a number of the sequence reads of the family that supports the reference allele that appears in the reference genome,
wherein the first probability and the second probability are each based on the first read data of the plurality of families, the second read data of the plurality of families, and the third read data of the plurality of families.Join the waitlist — get patent alerts
Track US2020013482A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.