US2023193360A1PendingUtilityA1

Determination of base modifications of nucleic acids

Assignee: UNIV HONG KONG CHINESEPriority: Aug 16, 2019Filed: Aug 19, 2022Published: Jun 22, 2023
Est. expiryAug 16, 2039(~13 yrs left)· nominal 20-yr term from priority
C12Q 1/6816C12Q 1/6851G16B 30/00G16B 20/00C12N 9/22G16B 40/10C12Q 2537/164C12Q 1/6869G16B 20/10G16B 40/20G16B 15/00C12Q 2565/601G16B 20/20C12Q 2600/154C12N 15/11C12Q 1/68G16B 40/00
75
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for using determination of base modification in analyzing nucleic acid molecules and acquiring data for analysis of nucleic acid molecules are described herein. Base modifications may include methylations. Methods to determine base modifications may include using features derived from sequencing. These features may include the pulse width of an optical signal from sequencing bases, the interpulse duration of bases, and the identity of the bases. Machine learning models can be trained to detect the base modifications using these features. The relative modification or methylation levels between haplotypes may indicate a disorder. Modification or methylation statuses may also be used to detect chimeric molecules.

Claims

exact text as granted — not AI-modified
1 . A method for detecting a methylation of a cytosine in a nucleic acid molecule, the method comprising:
 (a) receiving data acquired by sequencing a sample nucleic acid molecule by measuring pulses in an optical signal corresponding to nucleotides and obtaining, from the data, values for the following properties: 
 for each nucleotide: 
 an identity of the nucleotide, 
 a position of the nucleotide within the sample nucleic acid molecule, 
 a width of the pulse corresponding to the nucleotide, and 
 an interpulse duration representing a time between the pulse corresponding to the nucleotide and a pulse corresponding to a neighboring nucleotide; 
 
   (b) creating an input data structure, the input data structure comprising a window of the nucleotides sequenced in the sample nucleic acid molecule, wherein the input data structure includes, for each nucleotide within the window, the properties: 
 the identity of the nucleotide, 
 a position of the nucleotide with respect to a target position within the window, 
 the width of the pulse corresponding to the nucleotide, and 
 the interpulse duration; 
   (c) inputting the input data structure into a model, the model trained by: 
 receiving a first plurality of first data structures, each first data structure of the first plurality of first data structures corresponding to a respective window of nucleotides sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring pulses in the optical signal corresponding to the nucleotides, wherein the methylation has a known first state in a cytosine at a target position in each window of each first nucleic acid molecule, each first data structure comprising values for the same properties as the input data structure, 
 storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the cytosine at the target position, and 
 optimizing, using the plurality of first training samples, parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels when the first plurality of first data structures is input to the model, wherein an output of the model specifies whether the cytosine at the target position in the respective window has the methylation; and 
   (d) determining, using the model, whether the methylation is present inthe cytosine at the target position within the window in the input data structure.   
     
     
         2 . The method of  claim 1 , wherein:
 the input data structure is one input data structure of a plurality of input data structures,   the sample nucleic acid molecule is one sample nucleic acid molecule of a plurality of sample nucleic acid molecules,   the plurality of sample nucleic acid molecules is obtained from a biological sample of a subject, and   each input data structure corresponds to a respective window of nucleotides sequenced in a respective sample nucleic acid molecule of the plurality of sample nucleic acid molecules, and   the method further comprising:
 receiving the plurality of input data structures, 
 inputting the plurality of input data structures into the model, and 
 determining, using the model, whether the methylation is present in a 
 cytosine at a target location in the respective window of each input data structure.   
     
     
         3 - 5 . (canceled) 
     
     
         6 . The method of  claim 1 , wherein the model includes a machine learning model, a principal component analysis, a convolutional neural network, or a logistic regression. 
     
     
         7 . The method of  claim 1 , wherein:
 the window of nucleotides corresponding to the input data structure comprises nucleotides on a first strand of the sample nucleic acid molecule and nucleotides on a second strand of the sample nucleic acid molecule, and   the input data structure further comprises for each nucleotide within the window a value of a strand property, the strand property indicating the nucleotide being present on either the first strand or the second strand.   
     
     
         8 . The method of  claim 1 , wherein the nucleotides within the window are determined using a circular consensus sequence and without alignment of the sequenced nucleotides to a reference genome. 
     
     
         9 - 10 . (canceled) 
     
     
         11 . The method of  claim 1 , wherein the optical signal is a fluorescence signal from a dye-labeled nucleotide. 
     
     
         12 . The method of  claim 1 , wherein each window associated with the first plurality of first data structures comprises 13 consecutive nucleotides on a first strand of each first nucleic acid molecule. 
     
     
         13 . (canceled) 
     
     
         14 . The method of  claim 1 , further comprising:
 validating the model using a plurality of nucleic acid molecules, each including a first portion corresponding to a first reference sequence and a second portion corresponding to a second reference sequence, wherein the first portion has a first methylation pattern, and the second portion has a second methylation pattern.   
     
     
         15 . The method of  claim 14 , wherein the first portion is treated with a methylase. 
     
     
         16 . The method of  claim 15 , wherein the second portion corresponds to an unmethylated portion of the second reference sequence. 
     
     
         17 - 20 . (canceled) 
     
     
         21 . A method for detecting a methylation of a cytosine in a nucleic acid molecule, the method comprising:
 (a) receiving data acquired by sequencing a sample nucleic acid molecule by measuring pulses in an optical signal corresponding to nucleotides and obtaining, from the data, values for the following properties:
 for each nucleotide:
 an identity of the nucleotide, 
 a position of the nucleotide within the sample nucleic acid molecule, 
 a width of the pulse corresponding to the nucleotide, and 
 an interpulse duration representing a time between the pulse corresponding to the nucleotide and a pulse corresponding to a neighboring nucleotide; 
 
   (b) creating an input data structure, the input data structure comprising a window of the nucleotides sequenced in the sample nucleic acid molecule, wherein the window comprises 6 consecutive nucleotides upstream of a nucleotide at a position within the window and 6 consecutive nucleotides downstream of the nucleotide at the target position, wherein the input data structure includes, for each nucleotide within the window, the properties:
 the identity of the nucleotide, 
 a position of the nucleotide with respect to the target position , 
 the width of the pulse corresponding to the nucleotide, and 
 the interpulse duration; 
   (c) inputting the input data structure into a model, the model trained by:
 receiving a first plurality of first data structures, each first data structure of the first plurality of first data structures corresponding to a respective window of nucleotides sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring pulses in the optical signal corresponding to the nucleotides, wherein the methylation has a known first state in a cytosine at a target position in each window of each first nucleic acid molecule, the methylation being 5mC (5-methylcytosine), each first data structure comprising values for the same properties as the input data structure, 
 storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the cytosine at the target position, and 
 optimizing, using the plurality of first training samples, parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels when the first plurality of first data structures is input to the model, wherein an output of the model specifies whether the cytosine at the target position in the respective window has the methylation; and 
   (d) determining, using the model, whether the 5mC methylation is present in the cytosine at the target position within the window in the input data structure.   
     
     
         22 . The method of  claim 21 , wherein the model includes a machine learning model, a principal component analysis, a convolutional neural network, or a logistic regression. 
     
     
         23 . A computer product comprising a non-transitory computer readable medium storing a plurality of instructions that when executed control a computer system to perform a method for detecting a methylation of a cytosine in a nucleic acid molecule, the method comprising:
 (a) receiving data acquired by sequencing a sample nucleic acid molecule by measuring pulses in an optical signal corresponding to nucleotides and obtaining, from the data, values for the following properties:
 for each nucleotide:
 an identity of the nucleotide, 
 a position of the nucleotide within the sample nucleic acid molecule, 
 a width of the pulse corresponding to the nucleotide, and 
 an interpulse duration representing a time between the pulse corresponding to the nucleotide and a pulse corresponding to a neighboring nucleotide; 
 
   (b) creating an input data structure, the input data structure comprising a window of the nucleotides sequenced in the sample nucleic acid molecule, wherein the input data structure includes, for each nucleotide within the window, the properties:
 the identity of the nucleotide, 
 a position of the nucleotide with respect to a target position within the window, 
 the width of the pulse corresponding to the nucleotide, and 
 the interpulse duration; 
   (c) inputting the input data structure into a model, the model trained by:
 receiving a first plurality of first data structures, each first data structure of the first plurality of first data structures corresponding to a respective window of nucleotides sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring pulses in the optical signal corresponding to the nucleotides, wherein the methylation has a known first state in a cytosine at a target position in each window of each first nucleic acid molecule, each first data structure comprising values for the same properties as the input data structure, 
 storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the cytosine at the target position, and 
 optimizing, using the plurality of first training samples, parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels when the first plurality of first data structures is input to the model, wherein an output of the model specifies whether the cytosine at the target position in the respective window has the methylation; and 
   (d) determining, using the model, whether the methylation is present in the cytosine at the target position within the window in the input data structure.   
     
     
         24 . The computer product of  claim 23 , wherein the methylation is 5mC (5-methylcytosine). 
     
     
         25 . The computer product of  claim 23 , wherein:
 the input data structure is one input data structure of a plurality of input data structures,   the sample nucleic acid molecule is one sample nucleic acid molecule of a plurality of sample nucleic acid molecules,   the plurality of sample nucleic acid molecules is obtained from a biological sample of a subject, and   each input data structure corresponds to a respective window of nucleotides sequenced in a respective sample nucleic acid molecule of the plurality of sample nucleic acid molecules, and   the method further comprising: 
 receiving the plurality of input data structures, 
 inputting the plurality of input data structures into the model, and 
   determining, using the model, whether the methylation is present in a cytosine at a target location in the respective window of each input data structure.   
     
     
         26 . (canceled) 
     
     
         27 . The computer product of  claim 25 , wherein:
 the plurality of sample nucleic acid molecules aligns to a plurality of genomic regions,   for each genomic region of the plurality of genomic regions:
 a number of sample nucleic acid molecules is aligned to the genomic region, 
   the number of sample nucleic acid molecules is greater than a cutoff number.   
     
     
         28 . The computer product of  claim 23 , wherein the model includes a machine learning model, a principal component analysis, a convolutional neural network, or a logistic regression. 
     
     
         29 - 30 . (canceled) 
     
     
         31 . The method of  claim 1 , wherein the window of the input data structure has a different number of consecutive nucleotides upstream of the nucleotide at the target position than the number of consecutive nucleotides downstream of the nucleotide at the target position. 
     
     
         32 . The method of  claim 1 , wherein the window of the input data structure comprises 10 consecutive nucleotides upstream of the nucleotide at the target position and 10 consecutive nucleotides downstream of the nucleotide at the target position. 
     
     
         33 . The method of  claim 1 , wherein the window of the input data structure comprises 21 consecutive nucleotides upstream of the nucleotide at the target position and 21 consecutive nucleotides downstream of the nucleotide at the target position. 
     
     
         34 . The method of  claim 21 , wherein the optical signal is a fluorescence signal from a dye-labeled nucleotide. 
     
     
         35 . The computer product of  claim 23 , wherein the optical signal is a fluorescence signal from a dye-labeled nucleotide. 
     
     
         36 . The method of  claim 21 , wherein the nucleotides within the window are determined using a circular consensus sequence and without alignment of the sequenced nucleotides to a reference genome. 
     
     
         37 . The computer product of  claim 23 , wherein the nucleotides within the window are determined using a circular consensus sequence and without alignment of the sequenced nucleotides to a reference genome. 
     
     
         38 . The method of  claim 21 , wherein each window associated with the first plurality of first data structures comprises 13 consecutive nucleotides on a first strand of each first nucleic acid molecule. 
     
     
         39 . The computer product of  claim 23 , wherein each window associated with the first plurality of first data structures comprises 13 consecutive nucleotides on a first strand of each first nucleic acid molecule. 
     
     
         40 . The method of  claim 21 , wherein the window of the input data structure has a different number of consecutive nucleotides upstream of the nucleotide at the target position than the number of consecutive nucleotides downstream of the nucleotide at the target position. 
     
     
         41 . The computer product of  claim 23 , wherein the window of the input data structure has a different number of consecutive nucleotides upstream of the nucleotide at the target position than the number of consecutive nucleotides downstream of the nucleotide at the target position. 
     
     
         42 . The method of  claim 21 , wherein the window of the input data structure comprises 10 consecutive nucleotides upstream of the nucleotide at the target position and 10 consecutive nucleotides downstream of the nucleotide at the target position. 
     
     
         43 . The computer product of  claim 23 , wherein the window of the input data structure comprises 10 consecutive nucleotides upstream of the nucleotide at the target position and 10 consecutive nucleotides downstream of the nucleotide at the target position.

Join the waitlist — get patent alerts

Track US2023193360A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.