US2022205037A1PendingUtilityA1

Methods and compositions for analyzing nucleic acid

Assignee: SEQUENOM INCPriority: May 21, 2012Filed: Mar 10, 2022Published: Jun 30, 2022
Est. expiryMay 21, 2032(~5.8 yrs left)· nominal 20-yr term from priority
G16B 40/10G16B 40/00G16B 30/00G16B 30/20G16B 30/10C12Q 1/6869
78
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Technology provided herein relates in part to methods, processes, compositions and apparatuses for analyzing nucleic acid.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system comprising one or more processors and memory, which memory comprises instructions executable by the one or more processors and which memory comprises counts of thousands to millions of nucleotide sequence reads mapped to genomic portions of a reference genome, which sequence reads are reads of circulating cell-free nucleic acid from a test sample from a pregnant female bearing a fetus and which sequence reads are generated by a non-loci-specific massively parallel sequencing process, and which instructions executable by the one or more processors are configured to:
 (a) measure the lengths of nucleic acid fragments from which the sequence reads are generated;   (b) select counts for the thousands to millions of sequence reads from nucleic acid fragments that are shorter than about 150 bases to about 160 bases, thereby generating selected counts of the thousands to millions of sequence reads, wherein the selected counts are enriched for counts of reads from fetal nucleic acid; and   (c) normalize the selected counts of the thousands to millions of sequence reads for the test sample according to an experimental bias for the test sample and an experimental bias for each of multiple samples from multiple pregnant females, thereby generating normalized selected counts, wherein experimental bias in the normalized selected counts is reduced.   
     
     
         2 . The system of  claim 1 , wherein the instructions executable by the one or more processors are configured to select counts for the thousands to millions of sequence reads from nucleic acid fragments that are shorter than about 150 bases. 
     
     
         3 . The system of  claim 1 , wherein the instructions executable by the one or more processors are configured to select counts for the thousands to millions of sequence reads from nucleic acid fragments that are shorter than about 160 bases. 
     
     
         4 . The system of  claim 1 , wherein the lengths of the nucleic acid are measured according to positions of mapped sequence reads obtained from a paired-end sequencing process. 
     
     
         5 . The system of  claim 1 , wherein sequence read length is about 15 bases to about 25 bases. 
     
     
         6 . The system of  claim 1 , wherein the test sample is blood, serum, or plasma. 
     
     
         7 . The system of  claim 1 , wherein the experimental bias for the test sample in (c) is a guanine and cytosine (GC) bias, and the experimental bias for each of multiple samples from multiple pregnant females in (c) is a guanine and cytosine (GC) bias. 
     
     
         8 . The system of  claim 7 , wherein the normalizing in (c) comprises 1) determining a guanine and cytosine (GC) bias coefficient for the test sample based on a fitted relation between (i) the counts of the sequence reads mapped to each of the genomic portions and (ii) GC content for each of the genomic portions, wherein the GC bias coefficient is a slope for a linear fitted relation or a curvature estimation for a non-linear fitted relation; and 2) determining a fitted relation, for each of the genomic portions, between (i) a GC bias coefficient for each of the multiple samples from multiple pregnant females and (ii) counts of sequence reads mapped to each of the genomic portions for the multiple samples. 
     
     
         9 . The system of  claim 1 , wherein the instructions executable by the one or more processors are further configured to map the sequence reads to the portions of the reference genome, and count the mapped sequence reads, thereby generating the counts of the sequence reads mapped to the portions of the reference genome. 
     
     
         10 . The system of  claim 1 , wherein the sequence reads are mapped to multiple chromosome portions from multiple chromosomes of the reference genome. 
     
     
         11 . The system of  claim 1 , wherein the sequence reads are mapped to portions of a complete human reference genome. 
     
     
         12 . The system of  claim 1 , wherein the portions of the reference genome comprise portions that are about 10 kilobases (kb) to about 100 kb. 
     
     
         13 . The system of  claim 1 , wherein the memory comprises counts of millions of nucleotide sequence reads mapped to the genomic portions of the reference genome. 
     
     
         14 . The system of  claim 1 , wherein the sequencing process is performed at about 0.1-fold coverage, about 0.2-fold coverage, about 0.3-fold coverage, about 0.4-fold coverage, about 0.5-fold coverage, about 0.6-fold coverage, about 0.7-fold coverage, about 0.8-fold coverage, about 0.9-fold coverage, about 1-fold coverage, or greater than 1-fold coverage. 
     
     
         15 . The system of  claim 1 , wherein sequence coverage of reads from nucleic acid fragments shorter than about 150 bases to about 160 bases is reduced relative to sequence coverage of reads not restricted by fragment length. 
     
     
         16 . The system of  claim 1 , wherein sequence read count of reads from nucleic acid fragments shorter than about 150 bases to about 160 bases is reduced relative to sequence read count of reads not restricted by fragment length.

Join the waitlist — get patent alerts

Track US2022205037A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.