Method and system for DNA analysis
Abstract
The present invention pertains to a process for automatically analyzing nucleic acid samples. Specifically, the process comprises the steps of forming electrophoretic data of DNA samples with DNA ladders; comparing these data; transforming the coordinates of the DNA sample's data into DNA length coordinates; and analyzing the DNA sample in length coordinates. This analysis is useful for automating fragment analysis and quality assessment. The automation enables a business model based on usage, since it replaces (rather than assists) labor. This analysis also provides a mechanism whereby data generated on different instruments can be confidently compared. Genetic applications of this invention include gene discovery, genetic diagnosis, and drug discovery. Forensic applications include identifying people and their relatives, catching perpetrators, analyzing DNA mixtures, and exonerating innocent suspects.
Claims
exact text as granted — not AI-modified1 . A method for analyzing a nucleic acid sample comprised of the steps:
(a) forming labeled DNA sample fragments from a nucleic acid sample; (b) size separating said sample fragments with size standard fragments, and detecting the fragments to form a sample signal and a size standard signal; (c) transforming the sample signal into size coordinates using the size standard signal; (d) analyzing the nucleic acid sample in size coordinates to form data; and (e) using a computer to apply at least seven different rules to check the data for possible artifacts, where at least one of the rules compares observed measures of the data against expected data behavior.
2 . A method for analyzing a nucleic acid sample comprised of the steps:
(a) forming labeled DNA sample fragments from a nucleic acid sample; (b) size separating and detecting said sample fragments to form a sample signal; (c) forming labeled DNA ladder fragments corresponding to molecular lengths; (d) size separating and detecting said ladder fragments to form a ladder signal; (e) transforming the sample signal into an allelic ladder size coordinate using the ladder signal, where the lengths of the DNA sample fragments are expressed using integer values of base pairs; and (f) analyzing the nucleic acid sample signal in length coordinates.
3 . A method as described in claim 2 wherein after the analyzing step (f) there is the additional step of determining a length or amount of a fragment in the nucleic acid sample.
4 . A method as described in claim 3 wherein after the determining step there is the additional step of identifying an individual by DNA profiling.
5 . A system for analyzing a nucleic acid sample comprising:
(a) means for forming labeled DNA sample fragments from a nucleic acid sample; (b) means for size separating and detecting said sample fragments to form a sample signal, said separating and detecting means in communication with the sample fragments; (c) means for forming labeled DNA ladder fragments corresponding to molecular lengths; (d) means for size separating and detecting said ladder fragments to form a ladder signal, said separating and detecting means in communication with the ladder fragments; (e) means for transforming the sample signal into an allelic ladder size coordinate system length coordinates using the ladder signal, where the lengths of the DNA sample fragments are expressed using integer values of base pairs, said transforming means in communication with the signals; and (f) means for analyzing the nucleic acid sample signal in length coordinates, said analyzing means in communication with the transforming means.
6 . A method as described in claim 5 wherein the nucleic acid analysis identifies an individual by DNA profiling.
7 . A method as described in claim 1 wherein the rule computes a quality score on the data.
8 . A method as described in claim 1 wherein the rule models the steps of processing, computes appropriate quality scores at each step, and compares these observed data features with normative results.
9 . A method as described in claim 1 wherein the rule provides a computer diagnosis of problems with data signals and their quality.
10 . A method as described in claim 1 wherein the rule decides whether or not an experiment is comprised primarily of noise.
11 . A method as described in claim 1 wherein the rule determines whether or not data peaks fail to reach a predefined minimum height threshold.
12 . A method as described in claim 1 wherein the rule determines whether or not data peaks exceed a predefined maximum height threshold.
13 . A method as described in claim 1 wherein the rule determines whether or not designated peaks fail to comprise a predefined percentage of total signal.
14 . A method as described in claim 1 wherein the rule determines whether or not there is a conflict between genotype results derived from multiple scoring algorithms.
15 . A method as described in claim 1 wherein the rule determines whether or not a relative ratio of peak heights falls within a predefined range.
16 . A method as described in claim 1 wherein the rule determines whether or not a third peak contains too much signal relative to allele peaks.
17 . A method as described in claim 1 wherein the rule determines whether or not an allele peak is too far away from its allelic ladder peak.
18 . A method as described in claim 1 wherein the rule determines whether or not two allele peak sizes have a similar deviation from their respective ladder peak sizes.
19 . A method as described in claim 1 wherein the rule determines whether or not a control experiment is consistent with a known result.
20 . A method as described in claim 1 wherein the rule is statistically determined by distinguishing two score distributions according to a predetermined sensitivity or specificity.
21 . A method as described in claim 1 wherein the rule forms a probability for correct allele calls.
22 . A method as described in claim 1 wherein the rule is used to prioritize data review so that a person reviews worse data before better data.
23 . A method as described in claim 1 wherein the rule is used to focus data review so that highly confident data are not reviewed by a human operator.
24 . A method as described in claim 1 wherein the rule is used to focus data review so that unscorable data are not reviewed by a human operator.
25 . A method as described in claim 1 wherein the rule automatically determines what data artifacts are present, and then displays an appropriate visualization of the specific artifact based on this determination.
26 . A method as described in claim 1 wherein there is a plurality of at least seven different rules that are applied to the same data.Join the waitlist — get patent alerts
Track US2010010748A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.