US2020279620A1PendingUtilityA1
Methods and systems for interpretation and reporting of sequence-based genetic tests using pooled allele statistics
Est. expiryAug 15, 2034(~8 yrs left)· nominal 20-yr term from priority
G16B 50/20G16B 50/30G16B 50/10G16H 10/20G16B 50/00
60
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Disclosed herein are system, method, and computer program product embodiments for building a community database of allele counts. An embodiment operates by receiving human variant datasets derived from samples generated by distinct users, wherein the users consented to share pooled variant observations with other users; determining that a plurality of variant observations meet the inclusion criteria for a pool; and calculating one or more anonymized allele statistics from the pool.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for building a community database of variant observations, comprising:
receiving human variant datasets derived from a sample; storing the received human variant datasets in a knowledge base of genomic information, wherein the knowledge base comprises an ontology data structure that identifies a causal relationship between a genetic variant and a phenotype based on a combination of the genetic variant and modifier variant information, and wherein the modifier variant information identifies whether a modifier variant that modifies a severity of the phenotype is likely to exist; calculating one or more anonymized allele statistics from a plurality of variant observations identified from the knowledge base including the human variant datasets; and classifying each variant observation from the plurality of variant observations into a clinical significance category using the knowledge base and the anonymized allele statistics calculated based on the plurality of human variant datasets, wherein at least one of the receiving, storing, calculating, or classifying are performed by one or more computers.
2 . The method of claim 1 , further comprising:
determining that variant observations corresponding to the human variant datasets derived from the sample do not contribute to a pool of alleles based on at least one of a breadth of genome coverage of the sample, depth of coverage of the sample, quality of the sample, variant call quality, a phenotype associated with the sample, sample redundancy, variant counts, a trust metric for a source of the sample, community feedback, containing a well-established disease-causing variant, manual or automated quality control, or any combination thereof.
3 . The method of claim 1 , further comprising:
annotating the sample with one or more ethnicities using at least one of principal component analysis, a user-provided annotation, or any combination thereof.
4 . The method of claim 1 , wherein the calculating comprises:
calculating an allele frequency that is a ratio of a number of observed incidences of a given variant to a total number of variant observations corresponding to the human variant datasets derived from the sample, in the plurality of variant observations believed to have potential to measure the given variant.
5 . The method of claim 1 , further comprising:
providing to a user one or more anonymized allele statistics for one or more alleles; and excluding access to the anonymized allele statistics from users who have not provided consent during the receiving step.
6 . The method of claim 5 , wherein the providing comprises:
providing the anonymized allele statistics to the user via a web-based resource.
7 . The method of claim 1 , further comprising:
filtering variants from the received human variant datasets, using the anonymized allele statistics.
8 . The method of claim 1 , further comprising:
providing an incentive to one or more of the users for sharing the received human variant datasets, wherein the incentive is access to the anonymized allele statistics.
9 . The method of claim 1 , further comprising:
determining that variant observations corresponding to the human variant datasets derived from the sample do not contribute to a pool of alleles based on at least one of variant call quality, read depth, known association with a common technical error or failure mode, manual or automated quality control, or any combination thereof.
10 . A system, comprising:
a memory; and at least one processor coupled to the memory and configured to: receive human variant datasets derived from a sample; store the received human variant datasets in a knowledge base of genomic information, wherein the knowledge base comprises an ontology data structure that identifies a causal relationship between a genetic variant and a phenotype based on a combination of the genetic variant and modifier variant information, and wherein the modifier variant information identifies whether a modifier variant that modifies a severity of the phenotype is likely to exist; calculate one or more anonymized allele statistics from a plurality of variant observations identified from the knowledge base including the human variant datasets; and classify each variant observation from the plurality of variant observations into a clinical significance category using the knowledge base and the anonymized allele statistics calculated based on the plurality of human variant datasets.
11 . The system of claim 10 , wherein the processor is configured to:
determine that variant observations corresponding to the human variant datasets derived from the sample do not contribute to a pool of alleles based on at least one of a breadth of genome coverage of the sample, depth of coverage of the sample, quality of the sample, variant call quality, a phenotype associated with the sample, sample redundancy, variant counts, a trust metric for a source of the sample, community feedback, containing a well-established disease-causing variant, manual or automated quality control, or any combination thereof.
12 . The system of claim 10 , wherein the processor is configured to:
annotate the sample with one or more ethnicities using at least one of principal component analysis, a user-provided annotation, or any combination thereof.
13 . The system of claim 10 , wherein the processor is configured to:
calculate an allele frequency that is a ratio of a number of observed incidences of a given variant to a total number of variant observations corresponding to the human variant datasets derived from the sample, in the plurality of variant observations believed to have potential to measure the given variant.
14 . The system of claim 10 , wherein the processor is configured to:
provide to a user one or more anonymized allele statistics for one or more alleles; and exclude access to the anonymized allele statistics from users who have not provided consent during the receiving step.
15 . The system of claim 14 , wherein the processor is configured to:
provide the anonymized allele statistics to the user via a web-based resource.
16 . The system of claim 10 , wherein the processor is configured to:
filter variants from the received human variant datasets, using the anonymized allele statistics.
17 . The system of claim 10 , wherein the processor is configured to:
provide an incentive to one or more of the users for sharing the received human variant datasets.
18 . The system of claim 17 , wherein the incentive is access to the allele statistics.
19 . The system of claim 10 , wherein the processor is configured to:
determine that variant observations corresponding to the human variant datasets derived from the sample do not contribute to a pool of alleles based on at least one of variant call quality, read depth, known association with a common technical error or failure mode, manual or automated quality control, or any combination thereof.
20 . A non-transitory computer readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:
receiving human variant datasets derived from a sample; storing the received human variant datasets in a knowledge base of genomic information, wherein the knowledge base comprises an ontology data structure that identifies a causal relationship between a genetic variant and a phenotype based on a combination of the genetic variant and modifier variant information, and wherein the modifier variant information identifies whether a modifier variant that modifies a severity of the phenotype is likely to exist; calculating one or more anonymized allele statistics from a plurality of variant observations identified from the knowledge base including the human variant datasets; and classifying each variant observation from the plurality of variant observations into a clinical significance category using the knowledge base and the anonymized allele statistics calculated based on the plurality of human variant datasets.Join the waitlist — get patent alerts
Track US2020279620A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.