US2023402132A1PendingUtilityA1

Error Correction in Ancestry Classification

Assignee: 23ANDME INCPriority: Nov 8, 2012Filed: May 5, 2023Published: Dec 14, 2023
Est. expiryNov 8, 2032(~6.3 yrs left)· nominal 20-yr term from priority
G16B 40/00G06N 5/04G06N 20/00G06N 7/01G16B 40/20G06N 20/10
87
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Error correction in ancestry classification includes obtaining, from a classifier, initial ancestry classifications associated with portions of two phased haplotypes of a chromosome pair of an individual; performing error correction on an initial ancestry classification, including detecting a phasing error in the initial ancestry classifications; and outputting a corrected ancestry classification in which the phasing error is corrected.

Claims

exact text as granted — not AI-modified
1 . A method, implemented using a computer comprising one or more processors and memory, the method comprising:
 obtaining, from a classifier and by the one or more processors, a plurality of initial geographical ancestry classifications respectively associated with segments of two phased haplotypes of a chromosome pair of an individual, each phased haplotype being inherited from one of two parents of the individual, each segment of one of the two phased haplotypes corresponding to a segment of another of the two phased haplotypes;   obtaining, by the one or more processors, a Pair Hidden Markov Model (PHMM) in which observed states correspond to ordered pairs of the initial geographical ancestry classifications associated with two corresponding segments of the two phased haplotypes, and in which hidden states correspond to ordered pairs of underlying geographical ancestry classifications associated with the two corresponding segments of the two phased haplotypes;   determining, by the one or more processors and using the PHMM, two geographical ancestry classifications respectively associated with the two phased haplotypes;   correcting, by the one or more processors, the initial geographical ancestry classifications using the two geographical ancestry classifications; and   outputting, by the one or more processors, the initial geographical ancestry classifications as corrected.   
     
     
         2 . The method of  claim 1 , wherein determining the two geographical ancestry classifications comprises:
 determining a most likely sequence of the hidden states given the initial geographical ancestry classifications.   
     
     
         3 . The method of  claim 2 , wherein the two geographical ancestry classifications include two sequences of geographical ancestry classifications, and wherein the most likely sequence of the hidden states comprises the two sequences of geographical ancestry classifications. 
     
     
         4 . The method of  claim 2 , wherein determining the most likely sequence of hidden states comprises detecting a correlated prediction error that was caused by the classifier. 
     
     
         5 . The method of  claim 4 , wherein the initial geographical ancestry classification as corrected rectifies the correlated prediction error. 
     
     
         6 . The method of  claim 1 , wherein the PHMM is an Autoregressive Pair Hidden Markov Model (APHMM) in which the observed states are dependent on their corresponding hidden states and previous observed states. 
     
     
         7 . The method of  claim 1 , wherein correcting the initial geographical ancestry classifications comprises:
 obtaining a plurality of Pair Hidden Markov Models (PHMMs) including the PHMM, wherein each of the plurality of PHMMs corresponds to a distinct reference population; and   determining the initial geographical ancestry classifications as corrected based on a weighting of the plurality of PHMMs.   
     
     
         8 . The method of  claim 7 , wherein determining the initial geographical ancestry classifications as corrected comprises:
 applying the initial geographical ancestry classifications associated with segments of the two phased haplotypes to the plurality of PHMMs, and performing a Bayesian model averaging on outputs of the plurality of PHMMs.   
     
     
         9 . The method of  claim 1 , wherein obtaining the PHMM comprises:
 performing, by the one or more processors on the initial geographical ancestry classifications, dynamic programming to determine the PHMM.   
     
     
         10 . A system comprising one or more processors, and one or more memories configured to store instructions that, when executed by the one or more processors, cause the system to perform operations comprising:
 obtaining, from a classifier, a plurality of initial geographical ancestry classifications respectively associated with segments of two phased haplotypes of a chromosome pair of an individual, each phased haplotype being inherited from one of two parents of the individual, each segment of one of the two phased haplotypes corresponding to a segment of another of the two phased haplotypes;   obtaining a Pair Hidden Markov Model (PHMM) in which observed states correspond to ordered pairs of the initial geographical ancestry classifications associated with two corresponding segments of the two phased haplotypes, and in which hidden states correspond to ordered pairs of underlying geographical ancestry classifications associated with the two corresponding segments of the two phased haplotypes;   determining, using the PHMM, two geographical ancestry classifications respectively associated with the two phased haplotypes;   correcting the initial geographical ancestry classifications using the two geographical ancestry classifications; and   outputting the initial geographical ancestry classifications as corrected.   
     
     
         11 . The system of  claim 10 , wherein determining the two geographical ancestry classifications comprises:
 determining a most likely sequence of the hidden states given the initial geographical ancestry classifications.   
     
     
         12 . The system of  claim 11 , wherein the two geographical ancestry classifications include two sequences of geographical ancestry classifications, and wherein the most likely sequence of the hidden states comprises the two sequences of geographical ancestry classifications. 
     
     
         13 . The system of  claim 11 , wherein determining the most likely sequence of hidden states comprises detecting a correlated prediction error that was caused by the classifier. 
     
     
         14 . The system of  claim 10 , wherein correcting the initial geographical ancestry classifications comprises:
 obtaining a plurality of Pair Hidden Markov Models (PHMMs) including the PHMM, wherein each of the plurality of PHMMs corresponds to a distinct reference population; and   determining the initial geographical ancestry classifications as corrected based on a weighting of the plurality of PHMMs.   
     
     
         15 . The system of  claim 14 , wherein determining the initial geographical ancestry classifications as corrected comprises:
 applying the initial geographical ancestry classifications associated with segments of the two phased haplotypes to the plurality of PHMMs, and performing a Bayesian model averaging on outputs of the plurality of PHMMs.   
     
     
         16 . A non-transitory computer-readable medium storing program instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising:
 obtaining, from a classifier, a plurality of initial geographical ancestry classifications respectively associated with segments of two phased haplotypes of a chromosome pair of an individual, each phased haplotype being inherited from one of two parents of the individual, each segment of one of the two phased haplotypes corresponding to a segment of another of the two phased haplotypes;   obtaining a Pair Hidden Markov Model (PHMM) in which observed states correspond to ordered pairs of the initial geographical ancestry classifications associated with two corresponding segments of the two phased haplotypes, and in which hidden states correspond to ordered pairs of underlying geographical ancestry classifications associated with the two corresponding segments of the two phased haplotypes;   determining, using the PHMM, two geographical ancestry classifications respectively associated with the two phased haplotypes;   correcting the initial geographical ancestry classifications using the two geographical ancestry classifications; and   outputting the initial geographical ancestry classifications as corrected.   
     
     
         17 . The non-transitory computer-readable medium of  claim 16 , wherein determining the two geographical ancestry classifications comprises:
 determining a most likely sequence of the hidden states given the initial geographical ancestry classifications.   
     
     
         18 . The non-transitory computer-readable medium of  claim 17 , wherein the two geographical ancestry classifications include two sequences of geographical ancestry classifications, and wherein the most likely sequence of the hidden states comprises the two sequences of geographical ancestry classifications. 
     
     
         19 . The non-transitory computer-readable medium of  claim 17 , wherein determining the most likely sequence of hidden states comprises detecting a correlated prediction error that was caused by the classifier. 
     
     
         20 . The non-transitory computer-readable medium of  claim 16 , wherein correcting the initial geographical ancestry classifications comprises:
 obtaining a plurality of Pair Hidden Markov Models (PHMMs) including the PHMM, wherein each of the plurality of PHMMs corresponds to a distinct reference population; and   determining the initial geographical ancestry classifications as corrected based on a weighting of the plurality of PHMMs.

Join the waitlist — get patent alerts

Track US2023402132A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.