US2023061512A1PendingUtilityA1

Methods and systems for determining ancestral relatedness

Assignee: EMBARK VETERINARY INCPriority: Nov 18, 2019Filed: Aug 26, 2022Published: Mar 2, 2023
Est. expiryNov 18, 2039(~13.3 yrs left)· nominal 20-yr term from priority
Y02A90/10G16B 10/00G16B 20/20G16B 30/10G16B 20/40
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure provides methods of estimating a degree of ancestral relatedness between individuals. In an aspect, a method comprises receiving haplotype data comprising genetic markers shared among a population of individuals; dividing the haplotype data into segments based on the genetic markers; for each of the population of test individuals: (i) based on the genetic markers, matching segments of the haplotype data that are identical-by-descent between two individuals, (ii) for each of the matched segments: dividing the matched segment into discrete genomic intervals, scoring each of the discrete genomic intervals based on a degree of matching within or between the individuals, correcting the scores for consistency, and (iii) calculating a weighted sum over the discrete genomic intervals of the matched segment, based on the corrected scores and assigned weights; and (d) estimating the degree of ancestral relatedness between the individuals based on the weighted sums of the matched segments.

Claims

exact text as granted — not AI-modified
1 .- 129 . (canceled) 
     
     
         130 . A computer-implemented method for estimating a degree of ancestral relatedness between two individuals of a diploid population, comprising:
 (a) receiving haplotype data for a population of test individuals, the haplotype data comprising a plurality of genetic markers shared among the population of test individuals, wherein the plurality of genetic markers comprises at least about 1,000 distinct genetic markers;   (b) dividing the haplotype data into segments based on the plurality of genetic markers;   (c) for each of the population of test individuals:
 (i) based on the plurality of genetic markers, matching segments of the haplotype data that are identical-by-descent (IBD) between a first individual and a second individual among the population of test individuals, each of the matched segments having a size that is at least about 100 kilobase pairs (kbp); 
 (ii) for each of the matched segments between the first individual and the second individual:
 dividing the matched segment into a plurality of discrete genomic intervals; 
 
   scoring each of the plurality of discrete genomic intervals based on a degree of homozygosity matching of the discrete genomic interval within the first individual or the second individual, wherein the degree of homozygosity matching is a number of matching homozygous haplotypes within a discrete genomic interval within a single individual, thereby generating a plurality of scores;   and   assigning a plurality of weights to the plurality of discrete genomic intervals, based at least in part on the plurality of scores; and   (iii) calculating a weighted sum over the plurality of discrete genomic intervals of the matched segment, based on the plurality of scores and the plurality of weights; and   (d) estimating the degree of ancestral relatedness between the first individual and the second individual based on the weighted sums of the matched segments.   
     
     
         131 . The method of  claim 130 , wherein the diploid population is a mammal population. 
     
     
         132 . The method of  claim 130 , wherein the haplotype data is generated at least in part by processing genotype data of the population of test individuals using a haplotype phasing algorithm. 
     
     
         133 . The method of  claim 132 , wherein the haplotype phasing algorithm comprises a reference-based haplotype phasing algorithm comprising a Hidden Markov Model (HMM)-based search. 
     
     
         134 . The method of  claim 132 , wherein the genotype data is obtained at least in part by assaying biological samples obtained from the population of test individuals or derivatives thereof. 
     
     
         135 . The method of  claim 134 , wherein the assaying further comprises use of array hybridization. 
     
     
         136 . The method of  claim 134 , wherein the assaying further comprises sequencing the biological samples to generate a plurality of sequencing reads. 
     
     
         137 . The method of  claim 136 , wherein the assaying further comprises aligning the plurality of sequencing reads to a reference genome. 
     
     
         138 . The method of  claim 130 , wherein the plurality of genetic markers comprises at least about 10,000 distinct genetic markers. 
     
     
         139 . The method of  claim 130 , wherein each of the matched segments has a size that is at least about 500 kilobase pairs (kbp). 
     
     
         140 . The method of  claim 130 , wherein each of the matched segments comprises at least about 30 distinct genetic markers. 
     
     
         141 . The method of  claim 130 , further comprising dividing the matched segments such that the discrete genomic intervals of the plurality of discrete genomic intervals have an equal size. 
     
     
         142 . The method of  claim 130 , further comprising dividing the matched segments such that the discrete genomic intervals of the plurality of discrete genomic intervals have a variable size. 
     
     
         143 . The method of  claim 142 , wherein the variable size of a given discrete genomic interval of the plurality of discrete genomic intervals is determined based at least in part on a start position and an end position of IBD matches proximal to the given discrete genomic interval, a density of genetic markers in the given discrete genomic interval, a maximum number of markers for the given discrete genomic interval, a maximum length of the given discrete genomic interval, or a combination thereof. 
     
     
         144 . The method of  claim 130 , further comprising scoring each of the plurality of discrete genomic intervals based on the degree of homozygosity matching and a degree of pairwise matching of the discrete genomic interval between the first individual and the second individual, wherein the degree of pairwise matching is a number of matching haplotypes within a discrete genomic interval between two individuals. 
     
     
         145 . The method of  claim 144 , further comprising correcting the plurality of scores based on a consistency between degree of homozygosity matching and degree of pairwise matching. 
     
     
         146 . The method of  claim 145 , further comprising correcting pairwise matching scores based on a consistency with a corresponding homozygosity matching score. 
     
     
         147 . The method of  claim 144 , further comprising assigning the plurality of weights to the plurality of discrete genomic intervals, based at least in part on a plurality of identity states for two alleles in two diploid individuals,
 wherein zero weights are assigned to discrete genomic intervals with identity states indicative of no pairwise matching between two diploid individuals, and   wherein non-zero weights are assigned only to discrete genomic intervals with identity states indicative of non-zero pairwise matching between two diploid individuals.   
     
     
         148 . The method of  claim 147 , wherein the plurality of identity states comprises identity states selected from the following: 
       
         
           
                 
               
                     
                 
                   Identity state Probability Contribution to f xy  Contribution to r xy   
                 
                     
                 
                     
                 
                 
                 
                 
                 
                 
               
                     
                   
                     
                       
                       
                           
                           
                       
                     
                   
                   Δ 1    
                   1 
                   1 
                 
                     
                 
                     
                   a—b 
                   Δ 2   
                   0 
                   0 
                 
                     
                   c—d 
                     
                     
                     
                 
                     
                 
                     
                   
                     
                       
                       
                           
                           
                       
                     
                   
                   Δ 3   
                   1/2  
                   3/4  
                 
                     
                 
                     
                   a—b 
                   Δ 4   
                   0 
                   0 
                 
                     
                   c d 
                     
                     
                     
                 
                     
                 
                     
                   
                     
                       
                       
                           
                           
                       
                     
                   
                   Δ 5   
                   1/2  
                   3/4  
                 
                     
                 
                     
                   a b 
                   Δ 6   
                   0 
                   0 
                 
                     
                   c—d 
                     
                     
                     
                 
                     
                 
                     
                   
                     
                       
                       
                           
                           
                       
                     
                   
                   Δ 7   
                   1/2  
                   1  
                 
                     
                 
                     
                   
                     
                       
                       
                           
                           
                       
                     
                   
                   Δ 8   
                   1/4  
                   1/2  
                 
                     
                 
                     
                   a b 
                   Δ 9   
                   0 
                    0, 
                 
                     
                   c d 
                 
                     
                 
             
                
                
                
               
               
                
               
            
             
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
                
               
            
           
         
         wherein the first individual x has alleles a and b, 
         wherein the second individual y has alleles c and d, 
         wherein horizontal lines of the nine identity states indicate homozygosity in an individual from identity by descent, and 
         wherein the plurality of weights are assigned further based on a plurality of contributions to the relatedness r xy . 
       
     
     
         149 . The method of  claim 148 , wherein the degree of ancestral relatedness comprises a coefficient of relatedness. 
     
     
         150 . The method of  claim 149 , further comprising calculating the weighted sum over the plurality of discrete genomic intervals of the matched segment, wherein the weighted sum is expressed by: 
       
         
           
             
               
                 
                   r 
                   
                     x 
                     ⁢ 
                     y 
                   
                 
                 = 
                 
                   
                     
                       Δ 
                       ⁢ 
                       1 
                     
                     + 
                     
                       Δ 
                       ⁢ 
                       7 
                     
                     + 
                     
                       ( 
                       
                         0.75 
                         × 
                         Δ3 
                       
                       ) 
                     
                     + 
                     
                       ( 
                       
                         0.5 
                         × 
                         Δ8 
                       
                       ) 
                     
                   
                   L 
                 
               
               , 
             
           
         
         wherein Δ1, Δ3, Δ7, and Δ8 represent the sum total of the genome length assigned to one of the four match count states State 1, State 3, State 7, and State 8, respectively, 
         wherein State 1={Pairwise=4, Homozygous=2}, 
         wherein State 3={Pairwise=2, Homozygous=1}, 
         wherein State 7={Pairwise=2, Homozygous=0}, 
         wherein State 8={Pairwise=1, Homozygous=0}, and 
         wherein L is the total length of the genome considered. 
       
     
     
         151 . The method of  claim 147 , wherein the degree of ancestral relatedness comprises a coefficient of kinship. 
     
     
         152 . The method of  claim 151 , further comprising calculating the weighted sum over the plurality of discrete genomic intervals of the matched segment, wherein the weighted sum is expressed by: 
       
         
           
             
               
                 
                   k 
                   
                     x 
                     ⁢ 
                     y 
                   
                 
                 = 
                 
                   
                     
                       Δ 
                       ⁢ 
                       1 
                     
                     + 
                     
                       ( 
                       
                         0.5 
                         × 
                         
                           ( 
                           
                             Δ3 
                             + 
                             Δ7 
                           
                           ) 
                         
                       
                       ) 
                     
                     + 
                     
                       ( 
                       
                         0.25 
                         × 
                         Δ8 
                       
                       ) 
                     
                   
                   L 
                 
               
               , 
             
           
         
         wherein Δ1, Δ3, Δ7, and Δ8 represent the sum total of the genome length assigned to one of the four match count states State 1, State 3, State 7, and State 8, respectively, 
         wherein State 1={Pairwise=4, Homozygous=2}, 
         wherein State 3={Pairwise=2, Homozygous=1}, 
         wherein State 7={Pairwise=2, Homozygous=0}, 
         wherein State 8={Pairwise=1, Homozygous=0}, and 
         wherein L is the total length of the genome considered. 
       
     
     
         153 . The method of  claim 130 , wherein estimating the degree of ancestral relatedness between the first individual and the second individual comprises determining a degree of inbreeding of the first individual or the second individual. 
     
     
         154 . The method of  claim 153 , further comprising determining a familial relationship between the first individual and the second individual based at least in part on the degree of inbreeding of the first individual and the second individual. 
     
     
         155 . The method of  claim 154 , wherein the familial relationship is a parent-child relationship, a sibling relationship, an aunt/uncle-nephew/niece relationship, a cousin relationship, or a grandparent-grandchild relationship. 
     
     
         156 . The method of  claim 130 , further comprising generating a social connection between a first person associated with the first individual and a second person associated with the second individual, based at least in part on the estimated degree of ancestral relatedness between the first individual and the second individual. 
     
     
         157 . The method of  claim 130 , further comprising identifying a familial relationship between the first individual and the second individual based at least in part on the degree of ancestral relatedness, wherein the familial relationship is a parent-child relationship, a sibling relationship, an aunt/uncle-nephew/niece relationship, a cousin relationship, or a grandparent-grandchild relationship. 
     
     
         158 . A non-transitory computer readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements a method for estimating a degree of ancestral relatedness between two individuals of a diploid population, the method comprising:
 (a) receiving haplotype data for a population of test individuals, the haplotype data comprising a plurality of genetic markers shared among the population of test individuals, wherein the plurality of genetic markers comprises at least about 1,000 distinct genetic markers;   (b) dividing the haplotype data into segments based on the plurality of genetic markers;   (c) for each of the population of test individuals:
 (i) based on the plurality of genetic markers, matching segments of the haplotype data that are identical-by-descent (IBD) between a first individual and a second individual among the population of test individuals, each of the matched segments having a size that is at least about 100 kilobase pairs (kbp); 
 (ii) for each of the matched segments between the first individual and the second individual: 
   dividing the matched segment into a plurality of discrete genomic intervals;   scoring each of the plurality of discrete genomic intervals based on a degree of homozygosity matching of the discrete genomic interval within the first individual or the second individual, wherein the degree of homozygosity matching is a number of matching homozygous haplotypes within a discrete genomic interval within a single individual, thereby generating a plurality of scores;   and   assigning a plurality of weights to the plurality of discrete genomic intervals, based at least in part on the plurality of scores; and
 (iii) calculating a weighted sum over the plurality of discrete genomic intervals of the matched segment, based on the plurality of scores and the plurality of weights; and 
   (d) estimating the degree of ancestral relatedness between the first individual and the second individual based on the weighted sums of the matched segments.   
     
     
         159 . A computer system for estimating a degree of ancestral relatedness between two individuals of a diploid population, comprising:
 a database that is configured to store haplotype data for a population of test individuals, the haplotype data comprising a plurality of genetic markers shared among the population of test individuals, wherein the plurality of genetic markers comprises at least about 1,000 distinct genetic markers; and   one or more computer processors operatively coupled to the database, wherein the one or more computer processors are individually or collectively programmed to:
 (a) divide the haplotype data into segments based on the plurality of genetic markers; 
 (b) for each of the population of test individuals: 
   (i) based on the plurality of genetic markers, match segments of the haplotype data that are identical-by-descent (IBD) between a first individual and a second individual among the population of test individuals, each of the matched segments having a size that is at least about 100 kilobase pairs (kbp);   (ii) for each of the matched segments between the first individual and the second individual:
 divide the matched segment into a plurality of discrete genomic intervals; 
 score each of the plurality of discrete genomic intervals based on a degree of homozygosity matching of the discrete genomic interval within the first individual or the second individual, wherein the degree of homozygosity matching is a number of matching homozygous haplotypes within a discrete genomic interval within a single individual, thereby generating a plurality of scores; and 
 assign a plurality of weights to the plurality of discrete genomic intervals, based at least in part on the plurality of scores; and 
   (iii) calculate a weighted sum over the plurality of discrete genomic intervals of the matched segment, based on the plurality of scores and the plurality of weights; and
 (c) estimate the degree of ancestral relatedness between the first individual and the second individual based on the weighted sums of the matched segments.

Join the waitlist — get patent alerts

Track US2023061512A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.