Reducing error in predicted genetic relationships
Abstract
System, computer program products, and methods are disclosed for estimating a degree of ancestral relatedness between two individuals. The haplotype data for a population of individuals is divided into segment windows based on genetic markers, and matched segments for the haplotype data are generated. Each matched segment having a first cM width that exceeds a threshold cM width is included in counting the matched segments in each segment window. A weight associated with each segment window is estimated based on the count of matched segments in the associated segment window. A weighted sum of per-window cM widths for each matched segment is calculated based on the first cM width and the weights associated with the segment windows of the matched segment. The weighted sum of per-window cM widths are used to estimate a degree of ancestral relatedness between two individuals.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer implemented method, comprising:
receiving haplotype data for a set of individuals, the haplotype data including a plurality of genetic markers; dividing the haplotype data into segment windows; for each of at least two individuals in the set of individuals:
based on the genetic markers, matching segments of the haplotype data between the individual and other individuals in the set, each matched segment having a width,
counting the matched segments in each segment window, and
estimating a weight associated with each segment window based on the count of matched segments in the associated segment window;
calculating a weighted sum between the two individuals based on the width of each matched segment and the weights associated with the segment windows of the matched segment; and estimating a degree of ancestral relatedness between the two individuals based on the weighted sum between the two individuals.
2 . The computer-implemented method of claim 1 , wherein each matched segment has a width exceeding a threshold width, the threshold width is 5 centimorgan (cM), 6 cM, 7 cM, 8 cM, 9 cM, 10 cM, or any real number within a range of 5 cM to 10 cM.
3 . The computer-implemented method of claim 1 , wherein the weight associated with a segment window for individual A is approximated as:
w
i
A
=
P
rob
(
ι
RGH
)
Pro
H
)
Prob
=
k
i
)
,
wherein Prob( |RGH), Pro H), Prob =c) are estimates of a probability of an RGH segment given the count c of matched segments in the segment window i, a probability of an RGH segment in the segment window i, and the probability of the count c of matched segments in the segment window i, respectively.
4 . The computer-implemented method of claim 3 , wherein Prob 32 c) is approximated as:
Prob
=
c
c
>
0
)
=
(
n
c
)
B
(
c
+
α
,
n
-
c
+
β
)
B
(
α
,
β
)
,
wherein n is a size of the set of individuals and α, β are parameters of the Beta function B.
5 . The computer-implemented method of claim 4 , wherein α and β are estimated by using a maximum likelihood estimation of a joint distribution approximated as:
f
(
{
k
i
}
n
,
α
,
β
,
k
i
>
0
)
=
∏
i
=
1
K
(
n
k
i
-
1
)
B
(
k
i
-
1
+
α
,
n
-
k
i
+
1
+
β
)
B
(
α
,
β
)
.
6 . The computer-implemented method of claim 3 , wherein Pro H) is approximated by a maximum of:
1
Prob
(
RGH
)
Prob
=
c
)
.
7 . The computer-implemented method of claim 3 , wherein the Prob( |RGH) is based on a user specified threshold value V, and Prob |RGH, V) is approximated as:
Prob
(
RGH
,
V
,
c
>
0
)
=
(
n
c
)
B
(
c
+
α
,
n
-
c
+
β
)
B
(
α
,
β
)
/
Prob
(
0
≤
V
α
,
β
)
,
wherein n is a size of the set of individuals and α, β are parameters of the Beta function B.
8 . The computer-implemented method of claim 7 , wherein the user specified threshold value V may include two specified quantiles, a first specified quantile equal to 50%, 55%, 60%, 65%, 70%, or 75% and a second specified quantile equal to 75%, 80%, 85%, or 90%, to determine at least two estimates of Prob( |RGH, V).
9 . The computer-implemented method of claim 7 , wherein α and β is estimated by using a maximum likelihood estimation of a joint distribution approximated as:
f
(
{
m
i
}
n
,
α
,
β
,
V
,
m
i
>
0
)
=
∏
i
=
1
M
(
n
m
i
-
l
)
B
(
m
i
-
1
+
α
,
n
-
m
i
+
1
+
β
)
B
(
α
,
β
)
Prob
(
0
≤
V
α
,
β
)
.
10 . The computer-implemented method of claim 1 , wherein the weighted sum between two individuals A and B is calculated as:
c
M
2
A
,
B
=
∑
i
=
W
I
N
s
t
a
r
t
W
I
N
end
≤
K
w
i
A
·
w
i
B
·
cM
1
,
i
,
wherein cM 1,i is the width of a matched segment between individuals A and B for each segment window i that the matched segment spans starting at window WIN start and ending at window WIN end with w i A and w i B being the weights associated with the segment window i for individual A and B, respectively.
11 . The computer-implemented method of claim 1 , wherein estimating the weight comprises calculating temporary weights {tilde over (w)} c for the count c of matched segments in the associated segment window is approximated as:
w
~
c
=
1
,
if
c
≤
M
,
d
c
=
{
0
,
if
r
c
>
r
c
-
1
r
c
-
r
c
-
1
,
else
,
w
~
c
=
1
+
∑
i
=
M
+
1
c
d
i
,
if
c
>
M
,
wherein r c is the weight based on the count of matched segments c and approximated as:
r
c
=
P
rob
(
RGH
)
Pro
H
)
Prob
=
c
)
.
12 . The computer-implemented method of claim 1 , wherein the weights associated with each segment window decrease if the count of matched segments in the associated segment window increases.
13 . The computer-implemented method of claim 1 , wherein a size of the segment windows comprises 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 genetic markers or any number that falls within a range of 50 to 500 genetic markers.
14 . The computer-implemented method of claim 1 , further comprising:
responsive to matching segments of the haplotype data between the individual and other individuals in the set, removing close relatives from the set of individuals.
15 . The computer-implemented method of claim 14 , wherein close relatives comprise two individuals having a matched segment with a width larger than 60 cM.
16 . The computer-implemented method of claim 1 , wherein the set of individuals has a population size n and a number of genetic markers per segment window d, and a number of segment windows K is approximated as:
K=n/d.
17 . The computer-implemented method of claim 1 , wherein the degree of relatedness between two individuals comprises a probability that the two individuals are ancestrally related.
18 . The computer-implemented method of claim 1 , wherein the degree of relatedness between two individuals comprises a binary yes or no answer whether the two individuals are ancestrally related.
19 . A non-transitory computer readable medium for storing computer code comprising instructions, the instructions, when executed by one or more processors, cause the one or more processors to:
receive haplotype data for a set of individuals, the haplotype data including a plurality of genetic markers; divide the haplotype data into segment windows; for each of at least two individuals in the set of individuals:
based on the genetic markers, match segments of the haplotype data between the individual and other individuals in the set, each matched segment having a width,
count the matched segments in each segment window, and
estimate a weight associated with each segment window based on the count of matched segments in the associated segment window;
calculate a weighted sum between the two individuals based on the width of each matched segment and the weights associated with the segment windows of the matched segment; and estimate a degree of ancestral relatedness between the two individuals based on the weighted sum between the two individuals.
20 . A system comprising:
one or more processors; and a memory storing instructions that when executed by the one or more processors cause the one or more processors to perform steps comprising:
receiving haplotype data for a set of individuals, the haplotype data including a plurality of genetic markers;
dividing the haplotype data into segment windows;
for each of at least two individuals in the set of individuals:
based on the genetic markers, matching segments of the haplotype data between the individual and other individuals in the set, each matched segment having a width,
counting the matched segments in each segment window, and
estimating a weight associated with each segment window based on the count of matched segments in the associated segment window;
calculating a weighted sum between the two individuals based on the width of each matched segment and the weights associated with the segment windows of the matched segment; and
estimating a degree of ancestral relatedness between the two individuals based on the weighted sum between the two individuals.Join the waitlist — get patent alerts
Track US2020286591A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.