Genetic and genealogical analysis for identification of birth location and surname information
Abstract
A system identifies ancestral birth locations or surnames estimated to be associated with an individual's ancestors using an individual's genetic sample. The system identifies users who are genetic matches to the individual and determines whether and how often a birth location or surname appears in the pedigrees of those users. Birth locations or surnames that appear frequently throughout the pedigrees of genetically matching users may represent birth locations or surnames that are affiliated with the individual's ancestors. The system determines whether the frequency of appearance of a birth location or surname is statistically significant to eliminate biases for certain birth locations or surnames that appear more frequently than others. The birth location or surname may be provided to the individual based on an also-determined enrichment score.
Claims
exact text as granted — not AI-modified1 . A method comprising:
receiving a genetic dataset of an individual; identifying a set of related individuals who are related to the individual based on the genetic dataset, wherein the set of related individuals are identified through an Identity-By-Descent (IBD) estimation based on shared segments of DNA data between the genetic dataset of the individual and genetic datasets of the set of related individuals wherein the shared segments are from phased genetic data of the individual and the set of related individuals; identifying a genetic group associated with the individual, the genetic group containing the set of related individuals; identifying a surname associated with the genetic group, wherein the surname is determined to be a significant surname to the genetic group based on a frequency of the surname in the genetic group; and outputting the surname.
2 . The method of claim 1 , wherein the genetic group corresponds to a geographic region.
3 . The method of claim 1 , wherein the surname is associated with a significance score that indicates a significance level of the surname to the genetic group.
4 . The method of claim 3 , wherein the significance score is determined based on a match frequency and a background frequency associated with the surname.
5 . The method of claim 1 , further comprising:
identifying one or more additional surnames that are significant to the genetic group; outputting the one or more additional surnames.
6 . The method of claim 5 , wherein the one or more additional surnames are each associated with a significance score to the genetic group and the one or more additional surnames are outputted in an order based on the significance scores.
7 . The method of claim 1 wherein the phased genetic data comprises a pair of haplotypes for the individual.
8 . The method of claim 7 wherein the set of related individuals are IBD matches based on the pair of haplotypes for the individual and haplotypes of the set of related individuals.
9 . The method of claim 1 , wherein identifying a surname associated with the genetic group comprises:
accessing one or more of pedigrees of the set of related individuals, each pedigree comprising a genealogical graph of relatives for a member of the set of related individuals; identifying a frequency of the surname in the genetic group and in the one or more of pedigrees; and determining the surname is signficant based on the frequency.
10 . A non-transitory computer-readable medium comprising computer program code, the computer program code when executed by a processor causing the processor to perform steps comprising:
receiving a genetic dataset of an individual; identifying a set of related individuals who are related to the individual based on the genetic dataset, wherein the set of related individuals are identified through an Identity-By-Descent (IBD) estimation based on shared segments of DNA data between the genetic dataset of the individual and genetic datasets of the set of related individuals wherein the shared segments are from phased genetic data of the individual and the set of related individuals; identifying a genetic group associated with the individual, the genetic group containing the set of related individuals; identifying a surname associated with the genetic group, wherein the surname is determined to be a significant surname to the genetic group based on a frequency of the surname in the genetic group; and outputting the surname.
11 . The non-transitory computer-readable medium of claim 10 , wherein the genetic group corresponds to a geographic region.
12 . The non-transitory computer-readable medium of claim 10 , wherein the surname is associated with a significance score that indicates a significance level of the surname to the genetic group.
13 . The non-transitory computer-readable medium of claim 12 , wherein the significance score is determined based on a match frequency and a background frequency associated with the surname.
14 . The non-transitory computer-readable medium of claim 10 , wherein the steps further comprising:
identifying one or more additional surnames that are significant to the genetic group; outputting the one or more additional surnames.
15 . The non-transitory computer-readable medium of claim 14 , wherein the one or more additional surnames are each associated with a significance score to the genetic group and the one or more additional surnames are outputted in an order based on the significance scores.
16 . The non-transitory computer-readable medium of claim 10 , wherein the phased genetic data comprises a pair of haplotypes for the individual.
17 . The non-transitory computer-readable medium of claim 16 , wherein the set of related individuals are IBD matches based on the pair of haplotypes for the individual and haplotypes of the set of related individuals.
18 . The non-transitory computer-readable medium of claim 10 , wherein identifying a surname associated with the genetic group comprises:
accessing one or more of pedigrees of the set of related individuals, each pedigree comprising a genealogical graph of relatives for a member of the set of related individuals; identifying a frequency of the surname in the genetic group and in the one or more of pedigrees; and determining the surname is signficant based on the frequency.
19 . A computer system comprising:
one or more processors; and a non-transitory computer readable storage medium storing instructions, when executed by one or more processors, causing the one or more processors to perform steps comprising:
receiving a genetic dataset of an individual;
identifying a set of related individuals who are related to the individual based on the genetic dataset, wherein the set of related individuals are identified through an Identity-By-Descent (IBD) estimation based on shared segments of DNA data between the genetic dataset of the individual and genetic datasets of the set of related individuals wherein the shared segments are from phased genetic data of the individual and the set of related individuals;
identifying a genetic group associated with the individual, the genetic group containing the set of related individuals;
identifying a surname associated with the genetic group, wherein the surname is determined to be a significant surname to the genetic group based on a frequency of the surname in the genetic group; and
outputting the surname.
20 . The system of claim 19 , wherein identifying a surname associated with the genetic group comprises steps:
accessing one or more of pedigrees of the set of related individuals, each pedigree comprising a genealogical graph of relatives for a member of the set of related individuals; identifying a frequency of the surname in the genetic group and in the one or more of pedigrees; and determining the surname is significant based on the frequency.Join the waitlist — get patent alerts
Track US2021183474A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.