Allelotyping Methods for Massively Parallel Sequencing
Abstract
In one illustrative embodiment, an allelotyping method may include selecting a plurality of text strings that each represent a nucleotide sequence that was read by a massively parallel sequencing (MPS) instrument, where the nucleotide sequences represented by the selected plurality of text strings each correspond to a particular locus, comparing the selected plurality of text strings to one another to determine an abundance count for each unique text string included in the selected plurality of text strings, and determining one or more alleles for the particular locus by comparing the abundance count for each unique text string included in the selected plurality of text strings to an abundance threshold.
Claims
exact text as granted — not AI-modified1 . An allelotyping method comprising:
selecting a plurality of text strings that each represent a nucleotide sequence that was read by a massively parallel sequencing (MPS) instrument, wherein the nucleotide sequences represented by the selected plurality of text strings each correspond to a particular locus; comparing the selected plurality of text strings to one another to determine an abundance count for each unique text string included in the selected plurality of text strings; and determining one or more alleles for the particular locus by comparing the abundance count for each unique text string included in the selected plurality of text strings to an abundance threshold.
2 . The allelotyping method of claim 1 , wherein determining the one or more alleles for the particular locus comprises identifying one or more unique text strings that each represent a nucleotide sequence containing a short tandem repeat (STR).
3 . The allelotyping method of claim 1 , wherein determining the one or more alleles for the particular locus comprises identifying one or more unique text strings that each represent a nucleotide sequence containing a single nucleotide polymorphism (SNP).
4 . The allelotyping method of claim 1 , wherein comparing the abundance count for each unique text string included in the selected plurality of text strings to the abundance threshold comprises identifying the unique text string having a highest abundance count from among the selected plurality of text strings.
5 . The allelotyping method of claim 4 , wherein comparing the abundance count for each unique text string included in the selected plurality of text strings to the abundance threshold further comprises calculating whether a ratio of the abundance count for each unique text string compared to the highest abundance count exceeds the abundance threshold.
6 . The allelotyping method of claim 5 , wherein the abundance threshold is a percentage value in the range of 15% to 60%.
7 . The allelotyping method of claim 5 , wherein the abundance threshold is a percentage value configurable by a user.
8 . The allelotyping method of claim 5 , further comprising:
receiving a first text-based computer file comprising a plurality of text strings that each represent a nucleotide sequence that was read by the MPS instrument, prior to selecting the plurality of text strings for which the represented nucleotide sequences each correspond to the particular locus; and generating a second text-based computer file comprising each unique text string for which the ratio exceeds the abundance threshold, wherein a file size of the second text-based computer file is smaller than a file size of the first text-based computer file.
9 . The allelotyping method of claim 8 , wherein the second text-based computer file comprises one or more unique text strings that each represent a nucleotide sequence determined to be an allele for the particular locus.
10 . The allelotyping method of claim 8 , wherein the file size of the second text-based computer file is at least ten-thousand times smaller than the file size of the first text-based computer file.
11 . The allelotyping method of claim 1 , wherein the steps of (i) selecting the plurality of text strings for which the represented nucleotide sequences each correspond to a particular locus, (ii) comparing the selected plurality of text strings to one another to determine the abundance count for each unique text string included in the selected plurality of text strings, and (iii) determining one or more alleles for the particular locus by comparing the abundance counts to the abundance threshold are performed for each of a plurality of loci present in a sample that was read by the MPS instrument.
12 . The allelotyping method of claim 11 , wherein selecting the plurality of text strings for which the represented nucleotide sequences each correspond to one of the plurality of loci comprises:
determining whether each of a plurality of text strings generated by the MPS instrument when reading the sample includes text characters that represent a flanking nucleotide sequence associated with a particular locus; and selecting a plurality of text strings that include the text characters that represent the flanking nucleotide sequence associated with the particular locus.
13 . The allelotyping method of claim 12 , further comprising removing the text characters that represent the flanking nucleotide sequence from each of the selected plurality of text strings prior to comparing the selected plurality of text strings to one another to determine the abundance count for each unique text string included in the selected plurality of text strings.
14 . The allelotyping method of claim 1 , further comprising removing all text characters that do not represent a short tandem repeat (STR) from each of the selected plurality of text strings prior to comparing the selected plurality of text strings to one another to determine the abundance count for each unique text string included in the selected plurality of text strings.
15 . A computer-readable medium storing a text-based computer file, the text-based computer file comprising:
one or more records, each of the one or more records including:
a first text line;
a second text line comprising a text string representing a nucleotide sequence containing a single nucleotide polymorphism (SNP);
a third text line comprising a human-readable allele designation for the SNP; and
a fourth text line.
16 . The computer-readable medium of claim 15 , wherein, for each of the one or more records of the text-based file, the human-readable allele designation of the third text line comprises a number of attribute-value pairs that specify a first SNP state, a second SNP state, an abundance count of the first SNP state, an abundance count of the second SNP state, and a strand of the nucleotide sequence represented in the second text line.
17 . The computer-readable medium of claim 15 , wherein, for each of the one or more records of the text-based file, the human-readable allele designation of the third text line further comprises an attribute-value pair specifying a reference SNP identifier associated with the nucleotide sequence represented in the second text line.
18 . The computer-readable medium of claim 15 , wherein, for each of the one or more records of the text-based file, the first text line comprises a unique sequence identifier created by a massively parallel sequencing (VIPS) instrument when generating data related to the nucleotide sequence represented in the second text line.
19 . The computer-readable medium of claim 18 , wherein, for each of the one or more records of the text-based file, the first text line further comprises forensic metadata specifying one or more of a file format, a unique case identifier, a unique sample identifier, a unique laboratory identifier, and a unique technician identifier.
20 . The computer-readable medium of claim 15 , wherein, for each of the one or more records of the text-based file, the fourth text line comprises a text string representing quality scores associated with the nucleotide sequence represented in the second text line.Join the waitlist — get patent alerts
Track US2023197196A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.