Combined and transfer learning of a variant pathogenicity predictor using gapped and non-gapped protein samples
Abstract
The technology disclosed relates to training a pathogenicity predictor. In particular, the technology disclosed relates to accessing a gapped training set that includes respective gapped protein samples for respective positions in a proteome, accessing a non-gapped training set that includes non-gapped benign protein samples and non-gapped pathogenic protein samples, generating respective gapped spatial representations for the gapped protein samples, and generating respective non-gapped spatial representations for the non-gapped benign protein samples and the non-gapped pathogenic protein samples, training a pathogenicity predictor over one or more training cycles and generating a trained pathogenicity predictor, wherein each of the training cycles uses as training examples gapped spatial representations from the respective gapped spatial representations and non-gapped spatial representations from the respective non-gapped spatial representations, and using the trained pathogenicity classifier to determine pathogenicity of variants.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method of training a pathogenicity predictor, including:
accessing a gapped training set that includes respective gapped protein samples for respective positions in a proteome; accessing a non-gapped training set that includes non-gapped benign protein samples and non-gapped pathogenic protein samples; generating respective gapped spatial representations for the gapped protein samples, and generating respective non-gapped spatial representations for the non-gapped benign protein samples and the non-gapped pathogenic protein samples; training a pathogenicity predictor over one or more training cycles and generating a trained pathogenicity predictor, wherein each of the training cycles uses as training examples gapped spatial representations from the respective gapped spatial representations and non-gapped spatial representations from the respective non-gapped spatial representations; and using the trained pathogenicity classifier to determine pathogenicity of variants.
2 . The computer-implemented method of claim 1 , wherein the respective gapped protein samples are labelled with respective gapped ground truth sequences.
3 . The computer-implemented method of claim 2 , wherein a particular gapped ground truth sequence for a particular gapped protein sample has a benign label for a particular amino acid class that corresponds to a reference amino acid at a particular position in the particular gapped protein.
4 . The computer-implemented method of claim 3 , wherein the particular gapped protein sample has respective pathogenic labels for respective remaining amino acid classes that correspond to alternate amino acids at the particular position.
5 . The computer-implemented method of claim 1 , wherein a particular non-gapped benign protein sample includes a benign alternate amino acid at a particular position substituted by a benign nucleotide variant.
6 . The computer-implemented method of claim 5 , wherein a particular non-gapped pathogenic protein sample includes a pathogenic alternate amino acid at a particular position substituted by a pathogenic nucleotide variant.
7 . The computer-implemented method of claim 6 , wherein the particular non-gapped benign protein sample is labelled with a benign ground truth sequence that has a benign label for a particular amino acid class that corresponds to the benign alternate amino acid.
8 . The computer-implemented method of claim 7 , wherein the benign ground truth sequence respective masked labels for respective remaining amino acid classes that correspond to amino acids that are different from the benign alternate amino acid.
9 . The computer-implemented method of claim 8 , wherein the particular non-gapped pathogenic protein sample is labelled with a pathogenic ground truth sequence that has a pathogenic label for a particular amino acid class that corresponds to the pathogenic alternate amino acid.
10 . The computer-implemented method of claim 9 , wherein the pathogenic ground truth sequence has respective masked labels for respective remaining amino acid classes that correspond to amino acids that are different from the pathogenic alternate amino acid.
11 . The computer-implemented method of claim 1 , further including using a sample indicator to indicate to the pathogenicity predictor whether a current training example is a gapped spatial representation for a gapped protein sample, or a non-gapped spatial representation for a non-gapped protein sample.
12 . The computer-implemented method of claim 3 , further including masking the benign label for the particular amino acid class that corresponds to the reference amino acid at the particular position in the particular gapped protein.
13 . The computer-implemented method of claim 1 , wherein the non-gapped benign protein samples are derived from common human and non-human primate nucleotide variants.
14 . The computer-implemented method of claim 1 , wherein the non-gapped pathogenic protein samples are derived from combinatorically simulated nucleotide variants.
15 . The computer-implemented method of claim 1 , wherein the pathogenicity predictor generates an amino acid class-wise output sequence in response to processing a training example,
wherein the amino acid class-wise output sequence has amino acid class-wise pathogenicity scores.
16 . The computer-implemented method of claim 1 , further including measuring performance of the trained pathogenicity predictor between training cycles over a validation set.
17 . The computer-implemented method of claim 16 , wherein the validation set includes a pair of gapped and non-gapped spatial representations for each held-out protein sample.
18 . The computer-implemented method of claim 17 , wherein the trained pathogenicity predictor generates a first amino acid class-wise output sequence for the gapped spatial representation in the pair, and a second amino acid class-wise output sequence for the non-gapped spatial representation in the pair,
wherein a final pathogenicity score for a nucleotide variant that causes an amino acid substitution in a held-out protein sample is determined based on a combination of first and second pathogenicity scores for the amino acid substitution in the first and second amino acid class-wise output sequences.
19 . The computer-implemented method of claim 18 , wherein the final pathogenicity score is based on an average of the first and second pathogenicity scores.
20 . The computer-implemented method of claim 1 , wherein at least some of the training cycles use a same of number of gapped spatial representations and non-gapped spatial representations.
21 . The computer-implemented method of claim 1 , wherein at least some of the training cycles use batches of training examples that have a same of number of gapped spatial representations and non-gapped spatial representations.
22 . The computer-implemented method of claim 1 , wherein a masked label does not contribute to error determination, and therefore does not contribute to training of the pathogenicity predictor.
23 . The computer-implemented method of claim 22 , wherein the masked label is zeroed-out.
24 . The computer-implemented method of claim 1 , wherein the gapped spatial representations are weighted differently from the non-gapped spatial representations, such that a contribution of the gapped spatial representations to gradient updates applied to parameters of the pathogenicity predictor in response to the pathogenicity predictor processing the non-gapped spatial representations varies from a contribution of the non-gapped spatial representations to gradient updates applied to the parameters of the pathogenicity predictor in response to the pathogenicity predictor processing the non-gapped spatial representations.
25 . The computer-implemented method of claim 24 , wherein the variation is determined by pre-defined weights.
26 . A computer-implemented method of training a pathogenicity predictor, including:
starting with training a pathogenicity classifier on a gapped training set and generating a trained pathogenicity classifier; further training the trained pathogenicity classifier on a non-gapped training set and generating a retrained pathogenicity classifier; and using the retrained pathogenicity classifier to determine pathogenicity of variants.
27 . The computer-implemented method of claim 26 , further including measuring performance of the trained pathogenicity predictor between training cycles over a first validation set that includes only non-gapped spatial representations of held-out protein samples.
28 . The computer-implemented method of claim 27 , further including measuring performance of the retrained pathogenicity predictor between training cycles over a second validation set that includes gapped spatial representations and non-gapped spatial representations of held-out protein samples.
29 . The computer-implemented method of claim 28 , wherein the retrained pathogenicity predictor generates a first amino acid class-wise output sequence for the pair in response to processing the pair,
wherein a final pathogenicity score for a nucleotide variant that causes an amino acid substitution in a corresponding held-out protein sample is determined based on the first amino acid class-wise output sequence.
30 . A computer-implemented method of training a pathogenicity predictor, including:
accessing a gapped training set that includes respective gapped protein samples for respective positions in a proteome, wherein the respective gapped protein samples are labelled with respective gapped ground truth sequences, wherein a particular gapped ground truth sequence for a particular gapped protein sample has a benign label for a particular amino acid class that corresponds to a reference amino acid at a particular position in the particular gapped protein, and has respective pathogenic labels for respective remaining amino acid classes that correspond to alternate amino acids at the particular position; accessing a non-gapped training set that includes non-gapped benign protein samples and non-gapped pathogenic protein samples, wherein a particular non-gapped benign protein sample includes a benign alternate amino acid at a particular position substituted by a benign nucleotide variant, wherein a particular non-gapped pathogenic protein sample includes a pathogenic alternate amino acid at a particular position substituted by a pathogenic nucleotide variant, wherein the particular non-gapped benign protein sample is labelled with a benign ground truth sequence that has a benign label for a particular amino acid class that corresponds to the benign alternate amino acid, and respective masked labels for respective remaining amino acid classes that correspond to amino acids that are different from the benign alternate amino acid, and wherein the particular non-gapped pathogenic protein sample is labelled with a pathogenic ground truth sequence that has a pathogenic label for a particular amino acid class that corresponds to the pathogenic alternate amino acid, and respective masked labels for respective remaining amino acid classes that correspond to amino acids that are different from the pathogenic alternate amino acid; generating respective gapped spatial representations for the gapped protein samples, and generating respective non-gapped spatial representations for the non-gapped benign protein samples and the non-gapped pathogenic protein samples; training a pathogenicity predictor over one or more training cycles, and generating a trained pathogenicity predictor, wherein each of the training cycles uses as training examples gapped spatial representations from the respective gapped spatial representations, and non-gapped spatial representations from the respective non-gapped spatial representations; and using the trained pathogenicity classifier to determine pathogenicity of variants.Join the waitlist — get patent alerts
Track US2023108368A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.