Predicting variant pathogenicity from evolutionary conservation using three-dimensional (3d) protein structure voxels
Abstract
The technology disclosed relates to determining pathogenicity of nucleotide variants. In particular, the technology disclosed relates to specifying a particular amino acid at a particular position in a protein as a gap amino acid, and specifying remaining amino acids at remaining positions in the protein as non-gap amino acids, generating a gapped spatial representation of the protein that includes spatial configurations of the non-gap amino acids, and excludes a spatial configuration of the gap amino acid, determining an evolutionary conservation at the particular position of respective amino acids of respective amino acid classes based at least in part on the gapped spatial representation, and based at least in part on the evolutionary conservation of the respective amino acids, determining a pathogenicity of respective nucleotide variants that respectively substitute the particular amino acid with the respective amino acids in alternate representations of the protein.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method of determining pathogenicity of nucleotide variants, including:
specifying a particular amino acid at a particular position in a protein as a gap amino acid, and specifying remaining amino acids at remaining positions in the protein as non-gap amino acids; generating a gapped spatial representation of the protein that
includes spatial configurations of the non-gap amino acids, and
excludes a spatial configuration of the gap amino acid;
determining an evolutionary conservation at the particular position of respective amino acids of respective amino acid classes based at least in part on the gapped spatial representation; and based at least in part on the evolutionary conservation of the respective amino acids, determining a pathogenicity of respective nucleotide variants that respectively substitute the particular amino acid with the respective amino acids in alternate representations of the protein.
2 . The computer-implemented method of claim 1 , wherein the spatial configurations of the non-gap amino acids are encoded as amino acid class-wise distance channels,
wherein each of the amino acid class-wise distance channels has voxel-wise distance values for voxels in a plurality of voxels, and wherein the voxel-wise distance values specify distances from corresponding voxels in the plurality of voxels to atoms of the non-gap amino acids.
3 . The computer-implemented method of claim 2 , wherein the spatial configurations of the non-gap amino acids are determined based on spatial proximity between the corresponding voxels and the atoms of the non-gap amino acids.
4 . The computer-implemented method of claim 2 , wherein the spatial configuration of the gap amino acid is excluded from the gapped spatial representation by disregarding distances from the corresponding voxels to atoms of the gap amino acid when determining the voxel-wise distance values.
5 . The computer-implemented method of claim 4 , wherein the spatial configuration of the gap amino acid is excluded from the gapped spatial representation by disregarding spatial proximity between the corresponding voxels and the atoms of the gap amino acid.
6 . The computer-implemented method of claim 1 , wherein the particular amino acid is a reference amino acid that is a major allele of the protein.
7 . The computer-implemented method of claim 1 , wherein an evolutionary conservation predictor determines the evolutionary conservation by
processing, as input, the gapped spatial representation; and generating, as output, respective evolutionary conservation scores for the respective amino acids.
8 . The computer-implemented method of claim 7 , wherein the respective evolutionary conservation scores are rankable by magnitude.
9 . The computer-implemented method of claim 7 , further including classifying a nucleotide variant as pathogenic when an evolutionary conservation score generated by the evolutionary conservation predictor for a corresponding amino acid substitution is below a threshold.
10 . The computer-implemented method of claim 7 , further including classifying a nucleotide variant as pathogenic when an evolutionary conservation score generated by the evolutionary conservation predictor for a corresponding amino acid substitution is zero.
11 . The computer-implemented method of claim 7 , further including classifying a nucleotide variant as benign when an evolutionary conservation score generated by the evolutionary conservation predictor for a corresponding amino acid substitution is above a threshold.
12 . The computer-implemented method of claim 7 , further including classifying a nucleotide variant as benign when an evolutionary conservation score generated by the evolutionary conservation predictor for a corresponding amino acid substitution is non-zero.
13 . The computer-implemented method of claim 7 , wherein the evolutionary conservation predictor is trained on a conserved training set and a non-conserved training set.
14 . The computer-implemented method of claim 13 , wherein the conserved training set has respective conserved protein samples for respective conserved amino acids at respective positions in a proteome,
wherein the non-conserved training set has respective non-conserved protein samples for respective non-conserved amino acids at the respective positions.
15 . The computer-implemented method of claim 14 , wherein each of the respective positions has a set of conserved amino acids and a set of non-conserved amino acids.
16 . The computer-implemented method of claim 15 , wherein a particular set of conserved amino acids for a particular position in a particular protein in the proteome includes at least one major allele amino acid observed at the particular position across a plurality of species.
17 . The computer-implemented method of claim 16 , wherein the particular set of conserved amino acids includes one or more minor allele amino acids observed at the particular position across the plurality of species.
18 . The computer-implemented method of claim 17 , wherein a particular set of non-conserved amino acids for the particular position includes amino acids not in the particular set of conserved amino acids.
19 . The computer-implemented method of claim 18 , wherein the particular set of conserved amino acids and the particular set of non-conserved amino acids are identified based on evolutionary conservation profiles of homologous proteins of the plurality of species.
20 . The computer-implemented method of claim 19 , wherein the evolutionary conservation profiles of the homologous proteins are determined using a position-specific frequency matrix (PSFM).
21 . The computer-implemented method of claim 19 , wherein the evolutionary conservation profiles of the homologous proteins are determined using a position-specific scoring matrix (PSSM).
22 . The computer-implemented method of claim 16 , wherein the major allele amino acid is a reference amino acid.
23 . The computer-implemented method of claim 15 , wherein each of the respective positions has C conserved amino acids in the set of conserved amino acids,
wherein each of the respective positions has NC non-conserved amino acids in the set of non-conserved amino acids, where NC=20−C, wherein the conserved training set has CP conserved protein samples, where CP=a number of the respective positions*C, and wherein the non-conserved training set has NCP non-conserved protein samples, where NCP=the number of the respective positions*(20−C).
24 . The computer-implemented method of claim 23 , wherein the C ranges from one to ten.
25 . The computer-implemented method of claim 24 , wherein the C varies across the respective positions.
26 . The computer-implemented method of claim 25 , wherein the C is same for some of the respective positions.
27 . The computer-implemented method of claim 14 , wherein the respective conserved and non-conserved protein samples have respective gapped spatial representations generated by using respective reference amino acids at the respective positions as respective gap amino acids.
28 . The computer-implemented method of claim 7 , wherein the evolutionary conservation predictor trains on a particular conserved protein sample and estimates an evolutionary conservation of a particular conserved amino acid at a particular position in the particular conserved protein sample by
processing, as input,
a particular gapped spatial representation of the particular conserved protein sample,
wherein the particular gapped spatial representation is generated
by using a particular reference amino acid at the particular position as a gap amino acid, and
by using remaining amino acids at remaining positions in the particular conserved protein sample as non-gap amino acids; and
generating, as output, an evolutionary conservation score for the particular conserved amino acid.
29 . A non-transitory computer readable storage medium impressed with computer program instructions to determine pathogenicity of nucleotide variants, the instructions, when executed on a processor, implement a method comprising:
specifying a particular amino acid at a particular position in a protein as a gap amino acid, and specifying remaining amino acids at remaining positions in the protein as non-gap amino acids; generating a gapped spatial representation of the protein that
includes spatial configurations of the non-gap amino acids, and
excludes a spatial configuration of the gap amino acid;
determining an evolutionary conservation at the particular position of respective amino acids of respective amino acid classes based at least in part on the gapped spatial representation; and based at least in part on the evolutionary conservation of the respective amino acids, determining a pathogenicity of respective nucleotide variants that respectively substitute the particular amino acid with the respective amino acids in alternate representations of the protein.
30 . A system including one or more processors coupled to memory, the memory loaded with computer instructions to determine pathogenicity of nucleotide variants, the instructions, when executed on the processors, implement actions comprising:
specifying a particular amino acid at a particular position in a protein as a gap amino acid, and specifying remaining amino acids at remaining positions in the protein as non-gap amino acids; generating a gapped spatial representation of the protein that
includes spatial configurations of the non-gap amino acids, and
excludes a spatial configuration of the gap amino acid;
determining an evolutionary conservation at the particular position of respective amino acids of respective amino acid classes based at least in part on the gapped spatial representation; and based at least in part on the evolutionary conservation of the respective amino acids, determining a pathogenicity of respective nucleotide variants that respectively substitute the particular amino acid with the respective amino acids in alternate representations of the protein.Join the waitlist — get patent alerts
Track US2023108241A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.