Variant pathogenicity prediction using neural network
Abstract
The technology disclosed relates to constructing a computer-implemented method for variant classification. In particular, the method includes using a pathogenicity prediction neural network to process as input, (i) a reference protein sequence that has a first chain of amino acids with at least twenty amino acids, (ii) an alternative protein sequence aligned with the reference sequence, where the alternative protein sequence has a second chain of amino acids with at least twenty amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution, and (iii) a primate conservation profile generated using a primate cross-species multiple sequence alignment that aligns the reference protein sequence with other protein sequences from primate species. The method further includes based on the processing of the input by the neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method, including:
using a pathogenicity prediction neural network to process as input:
(i) a reference protein sequence that has a first chain of amino acids with at least twenty amino acids,
(ii) an alternative protein sequence aligned with the reference protein sequence, wherein the alternative protein sequence has a second chain of amino acids with at least twenty amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution,
(iii) a primate conservation profile generated using a primate cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from primate species,
(iv) a mammal conservation profile generated using a mammal cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from mammal species, and
(v) a vertebrate conservation profile generated using a vertebrate cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from vertebrate species; and
based on the processing of the input by the pathogenicity prediction neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
2 . A computer-implemented method, including:
using a pathogenicity prediction neural network to process as input:
(i) a reference protein sequence that has a first chain of amino acids with at least twenty amino acids,
(ii) an alternative protein sequence aligned with the reference protein sequence, wherein the alternative protein sequence has a second chain of amino acids with at least twenty amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution,
(iii) a primate conservation profile generated using a primate cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from primate species, and
(iv) a mammal conservation profile generated using a mammal cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from mammal species; and
based on the processing of the input by the pathogenicity prediction neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
3 . A computer-implemented method, including:
using a pathogenicity prediction neural network to process as input:
(i) a reference protein sequence that has a first chain of amino acids with at least twenty amino acids,
(ii) an alternative protein sequence aligned with the reference protein sequence, wherein the alternative protein sequence has a second chain of amino acids with at least twenty amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution,
(iii) a primate conservation profile generated using a primate cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from primate species, and
(iv) a vertebrate conservation profile generated using a vertebrate cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from vertebrate species; and
based on the processing of the input by the pathogenicity prediction neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
4 . A computer-implemented method, including:
using a pathogenicity prediction neural network to process as input:
(i) a reference protein sequence that has a first chain of amino acids,
(ii) an alternative protein sequence aligned with the reference protein sequence, wherein the alternative protein sequence has a second chain of amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution, and
(iii) a primate conservation profile generated using a cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from primate species; and
based on the processing of the input by the pathogenicity prediction neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
5 . A computer-implemented method, including:
using a pathogenicity prediction neural network to process as input:
(i) a reference protein sequence that has a first chain of amino acids,
(ii) an alternative protein sequence aligned with the reference protein sequence, wherein the alternative protein sequence has a second chain of amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution, and
(iii) a mammal conservation profile generated using a cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from mammal species; and
based on the processing of the input by the pathogenicity prediction neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
6 . A computer-implemented method, including:
using a pathogenicity prediction neural network to process as input:
(i) a reference protein sequence that has a first chain of amino acids,
(ii) an alternative protein sequence aligned with the reference protein sequence, wherein the alternative protein sequence has a second chain of amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution, and
(iii) a vertebrate conservation profile generated using a cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences from vertebrate species; and
based on the processing of the input by the pathogenicity prediction neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
7 . A computer-implemented method, including:
using a pathogenicity prediction neural network to process as input:
(i) a reference protein sequence that has a first chain of amino acids,
(ii) an alternative protein sequence aligned with the reference protein sequence, wherein the alternative protein sequence has a second chain of amino acids, and the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution, and
(iii) at least one conservation profile generated using a cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences; and
based on the processing of the input by the pathogenicity prediction neural network, generating as output a pathogenicity prediction for the nucleotide substitution.
8 . The computer-implemented method of claim 7 , wherein the conservation profile is a position-specific frequency matrix.
9 . The computer-implemented method of claim 7 , wherein the conservation profile is a position-specific score matrix.
10 . The computer-implemented method of claim 7 , wherein the conservation profile is a primate conservation profile generated using the cross-species multiple sequence alignment that aligns the reference protein sequence with the plurality of other protein sequences from primate species.
11 . The computer-implemented method of claim 7 , wherein the conservation profile is a mammal conservation profile generated using the cross-species multiple sequence alignment that aligns the reference protein sequence with the plurality of other protein sequences from mammal species.
12 . The computer-implemented method of claim 7 , wherein the conservation profile is a vertebrate conservation profile generated using the cross-species multiple sequence alignment that aligns the reference protein sequence with the plurality of other protein sequences from vertebrate species.
13 . The computer-implemented method of claim 7 , wherein the input further includes a three-state secondary structure prediction for each amino acid in the first chain of amino acids.
14 . The computer-implemented method of claim 13 , wherein a secondary structure neural network processes the conservation profile and generates the three-state secondary structure prediction.
15 . The computer-implemented method of claim 7 , wherein the input further includes a three-state solvent accessibility prediction for each amino acid in the first chain of amino acids.
16 . The computer-implemented method of claim 15 , wherein a solvent accessibility neural network processes the conservation profile and generates the three-state solvent accessibility prediction.
17 . A computer-implemented method, including:
processing as input:
(i) a reference protein sequence that has a first chain of amino acids,
(ii) an alternative protein sequence that has a second chain of amino acids, wherein the first and second chains of amino acids differ by a variant amino acid caused by a nucleotide substitution, and
(iii) at least one conservation profile generated using a cross-species multiple sequence alignment that aligns the reference protein sequence with a plurality of other protein sequences; and
based on the processing of the input, generating as output a pathogenicity prediction for the nucleotide substitution.Join the waitlist — get patent alerts
Track US2022237457A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.