US2023207064A1PendingUtilityA1

Inter-model prediction score recalibration during training

Assignee: ILLUMINA INCPriority: Dec 29, 2021Filed: Sep 16, 2022Published: Jun 29, 2023
Est. expiryDec 29, 2041(~15.4 yrs left)· nominal 20-yr term from priority
G06N 3/0464G16B 30/10G16B 20/00G16B 10/00G16B 20/20G16B 40/00G06N 20/00G16B 30/00G16B 40/20G06N 20/20G06N 3/08G16B 50/10G06F 18/2111G06F 18/2148G06F 18/2155G06N 3/126G16B 40/30G16B 20/40G06N 3/045Y02A90/10G06N 3/044G06N 3/084G06N 3/047
74
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The technology disclosed relates to a system for inter-model prediction score recalibration. The system includes a first model that generates, based on evolutionary conservation summary statistics of amino acids in a reference protein sequence, a first set of pathogenicity scores with rankings for variants that mutate the reference sequence to alternate protein sequences. The system further includes a second model that generates, based on epistasis expressed by amino acid patterns spanning a multiple sequence alignment aligning the reference sequence to non-target sequences, a second set of pathogenicity scores with rankings for the variants. The system further includes a rank loss determination logic that determines a rank loss parameter by comparing the two sets of rankings, a loss function reconfiguration logic that reconfigures a loss function based on the rank loss parameter, and a training logic that uses the reconfigured loss function to train the first model.

Claims

exact text as granted — not AI-modified
What we claim is: 
     
         1 . A system, comprising:
 a first model configured to generate, based in part on evolutionary conservation summary statistics of amino acids in a reference target protein sequence, a first set of pathogenicity scores for a set of variants that mutate the reference target protein sequence to a set of alternate protein sequences, wherein the first set of pathogenicity scores has a first set of score rankings;   a second model configured to generate, based in part on epistasis expressed by amino acid patterns spanning a multiple sequence alignment that aligns the reference target protein sequence to a plurality of non-target protein sequences, a second set of pathogenicity scores for the set of variants, wherein the second set of pathogenicity scores has a second set of score rankings;   a rank loss determination logic configured to determine a rank loss parameter based on a comparison of the first set of score rankings against the second set of score rankings;   a loss function reconfiguration logic configured to reconfigure a loss function based on the rank loss parameter; and   a training logic configured to use the reconfigured loss function to train the first model.   
     
     
         2 . The system of  claim 1 , wherein the second model processes respective alternate protein sequences in the set of alternate protein sequences as respective inputs and generates respective pathogenicity scores in the second set of pathogenicity scores as respective outputs. 
     
     
         3 . The system of  claim 2 , wherein the second model is pre-trained to process the multiple sequence alignment as an input and generate a reconstruction of the multiple sequence alignment as an output. 
     
     
         4 . The system of  claim 3 , wherein the second model represents a reconstruction of a given alternate protein sequence as base-wise probability scores for each amino acid in the given alternate protein sequence. 
     
     
         5 . The system of  claim 4 , wherein a joint probability determined from the base-wise probability scores is used as a pathogenicity score for a given variant that mutates the reference target protein sequence to the given alternate protein sequence. 
     
     
         6 . The system of  claim 1 , wherein respective coefficient and latent space configurations of the second model are pre-trained to process and reconstruct respective multiple sequence alignments that have respective reference target protein sequences as respective query sequences. 
     
     
         7 . The system of  claim 6 , wherein the second model has a particular coefficient and latent space configuration corresponding to the reference target protein sequence. 
     
     
         8 . The system of  claims 1 , wherein the second model has one to twenty thousand coefficient and latent space configurations corresponding to one to twenty thousand reference protein sequences in human proteome. 
     
     
         9 . The system of  claim 1 , wherein the rank loss determination logic is further configured to determine the rank loss parameter based on a combination of the first set of score rankings and the second set of score rankings. 
     
     
         10 . The system of  claim 9 , wherein the combination is a weighted combination. 
     
     
         11 . The system of  claim 10 , wherein weights used to generate the weighted combination are preset. 
     
     
         12 . The system of  claim 11 , wherein the weights are differentiable and learned in a re-ranking layer that is trained as part of the training of the first model. 
     
     
         13 . The system of  claim 1 , wherein the second model is a variational autoencoder (VAE). 
     
     
         14 . The system of  claim 1 , wherein the second model is a generative adversarial network (GAN). 
     
     
         15 . The system of  claims 1 , further configured to comprise:
 a third model configured to generate, based in part on the epistasis expressed by the amino acid patterns spanning the multiple sequence alignment, a third set of pathogenicity scores for the set of variants, wherein the third set of pathogenicity scores has a third set of score rankings;   the rank loss determination logic further configured to determine the rank loss parameter based on a comparison of the first set of score rankings, the second set of score rankings, and the third set of score rankings;   the loss function reconfiguration logic further configured to reconfigure the loss function based on the rank loss parameter; and   the training logic further configured to use the reconfigured loss function to train the first model.   
     
     
         16 . The system of  claim 15 , wherein the third model is a Transformer-based model. 
     
     
         17 . The system of  claim 15 , wherein the rank loss determination logic is further configured to determine the rank loss parameter based on a combination of the first set of score rankings, the second set of score rankings, and the third set of score rankings. 
     
     
         18 . The system of  claim 17 , wherein the combination is a weighted combination. 
     
     
         19 . The system of  claim 18 , wherein weights used to generate the weighted combination are preset. 
     
     
         20 . The system of  claim 19 , wherein the weights are differentiable and learned as part of the training of the first model. 
     
     
         21 . The system of  claim 20 , wherein the weights are differentiable and learned in stacked re-ranking layers that are trained as part of the training of the first model using activation functions that generate non-linear combinations of the first set of score rankings, the second set of score rankings, and the third set of score rankings. 
     
     
         22 . The system of  claim 15 , further configured to comprise:
 a fourth model configured to generate, based in part on masked representations of the evolutionary conservation summary statistics, a fourth set of pathogenicity scores for the set of variants, wherein the masked representations mask evolutionary conservation summary statistic data about at least one amino acid in the alternate protein sequences, and wherein the fourth set of pathogenicity scores has a fourth set of score rankings;   the rank loss determination logic further configured to determine the rank loss parameter based on a comparison of the first set of score rankings, the second set of score rankings, and the fourth set of score rankings;   the loss function reconfiguration logic further configured to reconfigure the loss function based on the rank loss parameter; and   the training logic further configured to use the reconfigured loss function to train the first model.   
     
     
         23 . The system of  claim 22 , further configured to comprise:
 the rank loss determination logic further configured to determine the rank loss parameter based on a comparison of the first set of score rankings, the second set of score rankings, the third set of score rankings, and the fourth set of score rankings;   the loss function reconfiguration logic further configured to reconfigure the loss function based on the rank loss parameter; and   the training logic further configured to use the reconfigured loss function to train the first model.   
     
     
         24 . The system of  claims 22 , wherein the training logic is further configured to use the reconfigured loss function to train the fourth model. 
     
     
         25 . The system of  claims 22 , further configured to comprise:
 the loss function reconfiguration logic further configured to reconfigure, based on the rank loss parameter, a first loss function for the first model and a fourth loss function for the fourth model; and   the training logic further configured to use the reconfigured first function to train the first model, and to use the reconfigured fourth function to train the fourth model.   
     
     
         26 . The system of  claim 1 , wherein the first model is further configured to generate, based in part on three-dimensional (3D) structural representations of amino acids in the reference target protein sequence, the first set of pathogenicity scores. 
     
     
         27 . The system of  claim 1 , wherein the first model is further configured to generate, based in part on the reference target protein sequence, the first set of pathogenicity scores. 
     
     
         28 . The system of  claims 1 , wherein the first model is further configured to generate, based in part on the alternate protein sequences, the first set of pathogenicity scores. 
     
     
         29 . The system of  claim 22 , wherein the fourth model is further configured to generate, based in part on masked representations of the 3D structural representations of the amino acids in the reference target protein sequence, the fourth set of pathogenicity scores, wherein the masked representations of the 3D structural representations mask 3D structural data about at least one amino acid in the reference target protein sequence. 
     
     
         30 . The system of  claim 22 , wherein the fourth model is further configured to generate, based in part on a masked representation of the reference target protein sequence, the fourth set of pathogenicity scores, wherein the masked representation masks at least one amino acid in the reference target protein sequence. 
     
     
         31 . The system of  claim 22 , wherein the fourth model is further configured to generate, based in part on masked representations of the alternate protein sequences, the fourth set of pathogenicity scores, wherein the masked representations of the alternate protein sequences mask at least one amino acid in the reference target protein sequence. 
     
     
         32 . The system of  claims 1 , wherein the evolutionary conservation summary statistics are determined from evolutionary profiles. 
     
     
         33 . The system of  claim 32 , wherein the evolutionary profiles include position-specific score matrices (PSSMs). 
     
     
         34 . The system of  claim 32 , wherein the evolutionary profiles include position-specific frequency matrices (PSFMs). 
     
     
         35 . The system of  claim 1 , wherein the reference target protein sequence is a sub-sequence in a region in the reference target protein sequence. 
     
     
         36 . The system of  claim 1 , wherein the alternate protein sequences are sub-sequences in regions in the alternate protein sequences.

Join the waitlist — get patent alerts

Track US2023207064A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.