US2025328756A1PendingUtilityA1

Privacy-conscious and robust detection of anomalies in collaborative and distributed learning

Assignee: CYPRESS SEMICONDUCTOR CORPPriority: Apr 19, 2024Filed: Apr 19, 2024Published: Oct 23, 2025
Est. expiryApr 19, 2044(~17.7 yrs left)· nominal 20-yr term from priority
G06N 7/01G06N 3/094G06N 3/088G06N 3/0455G06N 3/045G06N 3/047G06N 3/08
62
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes training, using data collected from sensors, a probabilistic neural network (NN) model including a set of model weights. The probabilistic NN model is trained to filter out data samples causing a threshold level of model uncertainty. The method includes training, at each cycle of training the probabilistic NN model and based on the set of model weights, a common estimator to generate gradient updates to the set of model weights that are to predict whether model updates from the sensors are anomalous. The method includes assigning, to each sensor, a trust coefficient value that estimates a level of trustworthiness of the model updates. The method includes transmitting the set of model weights to a subset of the sensors for which the trust coefficient value satisfies a threshold value.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 training, by a central computing device, using data collected from a plurality of sensors, a probabilistic neural network (NN) model comprising a set of model weights, wherein the probabilistic NN model is trained to filter out data samples causing a threshold level of model uncertainty;   training, at each cycle of training the probabilistic NN model and based on the set of model weights, a common estimator to generate gradient updates to the set of model weights that are to predict whether model updates from the plurality of sensors are anomalous;   assigning, to each sensor of the plurality of sensors, a trust coefficient value that estimates a level of trustworthiness of the model updates; and   transmitting the set of model weights, by the central computing device, to a subset of sensors of the plurality of sensors for which the trust coefficient value satisfies a threshold value.   
     
     
         2 . The method of  claim 1 , further comprising:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model after training a local instance of the probabilistic NN model;   evaluating, with the common estimator, the updated model weights to classify one or more of the updated model weights as anomalous; and   excluding, in retraining the common estimator, one or more of the updated model weights determined to be anomalous.   
     
     
         3 . The method of  claim 2 , further comprising retraining the common estimator using non-anomalous updated model weights from the central computing device and the subset of sensors. 
     
     
         4 . The method of  claim 1 , further comprising:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model after having trained a local instance of the probabilistic NN model;   evaluating, with the common estimator, the updated model weights to classify one or more of the updated model weights as anomalous; and   updating the trust coefficient value associated with the sensor based on whether each respective updated model weight is classified as anomalous.   
     
     
         5 . The method of  claim 4 , wherein updating the trust coefficient value comprises:
 decreasing the trust coefficient value in response to detecting an updated model weight, of the updated model weights, is anomalous; and   increasing the trust coefficient value in response to detecting an updated model weight, of the updated model weights, is non-anomalous.   
     
     
         6 . The method of  claim 4 , further comprising:
 determining a weighted aggregation of the updated model weights using, for each sensor of the subset of sensors, the updated trust coefficient value and an updated quantity of the data samples for each respective sensor;   modifying the updated model weights received from each respective sensor based on the weighted aggregation;   training the probabilistic NN model using the modified updated model weights to generate a set of updated model weights; and   transmitting the set of updated model weights to the subset of sensors for which the trust coefficient value satisfies the threshold value.   
     
     
         7 . The method of  claim 1 , wherein the common estimator is a variational autoencoder (VAE), the method further comprising:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model;   compressing, using an encoder, the updated model weights to a latent space of a lower dimension compared to that of the updated model weights;   remapping, using a decoder, the updated model weights from the latent space to generate reconstructed updated model weights;   determining reconstruction errors between the updated model weights and the reconstructed updated model weights; and   classifying, as anomalous, one or more of the updated model weights having a corresponding reconstruction error that satisfies a second threshold value.   
     
     
         8 . The method of  claim 1 , wherein the common estimator is a variational autoencoder (VAE), the method further comprising:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model;   compressing, using an encoder, the updated model weights to a latent space of a lower dimension compared to that of the updated model weights;   calculating, from the latent space, a mean of cluster centrals of the latent space using an existing dataset;   calculating a distance value between encoded samples, of the updated model weights, and the mean of cluster centrals; and   classifying as anomalous one or more of the updated model weights having a distance value that exceeds a threshold distance value.   
     
     
         9 . A method comprising:
 receiving, by a sensor of a plurality of sensors, from a central computing device, a local probabilistic neural network (NN) model having an initial set of model weights;   training the local probabilistic NN model, comprising:
 determining a subset of useable data samples by identifying those of a plurality of data samples having a model uncertainty below a threshold value; and 
 training the local probabilistic NN model with the useable data samples to generate updated model weights; and 
   transferring, by the sensor, the updated model weights to the central computing device for use in training a global probabilistic NN model.   
     
     
         10 . The method of  claim 9 , further comprising:
 receiving, from the central computing device, further updated weights based on further training of the global probabilistic NN model; and   further training the local probabilistic NN model using the further updated weights.   
     
     
         11 . The method of  claim 9 , wherein the local probabilistic NN model comprises an ensemble of classifiers or a set of Monte Carlo dropout samples, the method further comprising, during inference using the local probabilistic NN model:
 combining probabilities predicted by each individual classifier in the ensemble or the set of Monte Carlo dropout samples generated;   executing the local probabilistic NN model a plurality of times with dropout enabled, each time obtaining predictions using a different dropout mask; and   combining predictions for each class across the plurality of data samples to obtain the model uncertainty.   
     
     
         12 . The method of  claim 9 , wherein the local probabilistic NN model comprises an evidential deep learning (EDL) model, the method further comprising, during inference using the local probabilistic NN model:
 determining estimates of the model uncertainty for each data sample; and   excluding, from training the local probabilistic NN model, data samples for which the model uncertainty at least satisfies the threshold value.   
     
     
         13 . A non-transitory computer-readable storage medium storing instructions, which when executed, cause a processing device of a central computing device to perform operations comprising:
 training, using data collected from a plurality of sensors, a probabilistic neural network (NN) model comprising a set of model weights, wherein the probabilistic NN model is trained to filter out data samples causing a threshold level of model uncertainty;   training, at each cycle of training the probabilistic NN model and based on the set of model weights, a common estimator comprising gradient updates to the set of model weights that are to predict whether model updates from the plurality of sensors are anomalous;   assigning, to each sensor of the plurality of sensors, a trust coefficient value that estimates a level of trustworthiness of the model updates; and   causing the set of model weights to be transmitted to a subset of sensors of the plurality of sensors for which the trust coefficient value satisfies a threshold value.   
     
     
         14 . The non-transitory computer-readable storage medium of  claim 13 , wherein the operations further comprise:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model after training a local instance of the probabilistic NN model;   evaluating, with the common estimator, the updated model weights to classify one or more of the updated model weights as anomalous; and   excluding, in retraining the common estimator, one or more of the updated model weights determined to be anomalous.   
     
     
         15 . The non-transitory computer-readable storage medium of  claim 14 , wherein the operations further comprise retraining the common estimator using non-anomalous updated model weights from the central computing device and the subset of sensors. 
     
     
         16 . The non-transitory computer-readable storage medium of  claim 13 , wherein the operations further comprise:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model after having trained a local instance of the probabilistic NN model;   evaluating, with the common estimator, the updated model weights to classify one or more of the updated model weights as anomalous; and   updating the trust coefficient value associated with the sensor based on whether each respective updated model weight is classified as anomalous.   
     
     
         17 . The non-transitory computer-readable storage medium of  claim 16 , wherein updating the trust coefficient value comprises:
 decreasing the trust coefficient value in response to detecting an updated model weight, of the updated model weights, is anomalous; and   increasing the trust coefficient value in response to detecting an updated model weight, of the updated model weights, is non-anomalous.   
     
     
         18 . The non-transitory computer-readable storage medium of  claim 16 , wherein the operations further comprise:
 determining a weighted aggregation of the updated model weights using, for each sensor of the subset of sensors, the updated trust coefficient value and an updated quantity of the data samples for each respective sensor;   modifying the updated model weights received from each respective sensor based on the weighted aggregation;   training the probabilistic NN model using the modified updated model weights to generate a set of updated model weights; and   transmitting the set of updated model weights to the subset of sensors for which the trust coefficient value satisfies the threshold value.   
     
     
         19 . The non-transitory computer-readable storage medium of  claim 13 , wherein the common estimator is a variational autoencoder (VAE), the operations further comprising:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model;   compressing, using an encoder, the updated model weights to a latent space of a lower dimension compared to that of the updated model weights;   remapping, using a decoder, the updated model weights from the latent space to generate reconstructed updated model weights;   determining reconstruction errors between the updated model weights and the reconstructed updated model weights; and   classifying, as anomalous, one or more of the updated model weights having a corresponding reconstruction error that satisfies a second threshold value.   
     
     
         20 . The non-transitory computer-readable storage medium of  claim 13 , wherein the common estimator is a variational autoencoder (VAE), the operations further comprising:
 receiving, from a sensor of the subset of sensors, updated model weights for the probabilistic NN model;   compressing, using an encoder, the updated model weights to a latent space of a lower dimension compared to that of the updated model weights;   calculating, from the latent space, a mean of cluster centrals of the latent space using an existing dataset;   calculating a distance value between encoded samples, of the updated model weights, and the mean of cluster centrals; and   classifying as anomalous one or more of the updated model weights having a distance value that exceeds a threshold distance value.

Join the waitlist — get patent alerts

Track US2025328756A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.