US2023206131A1PendingUtilityA1

Clustering analysis for deduplication of training set samples for machine learning based computer threat analysis

Assignee: CYLANCE INCPriority: Nov 30, 2016Filed: Mar 6, 2023Published: Jun 29, 2023
Est. expiryNov 30, 2036(~10.3 yrs left)· nominal 20-yr term from priority
G06F 18/28G06F 18/232G06F 21/563G06F 21/564G06F 18/23G06N 20/00G05B 13/0265
68
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method, a system, and a computer program product for performing analysis of data to detect presence of malicious code are disclosed. Reduced dimensionality vectors are generated from a plurality of original dimensionality vectors representing features in a plurality of samples. The reduced dimensionality vectors have a lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors. A first plurality of clusters is determined by applying a first clustering algorithm to the reduced dimensionality vectors. A second plurality of clusters is determined by applying a second clustering algorithm to one or more clusters in the first plurality of clusters using the original dimensionality. An exemplar for a cluster in the second plurality of clusters is added to a training set, which is used to train a machine learning model for identifying a file containing malicious code.

Claims

exact text as granted — not AI-modified
What is claimed: 
     
         1 . A computer-implemented method comprising
 generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors;   first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors;   second determining a second plurality of clusters, the second determining comprising second applying a second clustering algorithm to one or more clusters of the first plurality of clusters using the original dimensionality;   selecting an exemplar for a cluster of the second plurality of clusters based on a date of creation for the samples corresponding to the second plurality of clusters; and   adding the selected exemplar for the cluster of the second plurality of clusters to a training set, the training set being used by at least one computing system to train a machine learning model for identifying a file containing malicious code.   
     
     
         2 . The method according to  claim 1 , wherein the first and second clustering algorithms are same. 
     
     
         3 . The method according to  claim 1 , where the generating of the reduced dimensionality vectors comprises applying a random projection to the original dimensionality vectors. 
     
     
         4 . The method according to  claim 3 , wherein the random projection approximately preserves all pairwise distances between the original dimensionality vectors. 
     
     
         5 . The method according to  claim 3 , wherein the random projection has a predetermined size. 
     
     
         6 . The method according to  claim 1 , wherein the adding further comprises selecting the exemplar corresponding to at least one of the following: a point in each cluster in the second plurality of clusters, a point approximately close to a center of a cluster in the second plurality of clusters, and a number of points in a cluster in the second plurality of cluster determined based on size of the cluster in the second plurality of clusters. 
     
     
         7 . The method according to  claim 6 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between points contained in the cluster of the second plurality of clusters are less than the predetermined radius. 
     
     
         8 . The method according to  claim 6 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between some points contained in the cluster of the second plurality of clusters are greater than the predetermined radius. 
     
     
         9 . The method according to  claim 6 , wherein the cluster of the second plurality of clusters has a predetermined minimum number of points. 
     
     
         10 . The method according to  claim 6 , wherein the exemplar corresponds to a sample outside of a cluster in the second plurality of clusters. 
     
     
         11 . A system comprising:
 at least one programmable processor; and   memory storing instructions which, when executed by the at least one programmable processor, execute operations comprising:
 generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors; 
 first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors; 
 second determining a second plurality of clusters, the second determining comprising second applying a second clustering algorithm to one or more clusters of the first plurality of clusters using the original dimensionality; 
 selecting, based on a predetermined fashion, an exemplar for a cluster of the second plurality of clusters; and 
 adding the selected exemplar for the cluster of the second plurality of clusters to a training set, the training set being used by at least one computing system to train a machine learning model for identifying a file containing malicious code. 
   
     
     
         12 . The system according to  claim 11 , wherein the first and second clustering algorithms are same. 
     
     
         13 . The system according to  claim 11 , where the generating of the reduced dimensionality vectors comprises applying a random projection to the original dimensionality vectors. 
     
     
         14 . The system according to  claim 13 , wherein the random projection approximately preserves all pairwise distances between the original dimensionality vectors. 
     
     
         15 . The system according to  claim 13 , wherein the random projection has a predetermined size. 
     
     
         16 . The system according to  claim 11 , wherein the adding further comprises selecting the exemplar corresponding to at least one of the following: a point in each cluster in the second plurality of clusters, a point approximately close to a center of a cluster in the second plurality of clusters, and a number of points in a cluster in the second plurality of cluster determined based on size of the cluster in the second plurality of clusters. 
     
     
         17 . The system according to  claim 16 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between points contained in the cluster of the second plurality of clusters are less than the predetermined radius. 
     
     
         18 . The system according to  claim 16 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between some points contained in the cluster of the second plurality of clusters are greater than the predetermined radius. 
     
     
         19 . The system according to  claim 16 , wherein the cluster of the second plurality of clusters has a predetermined minimum number of points and/or wherein the exemplar corresponds to a sample outside of a cluster in the second plurality of clusters. 
     
     
         20 . A computer-implemented method comprising:
 generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors;   first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors;   second determining a second plurality of clusters, the second determining comprising second applying a second clustering algorithm to one or more clusters of the first plurality of clusters using the original dimensionality;   selecting, for each cluster of the second plurality of clusters, a number of exemplars proportionally based on a spread of such cluster; and   adding the selected exemplars to a training set, the training set being used by at least one computing system to train a machine learning model.

Join the waitlist — get patent alerts

Track US2023206131A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.