Clustering analysis for deduplication of training set samples for machine learning based computer threat analysis
Abstract
A method, a system, and a computer program product for performing analysis of data to detect presence of malicious code are disclosed. Reduced dimensionality vectors are generated from a plurality of original dimensionality vectors representing features in a plurality of samples. The reduced dimensionality vectors have a lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors. A first plurality of clusters is determined by applying a first clustering algorithm to the reduced dimensionality vectors. A second plurality of clusters is determined by applying a second clustering algorithm to one or more clusters in the first plurality of clusters using the original dimensionality. An exemplar for a cluster in the second plurality of clusters is added to a training set, which is used to train a machine learning model for identifying a file containing malicious code.
Claims
exact text as granted — not AI-modifiedWhat is claimed:
1 . A computer-implemented method comprising
generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors; first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors; second determining a second plurality of clusters, the second determining comprising second applying a second clustering algorithm to one or more clusters of the first plurality of clusters using the original dimensionality; selecting an exemplar for a cluster of the second plurality of clusters based on a date of creation for the samples corresponding to the second plurality of clusters; and adding the selected exemplar for the cluster of the second plurality of clusters to a training set, the training set being used by at least one computing system to train a machine learning model for identifying a file containing malicious code.
2 . The method according to claim 1 , wherein the first and second clustering algorithms are same.
3 . The method according to claim 1 , where the generating of the reduced dimensionality vectors comprises applying a random projection to the original dimensionality vectors.
4 . The method according to claim 3 , wherein the random projection approximately preserves all pairwise distances between the original dimensionality vectors.
5 . The method according to claim 3 , wherein the random projection has a predetermined size.
6 . The method according to claim 1 , wherein the adding further comprises selecting the exemplar corresponding to at least one of the following: a point in each cluster in the second plurality of clusters, a point approximately close to a center of a cluster in the second plurality of clusters, and a number of points in a cluster in the second plurality of cluster determined based on size of the cluster in the second plurality of clusters.
7 . The method according to claim 6 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between points contained in the cluster of the second plurality of clusters are less than the predetermined radius.
8 . The method according to claim 6 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between some points contained in the cluster of the second plurality of clusters are greater than the predetermined radius.
9 . The method according to claim 6 , wherein the cluster of the second plurality of clusters has a predetermined minimum number of points.
10 . The method according to claim 6 , wherein the exemplar corresponds to a sample outside of a cluster in the second plurality of clusters.
11 . A system comprising:
at least one programmable processor; and memory storing instructions which, when executed by the at least one programmable processor, execute operations comprising:
generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors;
first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors;
second determining a second plurality of clusters, the second determining comprising second applying a second clustering algorithm to one or more clusters of the first plurality of clusters using the original dimensionality;
selecting, based on a predetermined fashion, an exemplar for a cluster of the second plurality of clusters; and
adding the selected exemplar for the cluster of the second plurality of clusters to a training set, the training set being used by at least one computing system to train a machine learning model for identifying a file containing malicious code.
12 . The system according to claim 11 , wherein the first and second clustering algorithms are same.
13 . The system according to claim 11 , where the generating of the reduced dimensionality vectors comprises applying a random projection to the original dimensionality vectors.
14 . The system according to claim 13 , wherein the random projection approximately preserves all pairwise distances between the original dimensionality vectors.
15 . The system according to claim 13 , wherein the random projection has a predetermined size.
16 . The system according to claim 11 , wherein the adding further comprises selecting the exemplar corresponding to at least one of the following: a point in each cluster in the second plurality of clusters, a point approximately close to a center of a cluster in the second plurality of clusters, and a number of points in a cluster in the second plurality of cluster determined based on size of the cluster in the second plurality of clusters.
17 . The system according to claim 16 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between points contained in the cluster of the second plurality of clusters are less than the predetermined radius.
18 . The system according to claim 16 , wherein the cluster of the second plurality of clusters has a predetermined radius, wherein pairwise distances between some points contained in the cluster of the second plurality of clusters are greater than the predetermined radius.
19 . The system according to claim 16 , wherein the cluster of the second plurality of clusters has a predetermined minimum number of points and/or wherein the exemplar corresponds to a sample outside of a cluster in the second plurality of clusters.
20 . A computer-implemented method comprising:
generating reduced dimensionality vectors from a plurality of original dimensionality vectors representing features in a plurality of samples, the reduced dimensionality vectors having lower dimensionality than an original dimensionality of the plurality of original dimensionality vectors; first determining a first plurality of clusters, the first determining comprising first applying a first clustering algorithm to the reduced dimensionality vectors; second determining a second plurality of clusters, the second determining comprising second applying a second clustering algorithm to one or more clusters of the first plurality of clusters using the original dimensionality; selecting, for each cluster of the second plurality of clusters, a number of exemplars proportionally based on a spread of such cluster; and adding the selected exemplars to a training set, the training set being used by at least one computing system to train a machine learning model.Join the waitlist — get patent alerts
Track US2023206131A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.