Systems and methods for probabilistic data classification
Abstract
A system for performing data classification operations. In one embodiment, the system comprises a file system configured to store a plurality of computer files and a scanning agent configured to traverse the file system and compile data regarding the attributes and content of the plurality of computer files. The system also comprises an index configured to store the data regarding attributes and content of the plurality of computer files and a file classifier configured to analyze the data regarding the attributes and content of the plurality of computer files and to classify the plurality of computer files into one or more categories based on the data regarding the attributes and content of the plurality of computer files. Results of the file classification operations can be used to set appropriate security permissions on files which include sensitive information or to control the way that a file is backed up or the schedule according to which it is archived.
Claims
exact text as granted — not AI-modified1 - 24 . (canceled)
25 . A system comprising:
one or more computing devices comprising computer hardware with one or more processors configured to:
access one or more data blocks of one or more files;
update or create, based on the one or more data blocks, index data associated with the one or more files with content of the one or more data blocks and at least one file attribute associated with the one or more files;
based, at least in part, on the updated index data associated with the one or more files, classify the one or more files as a member of a first category;
following an incremental or differential backup of the one or more files, access one or more modified data blocks of the one or more files,
wherein the one or more modified data blocks are data blocks that have been modified since the classification of the one or more files as a member of the first category;
update the index data associated with the one or more files based on the one or more modified data blocks; and
based, at least in part, on some of the content of the one or more files and at least one file attribute of the updated index data associated with the one or more files, classify the one or more files as a member of a second category.
26 . The system of claim 25 , wherein the one or more computing devices are further configured to alter security access restrictions of the one or more files based upon classification of being a member of at least one category.
27 . The system of claim 25 , wherein the one or more computing devices are further configured to alter a data backup schedule or a data migration plan of the one or more files based upon classification of being a member of at least one category.
28 . The system of claim 25 , wherein the one or more data blocks are secondary copies stored on one or more secondary storage devices.
29 . The system of claim 25 , wherein the one or more processors are further configured to:
determine a probability that the one or more files should be classified as a member of the first category; and determine that the probability satisfies a probability threshold for classifying the one or more files as a member of the first category, wherein the probability threshold is specified by a classification rule associated with the first category.
30 . The system of claim 29 , wherein the classification rule was computed using a training data set.
31 . The system of claim 25 , wherein the index data is stored separately from storage devices where the one or more files are stored.
32 . The system of claim 25 , wherein the at least one file attribute comprises information indicating file size, name, path, type, or date of creation or modification of the one or more files.
33 . The system of claim 25 , wherein the index data further comprises data indicating at least one classification category that the one or more files have been identified as being members of.
34 . The system of claim 25 , wherein the index data further comprises, for each of the one or more files, a list of keywords present in the one or more files and a frequency count for each keyword.
35 . A method for classifying files, the method comprising:
with one or more computing devices comprising computer hardware with one or more processors:
accessing one or more data blocks of one or more files;
updating or creating, based on the one or more data blocks, index data associated with the one or more files with content of the one or more data blocks and at least one file attribute associated with the one or more files;
based, at least in part, on the updated index data associated with the one or more files, classifying the one or more files as a member of a first category;
following an incremental or differential backup of the one or more files, accessing one or more modified data blocks of the one or more files,
wherein the one or more modified data blocks are data blocks that have been modified since the classification of the one or more files as a member of the first category;
updating the index data associated with the one or more files based on the one or more modified data blocks; and
based, at least in part, on some of the content of the one or more files and at least one file attribute of the updated index data associated with the one or more files, classifying the one or more files as a member of a second category.
36 . The method of claim 35 , the method further comprising altering security access restrictions of the one or more files based upon classification of being a member of at least one category.
37 . The method of claim 35 , the method further comprising altering a data backup schedule or a data migration plan of the one or more files based upon classification of being a member of at least one category.
38 . The method of claim 35 , wherein the one or more data blocks are secondary copies stored on one or more secondary storage devices.
39 . The method of claim 35 , the method further comprising:
determining a probability that the one or more files should be classified as a member of the first category; and determining that the probability satisfies a probability threshold for classifying the one or more files as a member of the first category, wherein the probability threshold is specified by a classification rule associated with the first category.
40 . The method of claim 39 , wherein the classification rule was computed using a training data set.
41 . The method of claim 35 , wherein the index data is stored separately from storage devices where the one or more files are stored.
42 . The method of claim 35 , wherein the at least one file attribute comprises information indicating file size, name, path, type, or date of creation or modification of the one or more files.
43 . The method of claim 35 , herein the index data further comprises data indicating at least one classification category that the one or more files have been identified as being members of.
44 . The method of claim 35 , wherein the index data further comprises, for each of the one or more files, a list of keywords present in the one or more files and a frequency count for each keyword.Join the waitlist — get patent alerts
Track US2021342368A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.