Systems and methods for pruning data by sampling
Abstract
Techniques provided herein allow for management of data. In various embodiments, systems and methods prune and retain data being managed by a data management system, where the managed data can include log data aggregated from one or more servers for analysis purposes. According to some embodiments, pruning can be triggered according to one or more constraints, such as the age of managed data (e.g., retain only 30 days of managed data) or the memory space required to store the managed data (e.g., retain only 100 GB worth of managed data). The constraints that trigger data pruning can be based on a data retention policy. When triggered, pruning can be performed on a fraction of the managed data stored based on the data retention policy (e.g., 3 days of full managed data, 27 days of pruned managed data). The pruning may be performed by sampling, at a desired rate, the managed data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method comprising:
identifying, by a computing system, an initial data subset from a data set for each of a plurality of time periods, from which at least some data elements will be removed by sampling; determining, by the computing system, a sampling rate for data element retention; identifying, by the computing system, a secondary data subset from the initial data subset for each of the plurality of time periods, based on sampling the initial data subset according to the sampling rate, the sampling rate applied to the initial data subset for each of the plurality of time periods; and removing, by the computing system, from the data set one or more data elements of the initial data subset for each of the plurality of time periods while retaining data elements of the secondary data subset for each of the plurality of time periods, wherein the sampling rate is determined such that a representative portion of the data set is retained when the one or more data elements of the initial data subset for each of the plurality of time periods are removed from the data set.
2 . The computer-implemented method of claim 1 , wherein the data set comprises log data.
3 . The computer-implemented method of claim 2 , wherein the log data is associated with operation of a social networking system.
4 . The computer-implemented method of claim 3 , wherein the log data comprises one or more time-stamped data elements regarding user activity occurring on the social networking system.
5 . The computer-implemented method of claim 1 , wherein the initial data subset is identified from the data set in response to detecting that a constraint for storing a data set has been exceeded.
6 . The computer-implemented method of claim 5 , wherein the constraint relates to one or more of: age of data elements in the data set or storage space occupied by data elements in the data set.
7 . The computer-implemented method of claim 5 , wherein the constraint is based on a data retention policy.
8 . The computer-implemented method of claim 1 , wherein the data set comprises data sampled from a larger data set.
9 . The computer-implemented method of claim 1 , wherein the initial data subset for each of the plurality of time periods is identified according to a data retention policy.
10 . The computer-implemented method of claim 9 , wherein the data retention policy prohibits removal of data elements from the data set that have been maintained for less than a threshold period of time.
11 . The computer-implemented method of claim 1 , wherein the sampling rate is defined by a ratio of data elements.
12 . The computer-implemented method of claim 1 , wherein the sampling rate is determined based on a type of data element included in the data set.
13 . The computer-implemented method of claim 12 , wherein the data set comprises event log data and the type of data element is based on an event type.
14 . The computer-implemented method of claim 1 , wherein the data set is a database table.
15 . The computer-implemented method of claim 14 , wherein the sampling rate is determined based on a table type associated with the database table.
16 . The computer-implemented method of claim 1 , further comprising designating data of the secondary data subset as being data retained during a data removal process.
17 . The computer-implemented method of claim 1 , further comprising associating the sampling rate with data of the secondary data subset.
18 . The computer-implemented method of claim 1 , wherein the data set is being stored in an in-memory database.
19 . A computer system comprising:
at least one processor; and a memory storing instructions configured to instruct the at least one processor to perform:
identifying an initial data subset from a data set for each of a plurality of time periods, from which at least some data elements will be removed by sampling;
determining a sampling rate for data element retention;
identifying a secondary data subset from the initial data subset for each of the plurality of time periods, based on sampling the initial data subset according to the sampling rate, the sampling rate applied to the initial data subset for each of the plurality of time periods; and
removing from the data set one or more data elements of the initial data subset for each of the plurality of time periods while retaining data elements of the secondary data subset for each of the plurality of time periods, wherein the sampling rate is determined such that a representative portion of the data set is retained when the one or more data elements of the initial data subset for each of the plurality of time periods are removed from the data set.
20 . A non-transitory computer-storage medium storing computer-executable instructions that, when executed, cause a computer system to perform a computer-implemented method comprising:
identifying an initial data subset from a data set for each of a plurality of time periods, from which at least some data elements will be removed by sampling; determining a sampling rate for data element retention; identifying a secondary data subset from the initial data subset for each of the plurality of time periods, based on sampling the initial data subset according to the sampling rate, the sampling rate applied to the initial data subset for each of the plurality of time periods; and removing from the data set one or more data elements of the initial data subset for each of the plurality of time periods while retaining data elements of the secondary data subset for each of the plurality of time periods, wherein the sampling rate is determined such that a representative portion of the data set is retained when the one or more data elements of the initial data subset for each of the plurality of time periods are removed from the data set.Join the waitlist — get patent alerts
Track US2017147615A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.