Empirically providing data privacy with reduced noise
Abstract
An empirical approach to providing differential privacy includes applying a common statistical query to a set of databases to produce sample values, both with and without any particular entity's data. The probability density is empirically estimated by sorting the sample values to generate an empirical cumulative distribution function. The cumulative distribution function is differenced across approximately the square root of the number of sample points to get an empirical density function. The statistical query is empirically (ε, δ)-private if the empirical densities with and without any particular individual differ by a factor of no more than exp(ε), with the exception of a set for which the densities exceed that bound by a total of no more than δ.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
receiving a request to run a query on a set of databases, each database in the set of databases including data of a plurality of entities; running a plurality of sample queries based on the query to generate sample query results for the databases in the subset, each of the sample queries having a different entity's data removed relative to other sample queries in the plurality; determining, from the sample query results, that the query fails to meet one or more differential privacy requirements; responsive to determining the query fails to meet the one or more differential privacy requirements, adding noise to the query results until the query results do meet the one or more differential privacy requirements; and providing the query results in response to the request.
2 . The method of claim 1 , wherein determining whether the query meets one or more differential privacy requirements comprises calculating a total amount, δ, by which probability densities with and without each entity's data exceed a bound of differing by no more than a factor of exp(ε), where ε is a parameter indicating a maximum acceptable change in the determined probability for a query if an entity is excluded.
3 . The method of claim 1 , wherein determining that the query fails to meet the one or more differential privacy requirements comprises calculating empirical probability density functions from the sample query results.
4 . The method of claim 3 , wherein calculating the empirical probability density functions comprises:
sorting the sample query results to create empirical cumulative distribution functions of the sample query results with and without each of the plurality of entities; and differencing the empirical cumulative distribution functions using a spacing of a number of data points depending on a size of the data to yield the empirical probability density functions.
5 . The method of claim 4 , wherein the spacing is approximately the square root of a total number of data points in the data set.
6 . The method of claim 3 , wherein calculating the empirical probability density function comprises:
determining an adaptive kernel density estimation, wherein data points are replaced by kernels whose widths are selected to span a number of data points, the number of data points depending on a size of the data set; and
adding the kernels to calculate the density.
7 . The method of claim 6 , wherein the width of the adaptive kernel is such that it spans approximately the square root of a total number of data points in the data set.
8 . The method of claim 6 , wherein the kernel shape is rectangular, triangular, or Gaussian.
9 . The method of claim 3 , wherein there is some information in the databases that may be known to an adversary, and the empirical probability density is calculated conditional on that adversarial information falling into a series of buckets.
10 . The method of claim 1 , wherein the noise is Gaussian noise or Laplacian noise.
11 . The method of claim 1 , wherein the set of databases is an ordered sequence, and, responsive to empirical evidence indicating that the query results in the sequence have statistically significant autocorrelation, the method further comprises:
locally aggregating the sequence of databases into a shorter sequence of larger databases, until the query results no longer have statistically significant autocorrelation; and testing the differential privacy requirements based on the shorter sequence of larger databases.
12 . The method of claim 1 , wherein determining whether the query meets one or more differential privacy requirements comprises calculating a “total delta” Δ, which is defined as the accumulation Δ=1−Π i (1−δ i ) over entities, where δ i is a minimal δ that works in a privacy criterion for entity i.
13 . A non-transitory computer-readable storage medium comprising computer program code that, when executed by a computing system, causes the computing system to perform operations including:
receiving a request to run a query on a set of databases, each database in the set of databases including data of a plurality of entities; running a plurality of sample queries based on the query to generate sample query results for the databases in the subset, each of the sample queries having a different entity's data removed relative to other sample queries in the plurality; determining, from the sample query results, that the query fails to meet one or more differential privacy requirements; responsive to determining the query fails to meet the one or more differential privacy requirements, adding noise to the query results until the query results do meet the one or more differential privacy requirements; and providing the query results in response to the request.
14 . The non-transitory computer-readable storage medium of claim 13 , wherein determining whether the query meets one or more differential privacy requirements comprises calculating a total amount, δ, by which probability densities with and without each entity's data exceed a bound of differing by no more than a factor of exp(c), where c is a parameter indicating a maximum acceptable change in the determined probability for a query if an entity is excluded.
15 . The non-transitory computer-readable storage medium of claim 13 , wherein determining that the query fails to meet the one or more differential privacy requirements comprises calculating empirical probability density functions from the sample query results.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein calculating the empirical probability density functions comprises:
sorting the sample query results to create empirical cumulative distribution functions of the sample query results with and without each of the plurality of entities; and differencing the empirical cumulative distribution functions using a spacing of a number of data points depending on a size of the data to yield the empirical probability density functions.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein calculating the empirical probability density function comprises:
determining an adaptive kernel density estimation, wherein data points are replaced by kernels whose widths are selected to span a number of data points, the number of data points depending on a size of the data set; and
adding the kernels to calculate the density.
18 . The non-transitory computer-readable storage medium of claim 15 , wherein there is some information in the databases that may be known to an adversary, and the empirical probability density is calculated conditional on that adversarial information falling into a series of buckets.
19 . The non-transitory computer-readable storage medium of claim 13 , wherein the set of databases is an ordered sequence, and, responsive to empirical evidence indicating that the query results in the sequence have statistically significant autocorrelation, the operations further include:
locally aggregating the sequence of databases into a shorter sequence of larger databases, until the query results no longer have statistically significant autocorrelation; and testing the differential privacy requirements based on the shorter sequence of larger databases.
20 . The non-transitory computer-readable storage medium of claim 13 , wherein determining whether the query meets one or more differential privacy requirements comprises calculating a “total delta” Δ, which is defined as the accumulation Δ=1−Π i (1−δ i ) over entities, where δ i is a minimal δ that works in a privacy criterion for entity i.Join the waitlist — get patent alerts
Track US2023161782A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.