Fast adaptive similarity detection based on algorithm-specific performance
Abstract
A set of similarity detection algorithms and techniques for determining which signature calculation, sampling, and generation algorithms may be most beneficially applied to application related data are described herein. These algorithms work well with SSD caching software to product high speed, high accuracy, and low false-positive detections. Because the different algorithms may show different performance depending on data sets and different applications, to achieve optimal performance, a calibration process may be applied to each application and associated data set to select the best combination of signature computation and sampling technique. The new algorithms are also very fast with execution times an order of magnitude smaller than existing techniques. While some of the algorithms are presented using examples for the purpose of easy readability, these algorithms are very general and can be easily applied to broad range of cases.
Claims
exact text as granted — not AI-modified1 - 56 . (canceled)
57 . A method comprising:
comparing a count of false positive detections generated by a similarity detection algorithm to a false positive threshold value; increasing the false positive threshold value if the false positive detections are greater than the false positive threshold value; comparing a count of reference and associated blocks identified by the similarity detection algorithm to a similarity detection threshold value if the false positive detections are less than the false positive threshold value; and decreasing the false positive threshold value if the count of reference and associated blocks are less than the similarity detection threshold value.
58 . The method of claim 57 , further comprising determining a false positive detection.
59 . The method of claim 58 , wherein the determining a false positive detection comprises:
determining a presence of similarity between the reference block and the associated block; and determining that a delta between the reference block and the associated block is greater than a compression threshold value.
60 . The method of claim 57 wherein the similarity detection algorithm comprises determining similarity of a data block to at least one reference data block using at least a portion of the plurality of signatures.
61 . The method of claim 60 wherein the determining similarity includes comparing signature occurrence data for the data block to signature occurrence data for the reference block.
62 . The method of claim 60 wherein the determining similarity includes generating a wavelet transform for each data block and comparing a plurality of sub-signatures and wavelet transform coefficients of the wavelet transform for at least one data block and at least one reference block.
63 . The method of claim 60 wherein the determining similarity includes producing a histogram of a portion of the plurality of signatures.
64 . The method of claim 60 wherein the reference block comprises a block of data for which calculated signature popularity exceeds a threshold.
65 . The method of claim 64 wherein the threshold is a reference block popularity threshold.
66 . A system comprising:
a processor configured to compare a count of false positive detections generated by a similarity detection algorithm to a false positive threshold value; increase the false positive threshold value if the false positive detections are greater than the false positive threshold value; compare a count of reference and associated blocks identified by the similarity detection algorithm to a similarity detection threshold value if the false positive detections are less than the false positive threshold value; and increase the false positive threshold value if the count of reference and associated blocks are less than the similarity detection threshold value.
67 . The system of claim 66 wherein the processor is further configured to determine a false positive detection.
68 . The system of claim 67 wherein the false positive detection comprises:
determining a presence of similarity between the reference block and the associated block; and
determining that a delta between the reference block and the associated block is greater than a compression threshold value.
69 . The system of claim 66 wherein the processor is further configured to determine similarity of a data block to at least one reference data block using at least a portion of the plurality of signatures.
70 . The system of claim 69 wherein the determining similarity includes comparing signature occurrence data for the data block to signature occurrence data for the reference block.
71 . The system of claim 69 wherein the determining similarity includes generating a wavelet transform for each data block and comparing a plurality of sub-signatures and wavelet transform coefficients of the wavelet transform for at least one data block and at least one reference block.
72 . The system of claim 69 wherein the determining similarity includes producing a histogram of a portion of the plurality of signatures.
73 . The system of claim 69 wherein the reference block comprises a block of data for which calculated signature popularity exceeds a threshold.
74 . The system of claim 73 wherein the threshold is a reference block popularity threshold.
75 - 107 . (canceled)Join the waitlist — get patent alerts
Track US2017052895A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.