US2016283505A1PendingUtilityA1

Methods and apparatus for efficient compression and deduplication

Assignee: DELL PRODUCTS LPPriority: Nov 23, 2009Filed: Jun 8, 2016Published: Sep 29, 2016
Est. expiryNov 23, 2029(~3.3 yrs left)· nominal 20-yr term from priority
G06F 16/1744G06F 16/1748G06F 16/174G06F 17/30153G06F 17/30156
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Mechanisms are provided for performing efficient compression and deduplication of data segments. Compression algorithms are learning algorithms that perform better when data segments are large. Deduplication algorithms, however, perform better when data segments are small, as more duplicate small segments are likely to exist. As an optimizer is processing and storing data segments, the optimizer applies the same compression context to compress multiple individual deduplicated data segments as though they are one segment. By compressing deduplicated data segments together within the same context, data reduction can be improved for both deduplication and compression. Mechanisms are applied to compensate for possible performance degradation.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 partitioning a plurality of data files into a plurality of data segments;   compressing a first subset of data segments using a shared compression context, wherein a compressor examines patterns across the first subset of data segments to determine frequently occurring patterns during compression of the first subset of data segments; and   compressing a second subset of data segments using a plurality of non-shared compression contexts.   
     
     
         2 . The method of  claim 1 , wherein the plurality of data segments span a. plurality of files. 
     
     
         3 . The method of  claim 1 , wherein the plurality of files are determined to be container or non-container files based on the file types associated with the files. 
     
     
         4 . The method of  claim 1 , wherein the plurality of data files are partitioned using segment sizes dependent on file type. 
     
     
         5 . The method of  claim 3 , wherein non-container files are evaluated to identify whether the non-container files would benefit from more fine-grained. segmentation, wherein non-container files that would not benefit from more fine-grained segmentation have segment boundaries set to the file boundaries. 
     
     
         6 . The method of  claim 5 , wherein container files are recursively parsed to determine whether components of the container files are container or non-container components. 
     
     
         7  The method of  claim 5 , wherein deduplicating the plurality of files comprises generating a plurality of filemaps corresponding to the plurality of files. 
     
     
         8 . The method of  claim 5 , wherein deduplicating the plurality of files comprises generating a plurality of datastore suitcases. 
     
     
         9 . The method of  claim 8 , wherein the datastore suitcase further comprises a plurality of reference counts corresponding to a plurality of deduplicated data segments. 
     
     
         10 . The method of  claim 9 , wherein the plurality of &duplicated data segments are determined to be infrequently accessed if they have low associated reference counts. 
     
     
         11 . A system, comprising:
 an interface configured to receive a plurality of data files;   a processor configured to partition a plurality of data files into a data segments;   storage configured to maintain the plurality of data segments;   wherein a first subset of data segments from the plurality of data segments is compressed using a shared compression context, wherein a compressor examines patterns across the first subset of data segments to determine frequently occurring patterns during compression of the first subset of data segments, and wherein a second subset of data segments is compressed using a plurality of non-shared compression contexts.   
     
     
         12 . The system of  claim 11 , wherein the plurality of data segments span a plurality of files. 
     
     
         13 . The system of  claim 11 , wherein the plurality of files are determined to be container or non-container files based on the file types associated with the files. 
     
     
         14 . The system of  claim 11 , wherein the plurality of data files are partitioned using segment sizes dependent on file type. 
     
     
         15 . The system of  claim 13 , wherein non-container files are evaluated to identify whether the non-container files would benefit from more fine-grained segmentation, wherein non-container files that would not benefit from more fine-grained segmentation have segment boundaries set to the file boundaries. 
     
     
         16 . The system of  claim 15 , wherein container files are recursively parsed to determine whether components of the container files are container or non-container components. 
     
     
         17 . The system of  claim 15 , wherein deduplicating the plurality of files comprises generating a plurality of filemaps corresponding to the plurality of files. 
     
     
         18 . The system of  claim 15 , wherein deduplicating the plurality of files comprises generating a plurality of datastore suitcases. 
     
     
         19 . The system of  claim 18 , wherein the datastore suitcase further comprises a plurality of reference counts corresponding to a plurality of deduplicated data segments. 
     
     
         20 . A non-transitory computer readable medium storing instructions to cause a processor to execute a method, the method comprising:
 partitioning a plurality of data tiles into a plurality of data segments;   compressing a first subset of data segments using a shared compression context, wherein a compressor examines patterns across the first subset of data segments to determine frequently occurring patterns during compression of the first subset of data segments; and   compressing a second subset of data segments using a plurality of non-shared compression contexts.

Join the waitlist — get patent alerts

Track US2016283505A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.