US2016154816A1PendingUtilityA1

Scalable mechanism for detection of commonality in a deduplicated data set

Assignee: DELL PRODUCTS LPPriority: Oct 6, 2009Filed: Feb 5, 2016Published: Jun 2, 2016
Est. expiryOct 6, 2029(~3.2 yrs left)· nominal 20-yr term from priority
Inventors:Vinod Jayaraman
G06F 11/1453G06F 16/1748G06F 16/2365G06F 17/30156G06F 17/30371
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Mechanisms are provided for efficiently determining commonality in a deduplicated data set in a scalable manner regardless of the number of deduplicated files or the number of stored segments. Information is generated and maintained during deduplication to allow scalable and efficient determination of data segments shared in a particular file, other files sharing data segments included in a particular file, the number of files sharing a data segment, etc. Data need not be expanded or uncompressed. Deduplication processing can be validated and verified during commonality detection.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 creating a datastore suitcase when a processor processes a file for deduplication, the datastore suitcase including a data table and a datastore, the data table including data offsets and data reference counts, the datastore including data segments for the file,   wherein the data table is arranged to allow the processor to perform parallel reads of large amounts of data in the datastore.   
     
     
         2 . The method of  claim 1 , further comprising generating a filemap, the filemap indicating where a last data segment was referenced. 
     
     
         3 . The method of  claim 1 , wherein files sharing a particular data segment can be determined by accessing a last file entry for the particular data segment and traversing referencing filemaps until NULL entries are identified. 
     
     
         4 . The method of  claim 1 , wherein the datastore includes a last file entry. 
     
     
         5 . The method of  claim 1 , wherein all files having data segments in common with a particular file are identified by analyzing a filemap for the particular file and identifying last file entries in the datastore suitcase for the data segments in the particular file. 
     
     
         6 . The method of  claim 1 , wherein referencing filemaps are traversed until NULL entries are identified. 
     
     
         7 . The method of  claim 1 , wherein the datastore suitcase includes indices for the data offsets and the data segments. 
     
     
         8 . A system, comprising:
 a processor configured to create datastore suitcase the processor processes a file for deduplication, the datastore suitcase including a data table and a datastore, the data table including data offsets and data reference counts, the datastore including data segments for the file,   wherein the data table is arranged to allow the processor to perform parallel reads of large amounts of data in the datastore.   
     
     
         9 . The system of  claim 8 , wherein the processor is further configured to generate a filemap, the filemap indicating where a last data segment was referenced. 
     
     
         10 . The system of  claim 8 , wherein files sharing a particular data segment can be determined by accessing the last file entry for the particular data segment and traversing referencing filemaps until NULL entries are identified. 
     
     
         11 . The system of  claim 8 , wherein the datastore includes a last file entry. 
     
     
         12 . The system of  claim 8 , wherein all files having data segments in common with a particular file are identified by analyzing a filemap for the particular file and identifying last file entries in the datastore suitcase for the data segments in the particular file. 
     
     
         13 . The system of  claim 8 , wherein referencing filemaps are traversed until NULL entries are identified. 
     
     
         14 . The system of  claim 8 , wherein the datastore suitcase includes indices for the data offsets and the data segments. 
     
     
         15 . A non-transitory computer readable medium having computer code embodied therein, the computer readable medium, the computer code comprising instructions for:
 creating a datastore suitcase when a processor processes a file for deduplication, the datastore suitcase including a data table and a datastore, the data table including data offsets and data reference counts, the datastore including data segments for the file,   wherein the data table is arranged to allow the processor to perform parallel reads of large amounts of data in the datastore.   
     
     
         16 . The non-transitory computer readable medium of  claim 15 , wherein the computer code further includes instructions for generating a filemap, the filemap indicating where a last data segment was referenced. 
     
     
         17 . The non-transitory computer readable medium of  claim 15 , wherein files sharing a particular data segment can be determined by accessing the last file entry for the particular data segment and traversing referencing filemaps until NULL entries are identified. 
     
     
         18 . The non-transitory computer readable medium of  claim 15 , wherein the datastore includes a last file entry. 
     
     
         19 . The non-transitory computer readable medium of  claim 15 , wherein all files having data segments in common with a particular file are identified by analyzing a filemap for the particular file and identifying last file entries in the datastore suitcase for the data segments in the particular file. 
     
     
         20 . The non-transitory computer readable medium of  claim 15 , wherein referencing filemaps are traversed until NULL entries are identified.

Join the waitlist — get patent alerts

Track US2016154816A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.