US2022342852A1PendingUtilityA1

Distributed deduplicated storage system

Assignee: COMMVAULT SYSTEMS INCPriority: Dec 14, 2010Filed: Jul 7, 2022Published: Oct 27, 2022
Est. expiryDec 14, 2030(~4.4 yrs left)· nominal 20-yr term from priority
G06F 11/1453G06F 16/256G06F 3/067G06F 16/137G06F 16/22H04L 67/06G06F 16/174G06F 16/1748H04L 67/02
75
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A distributed, deduplicated storage system according to certain embodiments is arranged in a parallel configuration including multiple deduplication nodes. Deduplicated data is distributed across the deduplication nodes. The deduplication nodes can be networked together and communicate with one another according using a light-weight, customized communication scheme (e.g., a scheme based on FTP or HTTP). In some cases, deduplication management information including deduplication signatures and/or other metadata is stored separately from the deduplicated data in deduplication management nodes, improving performance and scalability.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method, the computer-implemented method comprising:
 in a distributed storage system comprising one or more hardware processors:
 identifying source data from a client device for storage; 
 splitting the source data into at least first and second data blocks; 
 transmitting in parallel the first data block to a first deduplication database media agent and the second data block to a second deduplication database media agent; 
 determining whether the first data block is stored in a secondary storage device by querying the first deduplication database media agent,
 wherein the first deduplication database media agent is associated with first deduplication management information, 
 wherein the first deduplication management information is stored separately from deduplicated data stored in secondary storage devices; and 
 
 determining whether the second data block is stored in a secondary storage device by querying the second deduplication database media agent,
 wherein the second deduplication database media agent is associated with second deduplication management information, 
 wherein the second deduplication management information is stored separately from the deduplicated data stored in the secondary storage devices. 
 
   
     
     
         2 . The computer-implemented method of  claim 1  further comprising receiving, from the client device, data for storage as a secondary copy to a secondary storage devices, wherein the client device is separate from the secondary storage devices. 
     
     
         3 . The computer-implemented method of  claim 2 , wherein the data transmitted is a data block of a file. 
     
     
         4 . The computer-implemented method of  claim 1  further comprising receiving, from the client device, data at a data media agent for storage as a secondary copy, wherein the client device is separate from the media agent. 
     
     
         5 . The computer-implemented method of  claim 4 , wherein the data transmitted to the data media agent comprises two or more files. 
     
     
         6 . The computer-implemented method of  claim 1 , wherein the first deduplication database media agent generates a first hash signature of the first data block, the second deduplication database media agent generates a second hash signature of the second data block. 
     
     
         7 . The computer-implemented method of  claim 1 , wherein the data media agent is a software module executing on the client device. 
     
     
         8 . The computer-implemented method of  claim 1 , wherein the client device generates a first hash signature of the first data block and a second hash signature of the second data block. 
     
     
         9 . The computer-implemented method of  claim 1 , wherein the distributed storage system comprises a plurality of data media agents,
 wherein the plurality of data media agents are in communication with one another and in communication with a plurality of deduplication database media agents.   
     
     
         10 . The computer-implemented method of  claim 9 , wherein a first data media agent of the plurality of data media agents provides a path identifier associated with the second data block to a second data media agent in the plurality of data media agents. 
     
     
         11 . The computer-implemented method of  claim 9 , wherein a first data media agent of the plurality of data media agents provides block offset information associated with the second data block to a second data media agent in the plurality of data media agents. 
     
     
         12 . The computer-implemented method of  claim 8 , wherein the first deduplication database media agent stores a copy of the second hash signature. 
     
     
         13 . The computer-implemented method of  claim 9 , wherein the plurality of data media agents communicate hash signatures and headers and links between one another without having a shared static mount path configuration. 
     
     
         14 . The computer-implemented method of  claim 9 , wherein the plurality of data media agents communicate hash signatures and headers without using network shares. 
     
     
         15 . The computer-implemented method of  claim 1 , wherein the first deduplication database media agent determines whether the first data block is stored in a secondary storage device by querying the first deduplication management information. 
     
     
         16 . The computer-implemented method of  claim 1 , wherein the second deduplication database media agent determines whether the second data block is stored in a secondary storage device by querying the second deduplication management information. 
     
     
         17 . The computer-implemented method of  claim 1 , wherein the client device is not part of the distributed storage system. 
     
     
         18 . A computer-implemented method, the computer-implemented method comprising:
 in a distributed storage system comprising one or more hardware processors:
 at a client computing device:
 identifying source data from the client device for storage in a secondary storage system,
 wherein the client computing device generates the source data stored at the client computing device, 
 
 splitting the source data into at least first and second data blocks, 
 transmitting in parallel the first data block to a first deduplication database media agent and the second data block to a second deduplication database media agent; 
 
 determining whether the first data block is stored in a secondary storage device by querying the first deduplication database media agent,
 wherein the first deduplication database media agent is associated with first deduplication management information, 
 wherein the first deduplication management information is stored separately from deduplicated data stored in secondary storage devices; and 
 
 determining whether the second data block is stored in a secondary storage device by querying the second database media agent,
 wherein the second deduplication database media agent is associated with second deduplication management information, 
 wherein the second deduplication management information is stored separately from the deduplicated data stored in the secondary storage devices.

Join the waitlist — get patent alerts

Track US2022342852A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.