US2020241781A1PendingUtilityA1
Method and system for inline deduplication using erasure coding
Est. expiryJan 29, 2039(~12.5 yrs left)· nominal 20-yr term from priority
H03M 13/373G06F 11/1076G06F 3/0641G06F 3/0608G06F 3/067G06F 3/0673G06F 3/065G06F 3/0619
50
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method for storing data includes obtaining data, applying an erasure coding procedure to the data to obtain a plurality of data chunks and a parity chunk, deduplicating the plurality of data chunks to obtain a plurality of deduplicated data chunks, storing, across a plurality of nodes, the plurality of deduplicated data chunks and the parity chunk, and tracking location information for each of the plurality of deduplicated data chunks and the parity chunk.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for storing data, the method comprising:
obtaining data; applying an erasure coding procedure to the data to obtain a plurality of data chunks and a parity chunk; deduplicating the plurality of data chunks to obtain a plurality of deduplicated data chunks; storing, across a plurality of nodes, the plurality of deduplicated data chunks and the parity chunk; and tracking location information for each of the plurality of deduplicated data chunks and the parity chunk.
2 . The method of claim 1 , further comprising:
obtaining second data; applying the erasure coding procedure to the second data to obtain a second plurality of data chunks and a second parity chunk; deduplicating the second plurality of data chunks to obtain a second plurality of deduplicated data chunks; storing, across the plurality of nodes and using the location information for at least one of the plurality of deduplicated data chunks, the second plurality of deduplicated data chunks and the second parity chunk.
3 . The method of claim 2 ,
wherein a first deduplicated data chunk of the first plurality of deduplicated data chunks is stored in a node of the plurality of nodes, wherein a second deduplicated data chunk of the second plurality of deduplicated data chunks is a modified version of the first deduplicated data chunk, and wherein storing, across the plurality of nodes and using the location information for at least one of the plurality of deduplicated data chunks, the second plurality of deduplicated data chunks and the second parity chunk comprises storing the second deduplicated data chunk on the node of the plurality of nodes.
4 . The method of claim 3 ,
wherein the plurality of data chunks and the parity chunk are associated with a first stripe; wherein the second plurality of data chunks and the second parity chunk is associated with a second stripe, wherein the second stripe is a modified version of the first stripe, wherein storing, across the plurality of nodes and using the location information for at least one of the plurality of deduplicated data chunks, the second plurality of deduplicated data chunks and the second parity chunk further comprises storing the parity chunk and the second parity chunk on a second node of the plurality of nodes.
5 . The method of claim 1 , wherein the erasure coding procedure is applied by a deduplicator executing on a node in an accelerator pool, wherein the plurality of nodes is located is a non-accelerator pool, and wherein a data cluster comprises the accelerator pool and the non-accelerator pool.
6 . The method of claim 1 , wherein applying the erasure coding procedure comprises:
dividing the data into data chunks; selecting, from the data chunks, the plurality of data chunks; and generating the parity chunk using the plurality of data chunks.
7 . The method of claim 1 , wherein the parity chunk comprises a P parity value.
8 . The method of claim 1 , wherein each of the plurality of nodes is in a separate fault domain.
9 . The method of claim 1 , wherein deduplicating the plurality of data chunks to obtain the plurality of deduplicated data chunks is performed after a parity value for the plurality of data chunks is calculated.
10 . A non-transitory computer readable medium comprising computer readable program code, which when executed by a computer processor enables the computer processor to perform a method for storing data, the method comprising:
obtaining data; applying an erasure coding procedure to the data to obtain a plurality of data chunks and a parity chunk; deduplicating the plurality of data chunks to obtain a plurality of deduplicated data chunks; storing, across a plurality of nodes, the plurality of deduplicated data chunks and the parity chunk; and tracking location information for each of the plurality of deduplicated data chunks and the parity chunk.
11 . The non-transitory computer readable medium of claim 10 , the method further comprising:
obtaining second data; applying the erasure coding procedure to the data to obtain a second plurality of data chunks and a second parity chunk; deduplicating the second plurality of data chunks to obtain a second plurality of deduplicated data chunks; storing, across the plurality of nodes and using the location information for at least one of the plurality of deduplicated data chunks, the second plurality of deduplicated data chunks and the second parity chunk.
12 . The non-transitory computer readable medium of claim 11 ,
wherein a first deduplicated data chunk of the first plurality of deduplicated data chunks is stored in a node of the plurality of nodes, wherein a second deduplicated data chunk of the second plurality of deduplicated data chunks is a modified version of the first deduplicated data chunk, and wherein storing, across the plurality of nodes and using the location information for at least one of the plurality of deduplicated data chunks, the second plurality of deduplicated data chunks and the second parity chunk comprises storing the second deduplicated data chunk on the node of the plurality of nodes.
13 . The non-transitory computer readable medium of claim 12 ,
wherein the plurality of data chunks and the parity chunk are associated with a first stripe; wherein the second plurality of data chunks and the second parity chunk is associated with a second stripe, wherein the second stripe is a modified version of the first stripe, wherein storing, across the plurality of nodes and using the location information for at least one of the plurality of deduplicated data chunks, the second plurality of deduplicated data chunks and the second parity chunk further comprises storing the parity chunk and the second parity chunk on a second node of the plurality of nodes.
14 . The non-transitory computer readable medium of claim 10 , wherein the erasure coding procedure is applied by a deduplicator executing on a node in an accelerator pool, wherein the plurality of nodes is located is a non-accelerator pool, and wherein a data cluster comprises the accelerator pool and the non-accelerator pool.
15 . The non-transitory computer readable medium of claim 10 , wherein applying the erasure coding procedure comprises:
dividing the data into data chunks; selecting, from the data chunks, the plurality of data chunks; and generating the parity chunk using the plurality of data chunks.
16 . The non-transitory computer readable medium of claim 10 , wherein the parity chunk comprises a P parity value.
17 . The non-transitory computer readable medium of claim 10 , wherein each of the plurality of nodes is in a separate fault domain.
18 . The non-transitory computer readable medium of claim 10 , wherein deduplicating the plurality of data chunks to obtain the plurality of deduplicated data chunks is performed after a parity value for the plurality of data chunks is calculated.
19 . A data cluster, comprising:
a plurality of data nodes comprising an accelerator pool and a non-accelerator pool, wherein the accelerator pool comprises a data node, and the non-accelerator pool comprises a plurality of data nodes; wherein the data node of the plurality node is programmed to:
obtain data;
apply an erasure coding procedure to the data to obtain a plurality of data chunks and a parity chunk;
deduplicate the plurality of data chunks to obtain a plurality of deduplicated data chunks;
store, across a plurality of nodes, the plurality of deduplicated data chunks and the parity chunk; and
track location information for each of the plurality of deduplicated data chunks and the parity chunk.
20 . The data cluster of claim 19 , wherein the node is further programmed to:
obtain second data; apply the erasure coding procedure to the second data to obtain a second plurality of data chunks and a second parity chunk; deduplicate the second plurality of data chunks to obtain a second plurality of deduplicated data chunks; and storing, across the plurality of nodes and using the location information for at least one of the plurality of deduplicated data chunks, the second plurality of deduplicated data chunks and the second parity chunk.Join the waitlist — get patent alerts
Track US2020241781A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.