Synchronized data duplication
Abstract
A system and method for data deduplication is presented. Data received from one or more computing systems is deduplicated, and the results of the deduplication process stored in a reference table. A representative subset of the reference table is shared among a plurality of systems that utilize the data deduplication repository. This representative subset of the reference table can be used by the computing systems to deduplicate data locally before it is sent to the repository for storage. Likewise, it can be used to allow deduplicated data to be returned from the repository to the computing systems. In some cases, the representative subset can be a proper subset wherein a portion of the referenced table is identified shared among the computing systems to reduce bandwidth requirements for reference-table synchronization.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A computer-implemented data deduplication method, the method comprising:
with one or more computing systems of a shared storage system comprising computer hardware and in networked communication with a plurality of computing systems that are physically separate from the shared storage system:
determining whether a first data segment included in data generated by an application executing on a first computing system of the plurality of computing systems is already stored in the shared storage system;
if the first data segment is not already stored in the shared storage system, updating a reference table to include an entry corresponding to the first data segment;
based on an analysis of data traffic received from one or more of the plurality of computing systems, determining at least a second computing system of the plurality of computing systems to which to transmit an updated partial instantiation of the reference table that includes the entry corresponding to the first data segment; and
transmitting the updated partial instantiation of the reference table from the shared storage system to the second computing system of the plurality of computing systems such that, subsequent to said transmitting, a partial instantiation of the reference table local to the second computing system includes the entry corresponding to the first data segment and a partial instantiation of the reference table local to a third computing system of the plurality of computing systems does not include the entry corresponding to the first data segment.
3 . The method of claim 2 , wherein the partial instantiations of the reference table local to the second and third computing systems are proper subsets of the reference table.
4 . The method of claim 2 , further comprising determining additional entries to include in the updated partial instantiation of the reference table based on at least one of data segment utilization rate and data segment size.
5 . The method of claim 2 , further comprising determining additional entries to include in the partial instantiation of the reference table based on a combination of data segment utilization rate and data segment size.
6 . The method of claim 5 , wherein the combination comprises a weighted combination of data segment utilization rate and data segment size.
7 . The method of claim 2 wherein said determining is in response to receiving the first data segment from the first computing system at the shared storage system.
8 . The method of claim 7 further comprising storing the first data segment in the shared storage system.
9 . The method of claim 2 further comprising, subsequent to said transmitting, receiving a signature corresponding to the first data segment from the second computing system at the shared storage system, without receiving the first data segment itself.
10 . A system, comprising:
a shared storage repository comprising computer memory; and a server system including one or more computing devices comprising computer hardware, the server system in networked communication with a plurality of computing systems which are physically separate from the server system, the server system configured to:
determine whether a first data segment included in data generated by an application executing on a first computing system of the plurality of computing systems is already stored in the shared storage repository;
if the first data segment is not already stored in the shared storage repository, update a reference table to include an entry corresponding to the first data segment;
based on an analysis of data traffic received from one or more of the plurality of computing systems, determine at least a second computing system of the plurality of computing systems to which to transmit an updated partial instantiation of the reference table that includes the entry corresponding to the first data segment; and
transmit the updated partial instantiation of the reference table from the server system to the second computing system of the plurality of computing systems such that, subsequent to the transmission of the updated partial instantiation of the reference table, a partial instantiation of the reference table local to the second computing system includes the entry corresponding to the first data segment and a partial instantiation of the reference table local to a third computing system of the plurality of computing systems does not include the entry corresponding to the first data segment.
11 . The system of claim 10 , wherein the partial instantiations of the reference table local to the second and third computing systems are proper subsets of the reference table.
12 . The system of claim 10 , wherein the server system is further configured to determine additional entries to include in the updated partial instantiation of the reference table based on at least one of data segment utilization rate and data segment size.
13 . The system of claim 10 , wherein the server system is further configured to determine additional entries to include in the partial instantiation of the reference table based on a combination of data segment utilization rate and data segment size.
14 . The system of claim 13 , wherein the combination comprises a weighted combination of data segment utilization rate and data segment size.
15 . The system of claim 10 wherein the server system receives the first data segment from the first computing system prior to determining whether the first data segment is included in the data generated by the application.
16 . The system of claim 15 wherein the server system is further configured to store the first data segment in the shared storage repository subsequent to receipt from the first computing system.
17 . The system of claim 10 wherein subsequent to transmission of the partial instantiation of the reference table to the second computing system, the server system receives from the second computing system a signature corresponding to the first data segment, and wherein the server system is configured to update the reference table in response to receipt of the signature.Join the waitlist — get patent alerts
Track US2015154220A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.