Confirming data consistency in a data storage environment
Abstract
A method for confirming replicated data at a data site, including utilizing a hash function, computing a first hash value based on first data at a first data site and utilizing the same hash function, computing a second hash value based on second data at a second data site, wherein the first data had previously been replicated from the first data site to the second data site as the second data. The method also includes comparing the first and second hash values to determine whether the second data is a valid replication of the first data. In additional embodiments, the first data may be modified based on seed data prior to computing the first hash value and the second data may be modified based on the same seed data prior to computing the second hash value. The process can be repeated to increase reliability of the results.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A method for confirming the validity of replicated data at a data storage site, the method comprising:
a) replicating first data from a first computer readable storage medium at a first data storage site as second data to a second computer readable storage medium at a second data storage site; b) utilizing a hash function, computing a first hash value based on the first data stored on the first computer readable storage medium at the first data storage site, the first hash value being smaller in size than the first data; c) utilizing the same hash function, computing a second hash value based on the second data stored on the second computer readable storage medium at the second data storage site, the second hash value being smaller in size than the second data; d) transmitting at least one of the first or second hash values via a computer network for comparing with the other of the first or second hash values, in place of retransmitting the larger sized first or second data via the computer network; and e) comparing the first and second hash values, in lieu of comparing the actual first and second data, to determine whether the second data is a valid replication of the first data, wherein a mismatch between the first and second hash values indicates that at least one of the first or second data storage sites includes invalid data.
22 . The method of claim 21 , wherein the first and second data storage sites are remotely connected by the computer network.
23 . The method of claim 21 , further comprising providing a data structure storing a plurality of hash functions computable by a computer processor, each being available for use by the first and second data storage sites.
24 . The method of claim 23 , further comprising selecting the hash function from the data structure storing a plurality of hash functions for utilization in computing the first and second hash values.
25 . The method of claim 21 , further comprising modifying the first data based on seed data prior to computing the first hash value and modifying the second data based on the seed data prior to computing the second hash value.
26 . The method of claim 25 , further comprising transmitting the seed data via the network from at least one of the first or second data storage sites to the other of the first or second data storage sites for use by both first and second data storage sites.
27 . The method of claim 21 , further comprising:
utilizing a second hash function, computing a third hash value based on the first data stored on the first computer readable storage medium at the first data storage site; utilizing the second hash function, computing a fourth hash value based on the second data stored on the second computer readable storage medium at the second data storage site; transmitting at least one of the third or fourth hash values via a computer network for comparing with the other of the third or fourth hash values; and comparing the third and fourth hash values, in lieu of comparing the actual first and second data, to determine whether the second data is a valid replication of the first data.
28 . The method of claim 25 , further comprising:
modifying the first and second data based on second seed data; utilizing the hash function, computing a third hash value based on the modified first data; utilizing the second hash function, computing a fourth hash value based on the modified second data; and comparing the third and fourth hash values to determine whether the second data remains a valid replication of the first data.
29 . The method of claim 21 , further comprising repeating steps b) through d) a plurality of times, each time utilizing a different hash function than for a previous time.
30 . The method of claim 29 , wherein the steps b) through d) are repeated according to a predetermined periodic cycle.
31 . The method of claim 21 , further comprising repeating steps b) through d) a plurality of times.
32 . The method of claim 31 , wherein the steps b) through d) are repeated according to a predetermined periodic cycle.
33 . An information handling system comprising:
a first data storage site comprising a computer readable storage medium storing first data, and in operable communication with a computer processor computing a first hash value based on the first data, utilizing a hash function; and a second data storage site comprising a computer readable storage medium storing data replicated from the first data storage site, and in operable communication with a computer processor computing a second hash value based on second data, utilizing the same hash function; wherein in lieu of comparison of the first and second data, the computed first and second hash values are compared as an approximation of whether the second data is a valid replication of the first data, wherein a mismatch between the first and second hash values indicates that at least one of the first or second data storage sites includes invalid data.
34 . The information handling system of claim 33 , wherein the first data storage site and the second data storage site are remotely connected via a computer network.
35 . The information handling system of claim 34 , wherein the first data storage site is configured to modify the first data based on seed data prior to computing the first hash value and the second data storage site is configured to modify the second data based on the seed data prior to computing the second hash value.
36 . The information handling system of claim 33 , wherein computer processor in operable communication with the first data storage site and the computer processor in operable communication with the second data storage site are the same computer processor.
37 . The information handling system of claim 36 , wherein the computer processor is remote to at least one of the first and second data storage sites.
38 . A method for confirming the validity of replicated data at a data storage site, the method comprising:
a) replicating first data from a first computer readable storage medium at a first data storage site as second data to a second computer readable storage medium at a second data storage site; b) utilizing a hash function, computing a first hash value based on a selected portion of the first data stored on the first computer readable storage medium at the first data storage site; c) utilizing the same hash function, computing a second hash value based on a selected portion of the second data stored on the second computer readable storage medium at the second data storage site, the selected portion of second data corresponding to the selected portion of first data; d) transmitting at least one of the first or second hash values via a computer network for comparing with the other of the first or second hash values, in place of retransmitting the first or second data via the computer network; e) comparing the first and second hash values, in lieu of comparing the actual first and second data, as an approximation of whether the selected portion of second data is a valid replication of the selected portion of first data; and f) repeating steps b) through e) a plurality of times, each time utilizing a different selected portion of the first data and corresponding selected portion of the second data than for a previous time, wherein during any repetition a mismatch between the first and second hash values indicates that at least one of the first or second data storage sites includes invalid data.
39 . The method of claim 38 , wherein the first and second data storage sites are remotely connected by the network.
40 . The method of claim 39 , wherein the steps b) through e) are repeated according to a predetermined periodic cycle, each subsequent repetition in a contiguous chain of repetitions resulting in a match of the first and second hash values increasing the likelihood that the second data is a valid replication of the first data.Join the waitlist — get patent alerts
Track US2016210307A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.