Reducing transfer of redundant data objects
Abstract
A method and system for reducing storage requirements and speeding up storage operations by reducing the storage of redundant data includes receiving a request that identifies one or more data objects to which to apply a storage operation. For each data object, the storage system determines if the data object contains data that matches another data object to which the storage operation was previously applied. If the data objects do not match, then the storage system performs the storage operation in a usual manner. However, if the data objects do match, then the storage system may avoid performing the storage operation.
Claims
exact text as granted — not AI-modifiedI/We claim:
1 . A method performed by a computer system, the method comprising:
receiving or identifying a request to perform a data storage operation on multiple data objects,
wherein the request to perform the data storage operation on the multiple data object is initiated according to a storage policy associated with the multiple data objects,
wherein the storage policy is a data structure comprising of a set of preferences or criteria for performing the data storage operation;
storing, on a random-access storage medium, a first copy of a first data object,
wherein the first copy is associated with a first logical location in the random-access storage medium;
determining that a second data object of the multiple data objects, is duplicative of the first copy; creating a reference to the first copy based in response to determining that the second data object is duplicative of the first copy; transferring the first copy from the random-access storage medium to another storage media; and storing, in a single-instance database, at least one entry associated with the transferred first copy.
2 . The method of claim 1 , further comprising:
receiving or accessing the multiple data objects from multiple, different logical locations within a computer network.
3 . The method of claim 1 ,
wherein storing the at least one entry to the transferred first copy comprises storing a reference count to track a number of references that refer to the transferred first copy.
4 . The method of claim 1 , wherein storing the at least one entry to the transferred first copy comprises a media identifier identifying a storage medium and an offset on which the transferred first copy is stored within the identified storage medium to the copy.
5 . The method of claim 1 , the method further comprising:
maintaining an index on the random-access storage medium, wherein the index includes, for each data object of the multiple data objects:
an identifier for the data object;
information indicating whether the data object is stored as a copy or a reference to a copy; and
an identifier of a source copy when the data object is stored as a reference to the source copy.
6 . The method of claim 1 , wherein at least some of the multiple data objects are of different types or formats.
7 . The method of claim 1 , wherein the storage policy further specifies which storage medium in a secondary storage system, the first copy is to be transferred to.
8 . The method of claim 1 , wherein the another storage media is a type of sequential storage media.
9 . A non-transitory, computer-readable medium containing instructions for controlling a computer system to execute a method of storing a copy on a sequential storage medium, the method comprising:
receiving or identifying a request to perform a data storage operation on multiple data objects,
wherein the request to perform the data storage operation on the multiple data object is initiated according to a storage policy associated with the multiple data objects,
wherein the storage policy is a data structure comprising of a set of preferences or criteria for performing the data storage operation;
storing, on a random-access storage medium, a first copy of a first data object,
wherein the first copy is associated with a first logical location in the random-access storage medium;
determining that a second data object of the multiple data objects, is duplicative of the first copy; creating a reference to the first copy based in response to determining that the second data object is duplicative of the first copy; transferring the first copy from the random-access storage medium to another storage media; and storing, in a single-instance database, at least one entry associated with the transferred first copy.
10 . The non-transitory, computer-readable medium of claim 9 , further comprising:
receiving or accessing the multiple data objects from multiple, different logical locations within a computer network.
11 . The non-transitory, computer-readable medium of claim 9 ,
wherein storing the at least one entry to the transferred first copy comprises storing a reference count to track a number of references that refer to the transferred first copy.
12 . The non-transitory, computer-readable medium of claim 9 , wherein storing the at least one entry to the transferred first copy comprises a media identifier identifying a storage medium and an offset on which the transferred first copy is stored within the identified storage medium to the copy.
13 . The non-transitory, computer-readable medium of claim 9 , the method further comprising:
maintaining an index on the random-access storage medium, wherein the index includes, for each data object of the multiple data objects:
an identifier for the data object;
information indicating whether the data object is stored as a copy or a reference to a copy; and
an identifier of a source copy when the data object is stored as a reference to the source copy.
14 . The non-transitory, computer-readable medium of claim 9 , wherein at least some of the multiple data objects are of different types or formats.
15 . The non-transitory, computer-readable medium of claim 9 , wherein the storage policy further specifies which storage medium in a secondary storage system, the first copy is to be transferred to.
16 . The non-transitory, computer-readable medium of claim 9 , wherein the another storage media is a type of sequential storage media.
17 . An information management system configured to reduce data transfer, the information management system comprising:
one or more computing devices comprising computer hardware configured to:
receive or identify a request to perform a data storage operation on multiple data objects,
wherein the request to perform the data storage operation on the multiple data object is initiated according to a storage policy associated with the multiple data objects,
wherein the storage policy is a data structure comprising of a set of preferences or criteria for performing the data storage operation;
store, on a random-access storage medium, a first copy of a first data object,
wherein the first copy is associated with a first logical location in the random-access storage medium;
determine that a second data object of the multiple data objects, is duplicative of the first copy;
create a reference to the first copy based in response to determining that the second data object is duplicative of the first copy;
transfer the first copy from the random-access storage medium to another storage media; and
store, in a single-instance database, at least one entry associated with the transferred first copy.Join the waitlist — get patent alerts
Track US2021208785A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.