Application-aware and remote single instance data management
Abstract
A method and system for reducing storage requirements and speeding up storage operations by reducing the storage of redundant data includes receiving a request that identifies one or more files or data objects to which to apply a storage operation. For each file or data object, the storage system determines if the file or data object contains data that matches another file or data object to which the storage operation was previously applied, based on awareness of the application that created the data object. If the data objects do not match, then the storage system performs the storage operation in a usual manner. However, if the data objects do match, then the storage system may avoid performing the storage operation with respect to the particular file or data object.
Claims
exact text as granted — not AI-modifiedI/We claim:
1 . A method for storing single instances of data comprising:
receiving a request to copy data, comprising one of files and data objects from a first storage location associated with a client computing system to a second storage location distinct from the first storage location, during a storage operation, wherein the storage operation comprises at least one of: a backup operation and an archive operation; copying the data based on a storage policy to generate a data copy comprising multiple instances of files and multiple instances of data objects; identifying each file and data object in the data copy; generating a unique identifier for each data object in the data copy and associating the generated unique identifier with each data object; comparing the data objects in the data copy based on the associated unique identifier and determining that at least two data objects comprises the same data based at least in part on the two data objects having the same unique identifier, and copying only a single instance of the two data objects to the second storage location; comparing the generated unique identifier associated with the data objects in the data copy against a previously stored unique identifier associated with a previously stored single instance data stored in the second storage location; determining that the generated unique identifier associated with the data object in the data store is the same as the previously stored unique identifier associated with the previously stored single instanced data stored in the second storage location and not copying the data object associated with the generated unique identifier to the second storage location and associating a pointer with the data object that points to the previously stored single instanced data.
2 . The method of claim 1 , wherein the storage policy specifies preferences and storage operations to be performed on the data.
3 . The method of claim 1 , wherein identifying each data object comprises at least one of: identifying permissions for the file or data object, a property of the file or data object, an access control list for the file or data object, an identifier for the file or data object, a size of the file or data object, a creation data for the file or data object, and an access date for the file or data object.
4 . The method of claim 1 , further comprising parsing through the file in the data copy to identify containers and extracting data objects from the containers.
5 . The method of claim 1 , further comprising identifying metadata associated with a file and parsing the file to identify containers and extracting data objects from the containers based on the identified metadata.
6 . The method of claim 1 , wherein the unique identifier comprises one of: a hash value, a message digest, checksum, digital fingerprint, and digital signature.
7 . The method of claim 1 , further comprises modifying the single instance data stored in the second location, wherein the modifying comprises one of: encrypting the single instance data and compressing the single instance data.
8 . The method of claim 1 , wherein the data comprising one of files and data objects are encrypted at the first storage location before being copied to the data copy.
9 . The method of claim 1 , wherein the data comprising one of files and data objects are compressed at the first storage location before being copied to the data copy.
10 . The method of claim 1 , wherein one of: the client computing system and a single instance database component generates the unique identifier that represents the file or data object.
11 . A method for copying data from a computer system at a first storage location to a second storage location, the method comprising:
receiving a request to copy data, comprising one of files and data objects from a first storage location associated with a client computing system to a second storage location distinct from the first storage location, during a storage operation, wherein the storage operation comprises at least one of: a backup operation and an archive operation; copying the data based on a storage policy to generate a data copy comprising multiple instances of files and multiple instances of data objects; identifying each data object in the data copy of the data wherein the identifying comprises extracting metadata; generating a unique identifier for each data object in the data copy and associating the generated unique identifier with each data object; identifying metadata associated with each of the data objects; comparing the data objects in the data copy based on the data and metadata associated with each of the data objects; determining that at least two data objects comprises the same data; determining that the at least two data objects comprises different metadata; storing only a single instance of the at least two data objects that are determined to comprise the same data; storing multiple instances of the metadata associated with the single instance of the at least two data objects; storing an association between the single instance data of the at least two data objects and the multiple instances of the metadata associated with the single instance to the second storage location.
12 . The method of claim 11 , wherein the extracted metadata comprises at least one of: permissions for the file or data object, a property of the file or data object, an access control list for the file or data object, an identifier for the file or data object, a size of the file or data object, a creation data for the file or data object, and an access date for the file or data object.
13 . The method of claim 11 , wherein the storage policy specifies preferences and storage operations to be performed on the data.
14 . The method of claim 11 , wherein identifying each data object comprises identifying the type of the data file, the size of the file or data object, and the structure of the data file.
15 . The method of claim 11 , further comprising parsing through the file in the data copy to identify containers and extracting data objects from the containers.
16 . The method of claim 11 , wherein the unique identifier comprises one of: a hash value, a message digest, checksum, digital fingerprint, and digital signature.
17 . The method of claim 11 , further comprises modifying the single instance data stored in the second location, wherein the modifying comprises one of: encrypting the single instance data and compressing the single instance data.
18 . The method of claim 11 , wherein the data comprising one of files and data objects are encrypted at the first storage location before being copied to the data copy.
19 . The method of claim 11 , wherein the data comprising one of files and data objects are compressed at the first storage location before being copied to the data copy.
20 . The method of claim 11 , wherein one of: the client computing system and a single instance database component generates the unique identifier that represents the file or data object.Join the waitlist — get patent alerts
Track US2021089502A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.