Using utilities injected into cloud-based virtual machines to speed up virtual machine backup operations
Abstract
An executable utility is injected into cloud-based virtual machines (VMs) that are subject to backups by a data storage management system tasked with protecting the cloud-based VMs and their associated data. The utility is injected into a target VM which is “live” and operating. The utility analyzes the VM’s live volume to discover data extents therein, and for each extent computes a respective checksum and determines whether the extent is a “hole.” Afterwards, checksums help identify changed data in successive snapshots of the live volume, so that only changed data will be read and backed up in incremental backups. Time is saved in performing the backup operation first by pre-warming the backup’s source volume in parallel with the utility analyzing the live volume, and second by skipping read operations for extents unchanged since a preceding backup. The resulting incremental backup operation is sped up as compared to prior art approaches.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method comprising:
generating a first data structure that, for each data extent in a first data volume of a virtual machine, comprises: (i) a corresponding indication of whether or not the data extent consists of null data, and (ii) a corresponding checksum of the data extent; generating a first snapshot of the first data volume at a first point in time; causing a second data volume to be initialized with one or more data extents that are based on the first snapshot, wherein the second data volume is distinct from the first data volume; wherein generating the first data structure and causing the second data volume to be initialized are performed by the virtual machine before an incremental backup operation for the second data volume starts; and performing the incremental backup operation, comprising:
based at least in part on information obtained from the first data structure, determining that, among the one or more data extents initialized at the second data volume, a set of data extents are changed compared to data that was backed up in a preceding backup operation of the second data volume, and
generating a secondary copy of the second data volume, based on reading the set of data extents from the second data volume, and further based on skipping reading from the second data volume one or more of: (A) data extents that are unchanged compared to the preceding backup operation, and (B) data extents that consist of null data,
wherein the secondary copy represents an incremental backup copy of the first data volume at the first point in time.
2 . The method of claim 1 , wherein at least some of the generating of the first data structure is concurrent with the second data volume being initialized.
3 . The method of claim 1 , wherein to initialize each data extent at the second data volume, a corresponding data extent of the first snapshot is copied to the second data volume.
4 . The method of claim 1 , wherein a utility injected into the virtual machine performs the generating of the first data structure.
5 . The method of claim 1 , wherein a utility injected into the virtual machine causes the second data volume to be initialized.
6 . The method of claim 1 , wherein a utility injected into the virtual machine prewarms the second data volume, wherein pre-warming the second data volume comprises initializing, at the second data volume, one or more data extents that are based on the first snapshot.
7 . The method of claim 1 , wherein in generating the first data structure, a checksum for null data extents is not calculated.
8 . The method of claim 1 , wherein the virtual machine operates in a cloud computing environment, and wherein the first data volume and the second data volume are configured in the cloud computing environment.
9 . The method of claim 8 , wherein a data storage service of the cloud computing environment initializes the data extents at the second data volume.
10 . The method of claim 1 , wherein determining that the set of data extents are changed is further based on a second data structure, wherein the second data structure comprises, for each data extent in the first data volume, a corresponding checksum that was populated into the second data structure at a second point in time, before the first point in time.
11 . A system comprising:
one or more non-transitory computer-readable media having computer-executable instructions stored thereon; and one or more hardware processors that, having executed the computer-executable instructions, configure the system to: generate a first data structure that, for each data extent in a first data volume of a virtual machine, comprises: (i) a corresponding indication of whether or not the data extent consists of null data, and (ii) a corresponding checksum of the data extent; generate a first snapshot of the first data volume at a first point in time; cause a second data volume to be initialized with one or more data extents that are based on the first snapshot, wherein the second data volume is distinct from the first data volume, and wherein the system is configured to generate at least some of the first data structure concurrently with the second data volume being initialized; and perform an incremental backup operation of the second data volume, comprising:
based at least in part on information obtained from the first data structure, determine that, among the one or more data extents initialized at the second data volume, a set of data extents are changed compared to data that was backed up in a preceding backup operation of the second data volume, and
generate a secondary copy of the second data volume, based on reading the set of data extents from the second data volume, and further based on skipping reading from the second data volume one or more of: (A) data extents that are unchanged compared to the preceding backup operation of the second data volume, and (B) data extents that consist of null data,
wherein the secondary copy represents an incremental backup copy of the first data volume at the first point in time.
12 . The system of claim 11 , wherein generating the first data structure and causing the second data volume to be initialized are performed before starting the incremental backup operation for the second data volume.
13 . The system of claim 11 , wherein the system is further configured to initialize each data extent at the second data volume by copying a corresponding data extent of the first snapshot to the second data volume.
14 . The system of claim 11 , wherein the system is further configured to generate the first data structure by executing a utility injected into the virtual machine.
15 . The system of claim 11 , wherein the system is further configured to execute a utility injected into the virtual machine, wherein the utility causes the second data volume to be initialized.
16 . The system of claim 11 , wherein the system is further configured to execute a utility injected into the virtual machine, wherein pre-warming the second data volume comprises initializing, at the second data volume, one or more data extents that are based on the first snapshot.
17 . The system of claim 11 , wherein the system is further configured not to calculate a checksum for null data extents in generating the first data structure.
18 . The system of claim 11 , wherein the virtual machine is configured to operate in a cloud computing environment, and wherein the first data volume and the second data volume are configured in the cloud computing environment.
19 . The system of claim 18 , wherein a data storage service of the cloud computing environment initializes the data extents at the second data volume.
20 . The system of claim 11 , wherein determining that the set of data extents are changed is further based on a second data structure, wherein the second data structure comprises, for each data extent in the first data volume, a corresponding checksum that was populated into the second data structure at a second point in time, before the first point in time.Join the waitlist — get patent alerts
Track US2023052584A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.