Adaptive compression and transmission for big data migration
Abstract
A method for optimizing migration efficiency of a data file over network is provided. Specifically, a total time of compression time of the data file, transfer time of the data file over the network, and decompression time of the data file, is minimized by adaptively selecting compression methods to compress each data block of the data file. For selecting a compression method for a data block, information entropy of the data block is analyzed, and a real status of computing and system resources is considered. Further, trade-off among the resource usage, compassion speed and compression ratio is made to calculate an optimized transmission solution over the network for each data block of the data file.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
identifying an information entropy of a first data block; receiving a resource status of a system, the system including a first computer, a second computer, and a communication channel between the first computer and the second computer; determining a first compression method to compress the first data block based at least in part on the information entropy and the resource status; and transferring a compressed data block over the communication channel from the first computer to the second computer; wherein: the compressed data block is the first data block compressed according to the first compression method.
2 . The method of claim 1 , further comprising:
generating the compressed data block according to the first compression method.
3 . The method of claim 1 , further comprising:
responsive to receipt of the compressed data block by the second computer, causing the compressed data block to be decompressed.
4 . The method of claim 1 , wherein the resource status indicates a current status of the system.
5 . The method of claim 1 , wherein the step of determining a first compression method is further based on a block size of the first data block and a predicted compression time (PCT) of the first data block.
6 . The method of claim 1 , further comprising:
identifying, for the first data block, a real compression time (RCT).
7 . The method of claim 6 , further comprising:
determining a suitability level of the first compression method for a current data block based at least in part on the difference between the RCT and a predicted compression time (PCT) of the first data block.
8 . The method of claim 1 , wherein the information entropy is Shannon entropy.
9 . A computer program product comprising a computer readable storage medium having a set of instructions stored therein that, when executed by a processor, causes the processor to compress adaptively and transmit big data by:
identifying an information entropy of a first data block; receiving a resource status of a system, the system including a first computer, a second computer, and a communication channel between the first computer and the second computer; determining a first compression method to compress the first data block based at least in part on the information entropy and the resource status; and transferring a compressed data block over the communication channel from the first computer to the second computer; wherein: the compressed data block is the first data block compressed according to the first compression method.
10 . The computer program product of claim 9 , further causing the processor to compress adaptively and transmit big data by:
generating the compressed data block according to the first compression method.
11 . The computer program product of claim 9 , further causing the processor to compress adaptively and transmit big data by:
responsive to receipt of the compressed data block by the second computer, causing the compressed data block to be decompressed.
12 . The computer program product of claim 9 , wherein the resource status indicates a current status of the system.
13 . The computer program product of claim 9 , wherein determining a first compression method is further based on a block size of the first data block and a predicted compression time (PCT) of the first data block.
14 . The computer program product of claim 9 , further causing the processor to compress adaptively and transmit big data by:
identifying, for the first data block, a real compression time (RCT); and determining a suitability level of the first compression method for a current data block based at least in part on the difference between the RCT and a predicted compression time (PCT) of the first data block.
15 . The computer program product of claim 9 , wherein the information entropy is Shannon entropy.
16 . A computer system comprising:
a processor set; and a computer readable storage medium; wherein: the processor set is structured, located, connected, and/or programmed to run program instructions stored on the computer readable storage medium; and the program instructions, when executed by the processor set, cause the processor set to compress adaptively and transmit big data by:
identifying an information entropy of a first data block;
receiving a resource status of a system, the system including a first computer, a second computer, and a communication channel between the first computer and the second computer;
determining a first compression method to compress the first data block based at least in part on the information entropy and the resource status; and
transferring a compressed data block over the communication channel from the first computer to the second computer;
wherein:
the compressed data block is the first data block compressed according to the first compression method.
17 . The computer system of claim 16 , further causing the processor to compress adaptively and transmit big data by:
generating the compressed data block according to the first compression method.
18 . The computer system of claim 16 , further causing the processor to compress adaptively and transmit big data by:
responsive to receipt of the compressed data block by the second computer, causing the compressed data block to be decompressed.
19 . The computer system of claim 16 , wherein determining a first compression method is further based on a block size of the first data block and a predicted compression time (PCT) of the first data block.
20 . The computer system of claim 16 , further causing the processor to compress adaptively and transmit big data by:
identifying, for the first data block, a real compression time (RCT); and determining a suitability level of the first compression method for a current data block based at least in part on the difference between the RCT and a predicted compression time (PCT) of the first data block.Join the waitlist — get patent alerts
Track US2017214773A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.