Cross-site high-availability distributed cloud storage system to provide multiple virtual channels between storage nodes
Abstract
Systems and methods are described for a cross-site high availability distributed storage system. According to one embodiment, a computer implemented method includes providing a remote direct memory access (RDMA) request for a RDMA stream, and generating, with an interconnect (IC) layer of the first storage node, multiple IC channels and associated IC requests for the RDMA request. The method further includes mapping an IC channel to a group of multiple transport layer sessions to split data traffic of the IC channel into multiple packets for the group of multiple transport layer sessions using an IC transport layer of the first storage node and assigning, with the IC transport layer, a unique transaction identification (ID) to each IC request and assigning a different data offset to each packet of a transport layer session.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer implemented method performed by one or more processing resources of a first storage node of a distributed cloud storage system, the computer implemented method comprising:
providing a remote direct memory access (RDMA) request for a RDMA stream; generating, with an interconnect (IC) layer of the first storage node, multiple IC channels and associating IC requests for the RDMA request; mapping an IC channel to a group of multiple transport layer sessions to split data traffic of the IC channel into multiple packets for the group of multiple transport layer sessions using an IC transport layer of the first storage node; and assigning, with the IC transport layer, a unique transaction identification (ID) to each IC request and assigning a different data offset to each packet of a transport layer session.
2 . The computer implemented method of claim 1 , wherein a transport layer session of the group of multiple transport layer sessions comprises a user datagram protocol (UDP) session.
3 . The computer implemented method of claim 1 , further comprising:
sending the packets of each transport layer session to a device driver of the first storage node; and providing, with the device driver, the packets of each transport layer session into a transmission queue based on the unique transaction ID of a header of each packet.
4 . The computer implemented method of claim 1 , wherein the IC transport layer provides a software implementation of RDMA.
5 . The computer implemented method of claim 1 , further comprising:
determining a number of transport layer sessions for each IC channel based on a maximum transfer size of the transport layer session to fully utilize a network bandwidth between the first storage node and a remote second storage node.
6 . The method of claim 1 , wherein the IC transport layer groups two or more transport layer sessions for each IC channel to increase network throughput between the first storage node and a remote second storage node.
7 . The method of claim 1 , wherein the IC layer, the IC transport layer, and the device driver form a software stack of the first storage node.
8 . A storage node for use in a multi-site distributed storage system, the storage node comprising:
one or more processing resources; and a non-transitory computer-readable medium coupled to the one or more processing resources, having stored therein instructions, which when executed by the one or more processing resources cause the one or more processing resources to: generate multiple remote direct memory access (RDMA) streams with each RDMA stream to handle multiple RDMA requests; map an IC channel for each RDMA stream to multiple transport layer sessions to split data traffic of each IC channel into multiple packets for multiple transport layer sessions of each IC channel of the storage node; and assign a unique transaction identification (ID) to each IC request of an IC channel.
9 . The storage node of claim 8 , wherein the transport layer session comprises a user datagram protocol (UDP) session.
10 . The storage node of claim 8 , wherein the instructions when executed by the one or more processing resources cause the one or more processing resources to:
provide the packets of each transport layer session into a transmission queue based on the unique transaction ID of a header of each packet.
11 . The storage node of claim 8 , wherein the instructions when executed by the one or more processing resources cause the one or more processing resources to:
assign a different data offset to each packet in order to enable reordering of packets at a remote storage node.
12 . The storage node of claim 8 , wherein the instructions when executed by the one or more processing resources cause the one or more processing resources to:
determine a number of transport layer sessions based on a maximum transfer size of the transport layer session.
13 . The storage node of claim 8 , wherein the instructions when executed by the one or more processing resources cause the one or more processing resources to:
group two or more transport layer sessions for each IC channel to increase network throughput between the storage node and a remote storage node.
14 . The storage node of claim 8 , wherein the instructions when executed by the processing resource cause the processing resource to:
perform a data transfer using the multiple transport layer sessions in order to replicate copies of a journal log of data transfer requests from the storage node to a remote storage node.
15 . A non-transitory computer-readable storage medium embodying a set of instructions, which when executed by one or more processing resources of a storage node of a distributed cloud storage system cause the one or more processing resources to:
receive, with receive queues of a device driver of the storage node, data traffic of multiple transport layer sessions for each interconnect (IC) channel and associated remote direct memory access (RDMA) stream from a remote storage node; map each transport layer session to a single receive queue of the device driver of the storage node; order, with an interconnect (IC) transport layer, IC requests of a first IC channel for the multiple transport layer sessions into a first per IC channel queue based on unique transaction identifiers (IDs) of the IC requests; and order, with the interconnect (IC) transport layer, packets for the first per IC channel queue based on data offsets of the packets per IC request.
16 . The non-transitory computer-readable storage medium of claim 15 wherein each transport layer session comprises a user datagram protocol (UDP) session.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein the IC transport layer and the device driver form a software stack of the storage node.
18 . The non-transitory computer-readable storage medium of claim 15 , wherein the set of instructions when executed by a first processing resource of the one or more processing resources cause the first processing resource to:
process packets of the first IC channel from the first per IC channel queue.
19 . The non-transitory computer-readable storage medium of claim 15 , wherein the set of instructions when executed by a second processing resource of the one or more processing resources cause the second processing resource to:
process packets of a second IC channel from a second per IC channel queue.
20 . The non-transitory computer-readable storage medium of claim 15 , wherein the set of instructions when executed by the one or more processing resources of the distributed cloud storage system cause the one or more processing resources to:
copy data from the per IC channel queue into non-volatile (NV) memory of the storage node.Join the waitlist — get patent alerts
Track US2022404980A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.