Data integration for distributed and massively parallel processing environments
Abstract
Methods and systems for large scale data integration in distributed or massively parallel environments comprises a development phase wherein the results of a proposed jobflow can be viewed by the user during development, including the results of upstream units where the data sources and data targets can be any of a variety of different platforms, and further comprises the use of remote agents proximate to those data sources and data targets with direct communication between the associated agents under the direction of a topologically central controller to provide, among other things, improved security, reduced latency, reduced bandwidth requirements, and faster throughput.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A technical system comprising:
one or more non-transitory computer readable mediums configured to store executable programmed modules; one or more processors, each of the one or more processors communicatively coupled with at least one of the non-transitory computer readable mediums, the one or more processors configured to:
send first data set extraction instructions to a first agent by a controller;
send first data set transformation instructions to a second agent by the controller;
extract a first data set from a first data system by the first agent in accordance with the first data set extraction instructions;
send the first data set to the second agent via a network by the first agent in accordance with the first data set extraction instructions;
load the first data set to a second data system by the second agent; and
provide the first data set transformation instructions to the second data system by the second agent.Join the waitlist — get patent alerts
Track US2022253454A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.