Stream processing in search data pipelines
Abstract
Architecture that decomposes of one or more monolithic data concepts into atomic concepts and related atomic concept dependencies, and provides streaming data processing that processes individual or separate (atomic) data concepts and defined atomic dependencies. The architecture can comprise data-driven data processing that enables the plug-in of new data concepts with minimal effort. Efficient processing of the data concepts is enabled by streaming only required data concepts and corresponding dependencies and enablement of the seamless configuration of data processing between stream processing systems and batch processing systems as a result of data concept decomposition. Incremental and non-incremental metric processing enables realtime access and monitoring of operational parameters and queries.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A method comprising:
identifying individual concepts of data from a monolithic unit of data; determining one or more dependencies associated with each of the individual concepts of data, wherein each individual concept is associated with at least one corresponding dependency and wherein each dependency interrelates each individual concept of data with at least one other individual concept of data; generating a dependency graph based upon the one or more dependencies associated with each individual concept of data; deriving a computation path based on the dependency graph; and processing one of the individual concepts of data and each dependency associated with the individual concept of data based upon the computation path in response to a request for data relating to the monolithic unit of data.
22 . The method of claim 21 , wherein each individual concept of data comprises an input and an output, and wherein the one or more dependencies are associated with each individual concept of data based upon the input and the output.
23 . The method of claim 22 , wherein the output of an individual concept of data is dependent upon the input, the input comprises at least one other individual concept of data, and the individual concepts of data are interrelated based upon the inputted individual concept of data generating the output.
24 . The method of claim 21 , wherein the dependency graph includes the one or more dependencies associated with each individual concept of data and each individual concept of data is connected to at least one other individual concept of data in the dependency graph.
25 . The method of claim 21 , wherein the computation path is used to determine an optimal path of accessing an individual concept of data based upon the one or more dependencies associated with each of the individual concepts of data.
26 . The method of claim 21 , wherein the computation path is determined based upon the one or more dependencies associated with each of the individual concepts of data, and wherein the one or more dependencies are associated with each individual concept of data based upon an input and an output associated with each individual concept of data.
27 . The method of claim 21 , wherein the one of the individual concepts of data and each dependency associated with the one of the individual concept of data is stream processed, and wherein the method further comprises batch processing at least one other individual concepts of data and each dependency associated with the individual concept of data based upon the computation path in response to a request for data relating to the monolithic unit of data.
28 . The method of claim 21 , further comprising:
receiving a new individual concept of data associated with the monolithic unit of data after the dependency graph has been generated, determining one or more dependencies associated with the new individual concept of data; and updating the dependency graph by adding the one or more dependencies associated with the new individual concept of data.
29 . The method of claim 21 , further comprising:
identifying an individual concept of data to be updated; removing the one or more dependencies associated with the individual concept of data to be updated; determining one or more new dependencies associated with the individual concept of data to be updated; updating the dependency graph based upon the one or more new dependencies associated with the individual concept of data to be updated; deriving a new computation path based on the updated dependency graph; and stream processing the individual concept of data to be updated and each dependency associated with the individual concept of data to be updated based upon the new computation path in response to a request for data relating to the monolithic unit of data.
30 . A system comprising:
an execution engine configured to:
identify atomic concepts from a monolithic unit of data,
determine one or more dependencies associated with each of the atomic concepts, wherein each dependency interrelates each atomic concept with at least one other atomic concept;
derive a computation path based on the one or more determined dependencies; and
a processing engine configured to:
process at least one atomic concept and the one or more dependencies associated with the atomic concept based upon the derived computation path.
31 . The system of claim 30 , wherein the execution engine is further configured to analyze an input and an output associated with each of the atomic concepts.
32 . The system of claim 31 , wherein the output of the atomic concept is dependent upon the input, the input comprises at least one other atomic concept, and wherein the atomic concepts are interrelated based upon the inputted atomic concept generating the output.
33 . The system of claim 30 , wherein the execution engine is further configured to generate a dependency graph based upon the one or more dependencies.
34 . The system of claim 33 , wherein the dependency graph includes the one or more dependencies associated with each atomic concept, and wherein each atomic concept is connected to at least one other atomic concept in the dependency graph.
35 . The system of claim 30 , wherein the computation path is used to determine an optimal path of accessing an atomic concept based upon the one or more dependencies associated with each of the atomic concepts.
36 . The system of claim 30 , wherein the computation path is determined based upon the one or more dependencies associated with each of the atomic concepts, and wherein the one or more dependencies are associated with each atomic concept based upon an input and an output associated with each atomic concept.
37 . The system of claim 30 , wherein the at least one atomic concepts and each dependency associated with the at least one atomic concept is stream processed by the processing engine, and where the processing engine is further configured to processing at least one other atomic concept and each dependency associated with the one other atomic concept based upon the computation path in response to a request for data relating to the monolithic unit of data.
38 . The system of claim 30 , wherein an execution engine is further configured to:
receive a new atomic concept associated with the monolithic unit of data after the computation path has been derived; determine one or more dependencies associated with the new atomic concept; and derive a new computation path based upon the one or more dependencies associated with the new atomic concept.
39 . The system of claim 30 , wherein an execution engine is further configured to:
identify a atomic concept to be updated; remove the one or more dependencies associated with the atomic concept to be updated; determine one or more new dependencies associated with the atomic concept to be updated; and derive a new computation path based on the one or more new dependencies associated with the atomic concept to be updated; and wherein the processing engine is further configured to: stream process the atomic concept to be updated and each dependency associated with the atomic concept to be updated based upon the new computation path in response to a request for data relating to the monolithic unit of data.
40 . A method for stream processing of a data, comprising:
identifying multiple individual signals from a monolith of data; associating dependencies that interrelate the individual signals of the monolith of data; and stream processing one or more individual signals and dependencies associated with the one or more individual signals in response to a data request relating to the monolith of data.Join the waitlist — get patent alerts
Track US2020293536A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.