Apparatus, systems, and methods for batch and realtime data processing
Abstract
A traditional data processing system is configured to process input data either in batch or in real-time. On one hand, a batch data processing system is limiting because the batch data processing often cannot take into account any data received during the batch data processing. On the other hand, a real-time data processing system is limiting because the real-time system often cannot scale. The real-time data processing system is often limited to dealing with primitive data types and/or a small amount of data. Therefore, it is desirable to address the limitations of the batch data processing system and the real-time data processing system by combining the benefits of the batch data processing system and the real-time data processing system into a single data processing system.
Claims
exact text as granted — not AI-modifiedWe claim:
1 .- 23 . (canceled)
24 . A method comprising:
generating a first summary data using a set of data, the first summary data includes a first entity identifier and a first value associated with the first entity identifier; generating a second summary data using the first set of data and a second set of data, the second summary data includes a second entity identifier and a second value associated with the second entity identifier; determining a difference between the first summary data and the second summary data; and updating the first summary data based upon the difference between the first summary data and the second summary data.
25 . The method of claim 24 , wherein the first set of data comprises bulk data input.
26 . The method of claim 25 , wherein the bulk data input comprises one or more of:
raw information received from one or more contributors; web-crawler data received from a web-crawler; or data received from a storage center.
27 . The method of claim 25 , wherein the second set of data comprises intermittent data.
28 . The method of claim 27 , wherein the intermittent data comprises real-time data submissions.
29 . The method of claim 27 , further comprising:
formatting the bulk data input into structured data; group a plurality of elements in the structured data; and generate an entity identifier for the plurality of elements.
30 . The method of claim 24 , wherein the first summary data comprises a first entity identifier and the second summary data comprises a second entity identifier, and wherein determining the difference between the first summary data and the second summary data comprises:
determining whether a first value of the first entity identifier and a second value of the second entity identifier are equal; and when the first value and second value are equal, comparing data associated with the first entity identifier and the second entity identifier.
31 . A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method comprising:
generating a first summary data using a set of data, the first summary data includes a first entity identifier and a first value associated with the first entity identifier; generating a second summary data using the first set of data and a second set of data, the second summary data includes a second entity identifier and a second value associated with the second entity identifier; determining a difference between the first summary data and the second summary data; and updating the first summary data based upon the difference between the first summary data and the second summary data.
32 . The non-transitory computer-readable storage medium of claim 31 , wherein the method further comprises:
generate a third data set by combining the first data set and the second data set; and generate third summary data for the third data set.
33 . The non-transitory computer-readable storage medium of claim 31 , wherein the first data set comprises one or more of:
raw information received from one or more contributors; web-crawler data received from a web-crawler; or data received from a storage center.
34 . The non-transitory computer-readable storage medium of claim 33 , wherein the second data set comprises data received from a user to correct information in the first data set.
35 . The non-transitory computer-readable storage medium of claim 34 , wherein the first summary data comprises a first entity identifier and the second summary data comprises a second entity identifier, and wherein determining the difference between the first summary data and the second summary data comprises:
determining whether a first value of the first entity identifier and a second value of the second entity identifier are equal; and when the first value and second value are equal, comparing data associated with the first entity identifier and the second entity identifier.
36 . The non-transitory computer-readable storage medium of claim 31 , wherein the method further comprises processing the first data set to generate a first structured data set.
37 . The non-transitory computer-readable storage medium of claim 36 , wherein the second data set comprises real-time data submissions, and wherein the method further comprises processing the second data set to generate a second structured data set in response to receiving the real-time data submissions.
38 . A system comprising:
at least one processor; and memory encoding computer-executable instructions that, when executed by the at least one processor, perform a method comprising:
generating a first summary data using a set of data, the first summary data includes a first entity identifier and a first value associated with the first entity identifier;
generating a second summary data using the first set of data and a second set of data, the second summary data includes a second entity identifier and a second value associated with the second entity identifier;
determining a difference between the first summary data and the second summary data; and
updating the first summary data based upon the difference between the first summary data and the second summary data.
39 . The system of claim 38 , wherein the first set of data comprises bulk data input.
40 . The system of claim 39 , wherein the bulk data input comprises one or more of:
raw information received from one or more contributors; web-crawler data received from a web-crawler; or data received from a storage center.
41 . The system of claim 39 , wherein the second set of data comprises intermittent data.
42 . The system of claim 41 , wherein the intermittent data comprises real-time data submissions.
43 . The system of claim 41 , further comprising:
formatting the bulk data input into structured data; group a plurality of elements in the structured data; and generate an entity identifier for the plurality of elements.
44 . The system of claim 24 , wherein the first summary data comprises a first entity identifier and the second summary data comprises a second entity identifier, and wherein determining the difference between the first summary data and the second summary data comprises:
determining whether a first value of the first entity identifier and a second value of the second entity identifier are equal; and when the first value and second value are equal, comparing data associated with the first entity identifier and the second entity identifier.Join the waitlist — get patent alerts
Track US2021374109A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.