Methods and systems for merging data sets
Abstract
Systems and methods for merging data sets are provided. Data sets are merged based upon a process which begins by sorting data sets. Data sets each include at least one data set key column storing at least one data set key column record. The key column record subsets include at least one data set key column record. Based upon the identification of the first and second key column record subsets, a working data set is assembled. The working data set includes at least the first and second key column record subset, a first and second last record indicator corresponding to the last record of the first and second key column record subset respectively, and a first and second position indicator associating the data set key column records with the data sets respectively. The working data set is sorted in accordance with a selected one or more key column record subsets.
Claims
exact text as granted — not AI-modified1 - 7 . (canceled)
8 . A method comprising:
assembling, via a computing device, a first working data set comprising a first portion of sorted first records from a first sorted data set and a second portion of sorted second records from a second sorted data set, the sorted first portion having a last first record and identifying a first number of sorted records and the sorted second portion having a last second record and identifying a second number of sorted records, said assembling based in part on location information of a respective data set and the respective number of sorted records in each respective record subset, the location information corresponds to respective position indicators that associate respective key columns within the sorted data sets with the data sets; sorting, via the computing device, the first working data set, identifying a last sorted working record in the first working data set, the last sorted working record corresponding to the first occurrence in the sorted first working data set of either the last first record or last second record; and identifying, via the computing device, the last sorted working record and any records following the last sorted working record that are equivalent to the last sorted working record as a sorting cut-off point.
9 . The method of claim 8 further comprising:
assembling a second working data set comprised of all first and second records in the first working data set following the sorting cut-off point, a third portion of sorted first records from a first sorted data set and a fourth portion of sorted second records from a second sorted data set.
10 . The method of claim 8 further comprising:
copying all of the records before the sorting cut-off point to a combined data set.
11 - 13 . (canceled)
17 - 22 . (canceled)
23 . The method of claim 8 , wherein the first number and the second number are chosen based on a memory capacity.
24 . The method of claim 23 , wherein the memory capacity corresponds to a cache size.
25 . The method of claim 8 , wherein the first number and the second number are equal.
26 . The method of claim 8 , further comprising:
associating with each record in the first working data set a position indicator, the position indicator identifying the data set and location within the data set from which the record was obtained.
27 . The method of claim 8 , further comprising:
associating with each record in the first working data set a position indicator, the position indicator identifying the data set and location within the data set from which the record was obtained.
28 . A non-transitory computer-readable storage medium having tangibly stored thereon computer executable instructions, that when executed by a computing device, performs a method comprising:
assembling a first working data set comprising a first portion of sorted first records from a first sorted data set and a second portion of sorted second records from a second sorted data set, the sorted first portion having a last first record and identifying a first number of sorted records and the sorted second portion having a last second record and identifying a second number of sorted records, said assembling based in part on location information of a respective data set and the respective number of sorted records in each respective record subset, the location information corresponds to respective position indicators that associate respective key columns within the sorted data sets with the data sets; sorting the first working data set, identifying a last sorted working record in the first working data set, the last sorted working record corresponding to the first occurrence in the sorted first working data set of either the last first record or last second record; and identifying the last sorted working record and any records following the last sorted working record that are equivalent to the last sorted working record as a sorting cut-off point.
29 . The non-transitory computer-readable storage medium of claim 28 further comprising:
assembling a second working data set comprised of all first and second records in the first working data set following the sorting cut-off point, a third portion of sorted first records from a first sorted data set and a fourth portion of sorted second records from a second sorted data set.
30 . The non-transitory computer-readable storage medium of claim 28 further comprising:
copying all of the records before the sorting cut-off point to a combined data set.
31 . The non-transitory computer-readable storage medium of claim 28 , wherein the first number and the second number are chosen based on a memory capacity.
32 . The non-transitory computer-readable storage medium of claim 31 , wherein the memory capacity corresponds to a cache size.
33 . The non-transitory computer-readable storage medium of claim 28 , wherein the first number and the second number are equal.
34 . The non-transitory computer-readable storage medium of claim 28 , further comprising:
associating with each record in the first working data set a position indicator, the position indicator identifying the data set and location within the data set from which the record was obtained.
35 . A system comprising:
a plurality of processors; a datastore that stores a plurality of data sets wherein each of the data sets includes at least one key column comprised of associated data records, each key column comprises information identifying a location of each of the associated data records within the respective data set; a request module implemented by at least one of said plurality of processors that requests a transformation of the associated data records of at least a portion of the plurality of the data sets stored within the datastore; and a data transformation module implemented by at least one of said plurality of processors that performs the steps of:
assembling a first working data set comprising a first portion of sorted first records from a first sorted data set and a second portion of sorted second records from a second sorted data set, the sorted first portion having a last first record and identifying a first number of sorted records and the sorted second portion having a last second record and identifying a second number of sorted records, said assembling based in part on location information of a respective data set and the respective number of sorted records in each respective record subset, the location information corresponds to respective position indicators that associate respective key columns within the sorted data sets with the data sets;
sorting the first working data set, identifying a last sorted working record in the first working data set, the last sorted working record corresponding to the first occurrence in the sorted first working data set of either the last first record or last second record; and
identifying the last sorted working record and any records following the last sorted working record that are equivalent to the last sorted working record as a sorting cut-off point.
36 . The system of claim 35 wherein the first number and second number are chosen by the data transformation module based on a memory available to the transformation module.
37 . The system of claim 36 wherein the memory capacity corresponds to a cache size of a cache used by the transformation module.
38 . The system of claim 35 wherein the first number and second number are equal.
39 . The system of claim 35 wherein at least two of the first records in the first record subset are duplicate records and the transformation module includes only one of the at least two duplicate records in the working data set and identifies the only one of the at least two duplicate records with a duplicate record indicator.Join the waitlist — get patent alerts
Track US2012131022A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.