US2012131022A1PendingUtilityA1

Methods and systems for merging data sets

Assignee: UPPALA RADHAKRISHNAPriority: Sep 14, 2007Filed: Jan 30, 2012Published: May 24, 2012
Est. expirySep 14, 2027(~1.1 yrs left)· nominal 20-yr term from priority
G06F 7/32
32
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for merging data sets are provided. Data sets are merged based upon a process which begins by sorting data sets. Data sets each include at least one data set key column storing at least one data set key column record. The key column record subsets include at least one data set key column record. Based upon the identification of the first and second key column record subsets, a working data set is assembled. The working data set includes at least the first and second key column record subset, a first and second last record indicator corresponding to the last record of the first and second key column record subset respectively, and a first and second position indicator associating the data set key column records with the data sets respectively. The working data set is sorted in accordance with a selected one or more key column record subsets.

Claims

exact text as granted — not AI-modified
1 - 7 . (canceled) 
     
     
         8 . A method comprising:
 assembling, via a computing device, a first working data set comprising a first portion of sorted first records from a first sorted data set and a second portion of sorted second records from a second sorted data set, the sorted first portion having a last first record and identifying a first number of sorted records and the sorted second portion having a last second record and identifying a second number of sorted records, said assembling based in part on location information of a respective data set and the respective number of sorted records in each respective record subset, the location information corresponds to respective position indicators that associate respective key columns within the sorted data sets with the data sets;   sorting, via the computing device, the first working data set, identifying a last sorted working record in the first working data set, the last sorted working record corresponding to the first occurrence in the sorted first working data set of either the last first record or last second record; and   identifying, via the computing device, the last sorted working record and any records following the last sorted working record that are equivalent to the last sorted working record as a sorting cut-off point.   
     
     
         9 . The method of  claim 8  further comprising:
 assembling a second working data set comprised of all first and second records in the first working data set following the sorting cut-off point, a third portion of sorted first records from a first sorted data set and a fourth portion of sorted second records from a second sorted data set. 
 
     
     
         10 . The method of  claim 8  further comprising:
 copying all of the records before the sorting cut-off point to a combined data set. 
 
     
     
         11 - 13 . (canceled) 
     
     
         17 - 22 . (canceled) 
     
     
         23 . The method of  claim 8 , wherein the first number and the second number are chosen based on a memory capacity. 
     
     
         24 . The method of  claim 23 , wherein the memory capacity corresponds to a cache size. 
     
     
         25 . The method of  claim 8 , wherein the first number and the second number are equal. 
     
     
         26 . The method of  claim 8 , further comprising:
 associating with each record in the first working data set a position indicator, the position indicator identifying the data set and location within the data set from which the record was obtained.   
     
     
         27 . The method of  claim 8 , further comprising:
 associating with each record in the first working data set a position indicator, the position indicator identifying the data set and location within the data set from which the record was obtained.   
     
     
         28 . A non-transitory computer-readable storage medium having tangibly stored thereon computer executable instructions, that when executed by a computing device, performs a method comprising:
 assembling a first working data set comprising a first portion of sorted first records from a first sorted data set and a second portion of sorted second records from a second sorted data set, the sorted first portion having a last first record and identifying a first number of sorted records and the sorted second portion having a last second record and identifying a second number of sorted records, said assembling based in part on location information of a respective data set and the respective number of sorted records in each respective record subset, the location information corresponds to respective position indicators that associate respective key columns within the sorted data sets with the data sets;   sorting the first working data set, identifying a last sorted working record in the first working data set, the last sorted working record corresponding to the first occurrence in the sorted first working data set of either the last first record or last second record; and   identifying the last sorted working record and any records following the last sorted working record that are equivalent to the last sorted working record as a sorting cut-off point.   
     
     
         29 . The non-transitory computer-readable storage medium of  claim 28  further comprising:
 assembling a second working data set comprised of all first and second records in the first working data set following the sorting cut-off point, a third portion of sorted first records from a first sorted data set and a fourth portion of sorted second records from a second sorted data set. 
 
     
     
         30 . The non-transitory computer-readable storage medium of  claim 28  further comprising:
 copying all of the records before the sorting cut-off point to a combined data set. 
 
     
     
         31 . The non-transitory computer-readable storage medium of  claim 28 , wherein the first number and the second number are chosen based on a memory capacity. 
     
     
         32 . The non-transitory computer-readable storage medium of  claim 31 , wherein the memory capacity corresponds to a cache size. 
     
     
         33 . The non-transitory computer-readable storage medium of  claim 28 , wherein the first number and the second number are equal. 
     
     
         34 . The non-transitory computer-readable storage medium of  claim 28 , further comprising:
 associating with each record in the first working data set a position indicator, the position indicator identifying the data set and location within the data set from which the record was obtained.   
     
     
         35 . A system comprising:
 a plurality of processors;   a datastore that stores a plurality of data sets wherein each of the data sets includes at least one key column comprised of associated data records, each key column comprises information identifying a location of each of the associated data records within the respective data set;   a request module implemented by at least one of said plurality of processors that requests a transformation of the associated data records of at least a portion of the plurality of the data sets stored within the datastore; and   a data transformation module implemented by at least one of said plurality of processors that performs the steps of:
 assembling a first working data set comprising a first portion of sorted first records from a first sorted data set and a second portion of sorted second records from a second sorted data set, the sorted first portion having a last first record and identifying a first number of sorted records and the sorted second portion having a last second record and identifying a second number of sorted records, said assembling based in part on location information of a respective data set and the respective number of sorted records in each respective record subset, the location information corresponds to respective position indicators that associate respective key columns within the sorted data sets with the data sets; 
 sorting the first working data set, identifying a last sorted working record in the first working data set, the last sorted working record corresponding to the first occurrence in the sorted first working data set of either the last first record or last second record; and 
 identifying the last sorted working record and any records following the last sorted working record that are equivalent to the last sorted working record as a sorting cut-off point. 
   
     
     
         36 . The system of  claim 35  wherein the first number and second number are chosen by the data transformation module based on a memory available to the transformation module. 
     
     
         37 . The system of  claim 36  wherein the memory capacity corresponds to a cache size of a cache used by the transformation module. 
     
     
         38 . The system of  claim 35  wherein the first number and second number are equal. 
     
     
         39 . The system of  claim 35  wherein at least two of the first records in the first record subset are duplicate records and the transformation module includes only one of the at least two duplicate records in the working data set and identifies the only one of the at least two duplicate records with a duplicate record indicator.

Join the waitlist — get patent alerts

Track US2012131022A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.