System and Method for Organizing Data
Abstract
A system and method for organizing raw data from one or more sources uses an improved mechanism for identifying duplicate data between fields (e.g., columns) in the databases. The fields may be similar fields within a single database or similar or identical fields within a pair of databases and as organized as arrays or field vectors. The present invention sorts each of the field vectors and if necessary, partitions them by common value. A number of comparisons required to identify the duplicate data between the field vectors is reduced by feeding back a difference between the compared values. This difference is used to adjust indices into the field vectors for subsequent comparison.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for sorting data comprising:
receiving, by at least one computing processor, value to be sorted; determining, by the at least one computing processor, a first position in a vector where said value is to be included; retrieving, by the at least one computing processor, a vector value from said vector at said first position; determining, by the at least one computing processor, a difference between said value and said vector value; and determining, by the at least one computing processor, a new position in said vector based at least in part on said difference.
2 . The method of claim 1 , wherein said determining a new position comprises determining a new position in said vector based at least in part on said first position.
3 . A computer-implemented method for identifying duplicate data between a first vector and a second vector comprising:
sorting, by at least one computing processor, values in the first vector in an increasing order; partitioning, by the at least one computing processor, the sorted values in the first vector into first sets, wherein at least one of the first sets includes a plurality of sorted values that have a common value; sorting, by the at least one computing processor, values in the second vector in said increasing order; partitioning, by the at least one computing processor, the sorted values in the second vector into second sets, wherein at least one of the second sets includes a plurality of sorted values that have a common value; comparing, by the at least one computing processor, a first sorted value at a first index in a first one of the first sets of the partitioned first vector with a second sorted value at a second index in a first one of the second sets of the partitioned second vector; adjusting, by the at one computing processor, said first index to a next one of the first sets of the partitioned first vector if said first sorted value is less than said second sorted value; adjusting, by the at least one computing processor, said second index to a next one of the second sets of the partitioned second vector if said second sorted value is less than said first sorted value; and identifying, by the at least one computing processor, said first and second sorted values as duplicate data if said first sorted value is equal to said second sorted value.
4 . A computer-implemented method for identifying duplicate data between a first vector and a second vector comprising:
sorting, by at least one computing processor, values in the first vector in a decreasing order; partitioning, by the at least one computing processor, the sorted values in the first vector into first sets, wherein at least one of the first sets includes a plurality of sorted values that have a common value; sorting, by the at least one computing processor, values in the second vector in said decreasing order; partitioning, by the at least one computing processor, the sorted values in the second vector into second sets, wherein at least one of the second sets includes a plurality of members that have a common value; comparing, by the at least one computing processor, a first sorted value at a first index in a first one of the first sets of the partitioned first vector with a second sorted value at a second index in a first one of the second sets of the partitioned second vector; adjusting, by the at least one computing processor, said first index to a next one of the first sets of the partitioned first vector if said first sorted value is greater than said second sorted value; adjusting, by the at least one computing processor, said second index to a next one of the second sets of the partitioned second vector if said second sorted value is greater than said first sorted value; and identifying, by the at least one computing processor, said first and second sorted values as duplicate data if said first sorted value is equal to said second sorted value.
5 . A computer-implemented method for identifying duplicate data between a first vector and a second vector comprising:
sorting, by at least one computing processor, values in the first vector in an increasing order; partitioning, by the at least one computing processor, the sorted values in the first vector into first sets, each of the first sets having at least one sorted value, all sorted values in each of the first sets having a common value, wherein at least one of the first sets includes a plurality of sorted values; sorting, by the at least one computing processor, values in the second vector in said increasing order; partitioning, by the at least one computing processor, the sorted values in the second vector into second sets, each of the second sets having at least one sorted value, all sorted values in each of the second sets having a common value, wherein at least one of the second sets includes a plurality of sorted values; comparing, by the at least one computing processor, a first sorted value at a first index in a first one of the first sets of the partitioned first vector with a second sorted value at a second index in a first one of the second sets of the partitioned second vector; adjusting, by the at least one computing processor, said first index to a next one of the first sets of the partitioned first vector if said first sorted value is less than said second sorted value; adjusting, by the at least one computing processor, said second index to a next one of the second sets of the partitioned second vector if said second sorted value is less than said first sorted value; and identifying, by the at least one computing processor, said first and second sorted values as duplicate data if said first sorted value is equal to said second sorted value.
6 . The method of claim 5 , wherein at least one of the second sets includes a plurality of sorted values.
7 . The method of claim 1 , wherein said determining a difference between said value and said vector value comprises computing a difference between said value and said vector value.
8 . The method of claim 1 , wherein said determining a difference between said value and said vector value comprises determining a difference between said value and said vector value either by subtracting said value from said vector value or by subtracting said vector value from said value.
9 . The method of claim 3 , further comprising storing said duplicate data.
10 . The method of claim 4 , further comprising storing said duplicate data.
11 . The method of claim 5 , further comprising storing said duplicate data.
12 . A computer-implemented method comprising:
sorting, by at least one computing processor, values in a first vector in an increasing order; partitioning, by the at least one computing processor, the sorted values in the first vector into first sets, each of the first sets having at least one sorted value, all sorted values in each of the first sets sharing a common value, wherein at least one of the first sets includes a plurality of sorted values; sorting, by the at least one computing processor, values in a second vector in said increasing order; partitioning, by the at least one computing processor, the sorted values in the second vector into second sets, each of the second sets having at least one sorted value, all sorted values in each of the second sets having a common value; comparing, by the at least one computing processor, a first sorted value at a first index in a first one of the first sets of the partitioned first vector with a second sorted value at a second index in a first one of the second sets of the partitioned second vector, wherein the first one of the first sets of the partitioned first vector includes a plurality of sorted values; when said first sorted value is less than said second sorted value, adjusting said first index to a next one of the first sets of the partitioned first vector; when said second sorted value is less than said first sorted value, adjusting said second index to a next one of the second sets of the partitioned second vector; and when said first sorted value is equal to said second sorted value, adjusting said first index to a next one of the first sets of the partitioned first vector and adjusting said second index to a next one of the second sets of the partitioned second vector.
13 . The method of claim 12 , wherein the first sorted value is one of the at least one sorted value in the first one of the first sets of the partitioned first vector, and wherein the second sorted value is one of the at least one sorted value in the second one of the second sets of the partitioned second vector.
14 . The method of claim 1 , wherein determining a new position in said vector based at least in part on said difference comprises determining the new position in said vector based at least in part on a magnitude of said difference.Join the waitlist — get patent alerts
Track US2013297568A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.