US2020356579A1PendingUtilityA1

Data clustering based on candidate queries

Assignee: AB INITIO TECHNOLOGY LLCPriority: Nov 15, 2011Filed: Feb 3, 2020Published: Nov 12, 2020
Est. expiryNov 15, 2031(~5.3 yrs left)· nominal 20-yr term from priority
G06F 16/24534G06F 16/20G06F 16/278G06F 16/3338G06F 16/285
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Received data records, each including one or more values in one or more fields, are processed to identify a matched data cluster. The processing includes: for selected data records, generating a query from one or more values; identifying one or more candidate data records from the received data records using the query; determining whether or not the selected data record satisfies a cluster membership criterion for at least one candidate data cluster of one or more existing data clusters containing the candidate records; and selecting the matched data cluster from among one or more candidate data clusters based at least in part on a growth criterion for the candidate data clusters, or initializing the matched data cluster with the selected data record if the selected data record does not satisfy a cluster membership criterion for any of the existing data clusters or based on a result of the growth criterion.

Claims

exact text as granted — not AI-modified
1 . A method including:
 receiving data records, the received data records each including one or more values in one or more fields and   processing the received data records to identify a matched data cluster to associate with each received data record, the processing including,
 for selected data records from the received data records, generating a query from the one or more values included in the selected data record; 
 identifying one or more candidate data records from the received data records using the query; 
 determining that the selected data record satisfies a cluster membership criterion for a plurality of existing data clusters; 
 selecting the matched data cluster for the selected data record from the plurality of existing data clusters; and 
 providing an indication that the selected data record satisfies the cluster membership criterion for the plurality of existing clusters and of the selected matched data cluster to a user. 
   
     
     
         2 . The method of  claim 1 , wherein generating the query includes identifying tokens, wherein each of the tokens includes a value that is found in a field set, and wherein the field set includes a field of the selected data record. 
     
     
         3 . The method of  claim 1 , wherein generating the query includes identifying tokens, wherein each of the tokens includes a fragment of a value that is found in a field set that includes at least one field of the selected data record. 
     
     
         4 . The method of  claim 1 , wherein processing the received data records to identify a matched data cluster to associate with each received data record further includes sorting an initial set of the received data records based on a distinguishability criterion, wherein the distinguishability criterion determines a degree to which a value included in a particular data record distinguishes that particular data record from other data records. 
     
     
         5 . The method of  claim 1 , wherein selecting the matched data cluster for the selected data record from the plurality of existing data clusters includes comparing the selected data record to a representative data record from an existing data cluster, based on the comparison, calculating a comparison score, determining that the comparison score exceeds a first threshold, and as a result of having determined that the comparison score exceeds the first threshold, selecting the existing data cluster as the matched data cluster. 
     
     
         6 . The method of  claim 1 , wherein selecting the matched data cluster from the plurality of existing data clusters includes determining that the selected data record satisfies a cluster-membership criterion for a particular data cluster from among the existing data clusters and, based on the selected data record having satisfied the cluster-membership criterion, selecting the particular data cluster to be the matched data cluster. 
     
     
         7 . The method of  claim 1 , wherein identifying one or more candidate data records from the received data records includes comparing the query to queries in a data store and determining that the query maps to a first cluster, wherein the data store maps queries to existing clusters. 
     
     
         8 . The method of  claim 1 , further including receiving a request to map the selected data record to a second cluster and updating a data store to map the query to a second cluster instead of to a first cluster, wherein the data store maps queries to existing clusters. 
     
     
         9 . The method of  claim 1 , further including receiving a request to map the selected data record to a new cluster, updating a data store with a new cluster indicator, generating a new cluster, and assigning the selected data record to the new cluster, wherein the data store maps queries to existing clusters. 
     
     
         10 . The method of  claim 1 , further including receiving a request to confirm membership of the selected data record in a first cluster, storing information in a data store, updating the data store in response to a request that is associated with another data record, and maintaining membership of the selected data record in the first cluster notwithstanding the request that is associated with another data record, wherein the data store maps queries to existing clusters. 
     
     
         11 . The method of  claim 1 , further including receiving a first request, the first request being a request to exclude membership of the selected data record in a first cluster, in response to the first request, updating a data store to change membership of the selected data record, receiving a second request to update the data store, the second request being associated with another data record, updating the data store in response to the second request, wherein, as a result of information having been written into the data store when updating the data store in response to the first request, the selected data record continues to be excluded from membership in the first cluster and wherein the data store maps queries to existing clusters. 
     
     
         12 . The method of  claim 1 , further including receiving input from a user to approve a proposed association of one of the received data records with a matched data cluster. 
     
     
         13 . The method of  claim 1 , further including receiving, from a user, an instruction to modify an association of one of the received data records with a matched data cluster. 
     
     
         14 . The method of  claim 1 , further including selecting a particular data cluster to be the matched data cluster for the selected data record and storing information identifying an existing data cluster that was not selected as the matched data cluster for the selected data record. 
     
     
         15 . The method of  claim 1 , further including selecting the selected data record from an initial set of data records that have been sorted based on a distinguishability criterion that determines a degree to which a value included in a particular data record distinguishes that particular data record from other data records. 
     
     
         16 . The method of  claim 1 , wherein processing the received data records includes sorting the data records based on a number of fields that are populated with a particular value, wherein the number of fields so populated is indicative of an extent to which the value distinguishes a data record from other data records. 
     
     
         17 . The method of  claim 1 , wherein processing the received data records includes sorting the data records based on a number of tokens in one or more fields of the data records, wherein the number of tokens determines a degree to which a value included in a particular data record distinguishes that particular data record from other data records, wherein each of the tokens includes either a value or a fragment of a value that is found in a either a field of the selected data record or a combination of fields of the selected data record. 
     
     
         18 . A computing system, the computing system including
 an input device or port configured to receive data records, the received data records each including one or more values in one or more field and   at least one processor configured to process the received data records to identify a matched data cluster to associate with each received data record, the processing including, for selected data records from the received data records,
 generating a query from the one or more values included in the selected data record, 
 identifying one or more candidate data records from the received data records using the query, 
 determining that the selected data record satisfies a cluster membership criterion for a plurality of existing data clusters, 
 selecting the matched data cluster for the selected data record from the plurality of existing data clusters, and 
 providing an indication that the selected data record satisfies the cluster membership criterion for the plurality of existing clusters and of the selected matched data cluster to a user. 
   
     
     
         19 . A computer program that is stored on a computer-readable storage medium, the computer program including instructions for causing a computing system
 to receive data records and   to process the received data records to identify a matched data cluster to associate with each received data record,   wherein the received data records each include one or more values in one or more fields,   wherein the instructions for causing the computer system to process the received data records include instructions for causing the computer system, for selected data records from the received data records,
 to generate a query from the one or more values included in the selected data record, 
 to identify one or more candidate data records from the received data records using the query, 
 to determine that the selected data record satisfies a cluster membership criterion for a plurality of existing data clusters, 
 to select the matched data cluster for the selected data record from the plurality of existing data clusters, and 
 to provide an indication that the selected data record satisfies the cluster membership criterion for the plurality of existing clusters and of the selected matched data cluster to a user. 
   
     
     
         20 . A computing system, the computing system including
 means for receiving data records, the received data records each including one or more values in one or more field and   means for processing the received data records to identify a matched data cluster to associate with each received data record, the processing including, for selected data records from the received data records,
 generating a query from the one or more values included in the selected data record, 
 identifying one or more candidate data records from the received data records using the query, 
 determining that the selected data record satisfies a cluster membership criterion for a plurality of existing data clusters, 
 selecting the matched data cluster for the selected data record from the plurality of existing data clusters, and 
 providing an indication that the selected data record satisfies the cluster membership criterion for the plurality of existing clusters and of the selected matched data cluster to a user.

Join the waitlist — get patent alerts

Track US2020356579A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.