US2022092041A1PendingUtilityA1

Automatic entity resolution data cleaning

Assignee: GROUPON INCPriority: Mar 18, 2015Filed: Aug 30, 2021Published: Mar 24, 2022
Est. expiryMar 18, 2035(~8.6 yrs left)· nominal 20-yr term from priority
G06F 16/2365G06N 20/00G06F 16/215G06F 16/9024G06F 16/285
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In general, embodiments of the present invention provide systems, methods and computer readable media for automatic cleaning of entity resolution (ER) data persistently stored in a data repository.

Claims

exact text as granted — not AI-modified
1 - 24 . (canceled) 
     
     
         25 . A computer program product, stored on a non-transitory computer readable medium, comprising instructions that when executed on one or more computers cause the one or more computers to perform operations implementing automatic entity resolution (ER) data cleaning, the operations comprising:
 receiving a group of N references to a candidate ER error entity,   wherein the references and the candidate ER entity are persistent data stored in a data repository, and wherein the references each include attribute data describing the candidate ER entity; identifying a set of related references within the group of N references,   wherein each related reference is related to at least one other reference of the group of N references;   for each pair of references identified as a part of the set of the related references, calculating an ER score representing a likelihood that the pair of related references refers to the candidate ER error entity;   determining based on the calculation of each ER score, that an ER error has occurred in grouping the references;   selecting a set of the pairs of the related references for validation of their ER scores,   sending the selected set of the pairs to an oracle for validation of their ER scores;   receiving validated ER scores from the oracle, the validated ER scores indicative of a match or non-match and are associated with a value between 0 and 1;   adjusting at least one of the validated ER scores by calculating a paralyzed ER score including determining whether the validated ER score satisfies a match threshold; and   performing a recursive re-grouping process comprising:   re-grouping, using a grouping method, the set of related references based in part on their respectively associated ER scores forming additional new input data based on the re-grouping of the references, the respectively associated ER scores comprising the paralyzed ER score;   re-assigning the group of each of the N references based on the re-grouping; and   recalculating a pairwise matching of the additional new input data formed based on the re-grouping of the references.   
     
     
         26 . The computer program product of  claim 25 , wherein the group of N references is represented as a graph, wherein each reference is represented as a graph node, and wherein each graph edge represents the relationship between a pair of nodes connected by the edge. 
     
     
         27 . The computer program product of  claim 26 , wherein the ER score calculated for a pair of the related references is associated with the edge connecting the pair of related references. 
     
     
         28 . The computer program product of  claim 25 , wherein receiving the group of N references is preceded by selecting the candidate ER error entity based on an ER error score that represents the likelihood that the entity is described by erroneous ER data. 
     
     
         29 . The computer program product of  claim 28 , wherein the ER error score includes one or more of a count of the number of unique references in the group of N references and a count of the number of duplicates of the candidate ER error entity that are identified within the persistent data stored in the data repository. 
     
     
         30 . The computer program product of  claim 28 , wherein the group of N references is represented as a graph, and wherein the ER error score may be derived based in part on an analysis of the graph edges. 
     
     
         31 . The computer program product of  claim 25 , wherein calculation of each ER score is implemented using a machine learning algorithm, wherein a binary classifier, derived using supervised machine learning, is trained to return a result label of “match” or “no match” as a decision of whether or not an input pair of entity references describes the same entity, and wherein the result label is returned with a value of an ER score representative of a certainty in the decision. 
     
     
         32 . A method comprising:
 receiving a group of N references to a candidate ER error entity,   wherein the references and the candidate ER entity are persistent data stored in a data repository, and wherein the references each include attribute data describing the candidate ER entity; identifying a set of related references within the group of N references,   wherein each related reference is related to at least one other reference of the group of N references;   for each pair of references identified as a part of the set of the related references, calculating an ER score representing a likelihood that the pair of related references refers to the candidate ER error entity;   determining based on the calculation of each ER score, that an ER error has occurred in grouping the references;   selecting a set of the pairs of the related references for validation of their ER scores,   sending the selected set of the pairs to an oracle for validation of their ER scores;   receiving validated ER scores from the oracle, the validated ER scores indicative of a match or non-match and are associated with a value between 0 and 1;   adjusting at least one of the validated ER scores by calculating a paralyzed ER score including determining whether the validated ER score satisfies a match threshold; and   performing a recursive re-grouping process comprising:   re-grouping, using a grouping method, the set of related references based in part on their respectively associated ER scores forming additional new input data based on the re-grouping of the references, the respectively associated ER scores comprising the paralyzed ER score;   re-assigning the group of each of the N references based on the re-grouping; and   recalculating a pairwise matching of the additional new input data formed based on the re-grouping of the references.   
     
     
         33 . The method of  claim 32 , wherein the group of N references is represented as a graph, wherein each reference is represented as a graph node, and wherein each graph edge represents the relationship between a pair of nodes connected by the edge. 
     
     
         34 . The method of  claim 33 , wherein the ER score calculated for a pair of the related references is associated with the edge connecting the pair of related references. 
     
     
         35 . The method of  claim 32 , wherein receiving the group of N references is preceded by selecting the candidate ER error entity based on an ER error score that represents the likelihood that the entity is described by erroneous ER data. 
     
     
         36 . The method of  claim 35 , wherein the ER error score includes one or more of a count of the number of unique references in the group of N references and a count of the number of duplicates of the candidate ER error entity that are identified within the persistent data stored in the data repository. 
     
     
         37 . The method of  claim 35 , wherein the group of N references is represented as a graph, and wherein the ER error score may be derived based in part on an analysis of the graph edges. 
     
     
         38 . The method of  claim 32 , wherein calculation of each ER score is implemented using a machine learning algorithm, wherein a binary classifier, derived using supervised machine learning, is trained to return a result label of “match” or “no match” as a decision of whether or not an input pair of entity references describes the same entity, and wherein the result label is returned with a value of an ER score representative of a certainty in the decision. 
     
     
         39 . A system comprising one or more computers and one or more non-transitory storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations implementing automatic entity resolution (ER) data cleaning, the operations comprising:
 receiving a group of N references to a candidate ER error entity,   wherein the references and the candidate ER entity are persistent data stored in a data repository, and wherein the references each include attribute data describing the candidate ER entity; identifying a set of related references within the group of N references,   wherein each related reference is related to at least one other reference of the group of N references;   for each pair of references identified as a part of the set of the related references, calculating an ER score representing a likelihood that the pair of related references refers to the candidate ER error entity;   determining based on the calculation of each ER score, that an ER error has occurred in grouping the references;   selecting a set of the pairs of the related references for validation of their ER scores,   sending the selected set of the pairs to an oracle for validation of their ER scores;   receiving validated ER scores from the oracle, the validated ER scores indicative of a match or non-match and are associated with a value between 0 and 1;   adjusting at least one of the validated ER scores by calculating a paralyzed ER score including determining whether the validated ER score satisfies a match threshold; and   performing a recursive re-grouping process comprising:   re-grouping, using a grouping method, the set of related references based in part on their respectively associated ER scores forming additional new input data based on the re-grouping of the references, the respectively associated ER scores comprising the paralyzed ER score;   re-assigning the group of each of the N references based on the re-grouping; and   recalculating a pairwise matching of the additional new input data formed based on the re-grouping of the references.   
     
     
         40 . The system of  claim 39 , wherein the group of N references is represented as a graph, wherein each reference is represented as a graph node, and wherein each graph edge represents the relationship between a pair of nodes connected by the edge. 
     
     
         41 . The system of  claim 40 , wherein the ER score calculated for a pair of the related references is associated with the edge connecting the pair of related references. 
     
     
         42 . The system of  claim 39 , wherein receiving the group of N references is preceded by selecting the candidate ER error entity based on an ER error score that represents the likelihood that the entity is described by erroneous ER data. 
     
     
         43 . The system of  claim 42 , wherein the ER error score includes one or more of a count of the number of unique references in the group of N references and a count of the number of duplicates of the candidate ER error entity that are identified within the persistent data stored in the data repository. 
     
     
         44 . The system of  claim 42 , wherein the group of N references is represented as a graph, and wherein the ER error score may be derived based in part on an analysis of the graph edges. 
     
     
         45 . The system of  claim 39 , wherein calculation of each ER score is implemented using a machine learning algorithm, wherein a binary classifier, derived using supervised machine learning, is trained to return a result label of “match” or “no match” as a decision of whether or not an input pair of entity references describes the same entity, and wherein the result label is returned with a value of an ER score representative of a certainty in the decision.

Join the waitlist — get patent alerts

Track US2022092041A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.