US2013332467A1PendingUtilityA1

Linking Data Elements Based on Similarity Data Values and Semantic Annotations

Assignee: BORNEA MIHAELA ANCUTAPriority: Jun 8, 2012Filed: Jul 8, 2012Published: Dec 12, 2013
Est. expiryJun 8, 2032(~5.9 yrs left)· nominal 20-yr term from priority
G06F 16/951
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Data elements from data sources and having a data value set are linked by using hash functions to determine a dimensionally reduced instance signature for each data element based on all data values associated with that data element to yield a plurality of dimensionally reduced instance signatures of equivalent fixed size such that similarities among the data values in the data value sets across all data elements is maintained among the plurality of instance signatures. Candidate pairs of data elements to link are identified using the plurality of instance signatures in locality sensitive hash functions, and a similarity index is generated for each candidate pair using a pre-determined measure of similarity. Candidate pairs of data elements having a similarity index above a given threshold are linked.

Claims

exact text as granted — not AI-modified
1 . A system for linking data elements derived from data sources, the system comprising:
 a data element linking computing system in communication with a plurality of data sources, each data source comprising a plurality of data elements and each data element comprising an associated context within its data source and a data value set comprising a plurality of data values for that data element, the data element linking computing system comprising:   a database to store the data elements;   an instance signature module in communication with the database and configured to determine a dimensionally reduced instance signature for each data element based on all data values associated with that data element to yield a plurality of dimensionally reduced instance signatures of equivalent fixed size;   a plurality of data value hash functions used to generate the dimensionally reduced data instance signature for each data element such that similarities among the data values in the data value sets across all data elements is maintained among the plurality of instance signatures;   a data element pairing module in communication with the instance signature module and the database to identify candidate pairs of data elements to link using the plurality of instance signatures;   a similarity determination module in communication with the data element pairing module and the database to generate a similarity index for each candidate pair using a pre-determined measure of similarity; and   a data element linking module in communication with the similarity determination module to link candidate pairs of data elements having a similarity index above a given threshold.   
     
     
         2 . The system of  claim 1 , wherein each data source comprises a data collection comprising at least one of structured data and semi-structured data. 
     
     
         3 . (canceled) 
     
     
         4 . The system of  claim 1 , wherein the data value hash functions are compatible with the pre-determined measure of similarity used by the similarity determination module to generate the similarity index. 
     
     
         5 . The system of  claim 1 , wherein:
 the database further comprises the plurality of dimensionally reduced instance signatures, each dimensionally reduced instance signature comprising a hash value set having a quantity of hash values equal to the fixed size; and   the instance signature module utilizes a number of data value hash functions equal to the quantity of hash values.   
     
     
         6 . The system of  claim 1 , wherein each instance signature comprises a set of signature hash values comprising a plurality signature hash values equal in number to the plurality of data value hash functions. 
     
     
         7 . The system of  claim 1 , wherein the database further comprises a universal data value set comprising all data values from all data value sets; and
 an occurrence frequency module to determine a frequency of occurrence for each data value within the universal data value set and to assign a similarity weight to each data value based on the occurrence frequency associated with the data value such that the similarity weight is inversely proportional to the occurrence frequency.   
     
     
         8 . The system of  claim 1 , wherein:
 the database further comprises semantic annotations associated with data values in each data set; and   the system further comprises a semantic signature module in communication with the database and the data element pairing module to construct a semantic signature for each data set to yield a plurality of semantic signatures, the data element pairing module using the plurality of instance signatures in combination with the plurality of semantic signatures to identify the candidate pairs.   
     
     
         9 . The system of  claim 8 , wherein the database further comprises a semantic annotation determination module to assign semantic annotations to the data values in each data value set. 
     
     
         10 . The system of  claim 1 , wherein the candidate pairs comprise a subset of all pairs of data elements. 
     
     
         11 . The system of  claim 1 , wherein the data element pairing module further comprises a signature hashing module comprising at least one signature hash function configured to map each signature into one or more of a plurality of categories to generate a categorical signature for each instance signature associated with each data element and to compare categorical signatures to identify pairs of instance signatures mapping to common categories, the identified candidate pairs of data elements associated with the identified pairs of instance signatures. 
     
     
         12 . The system of  claim 11 , wherein the signature hash function comprises a locality sensitive hash function. 
     
     
         13 . The system of  claim 1 , wherein the pre-determined measure of similarity comprises jaccard similarity, cosine similarity or combinations thereof. 
     
     
         14 . A system for linking data elements derived from data sources, the system comprising:
 a data element linking computing system in communication with a plurality of data sources, each data source comprising a plurality of data elements and each data element comprising an associated context within its data source and a data value set comprising a plurality of data values for that data element, the data element linking computing system comprising:   a database to store the data elements, the database comprising semantic annotations associated with data values in each data set;   a semantic signature module in communication with the database to construct a semantic signature for each data set to yield a plurality of semantic signatures;   a data element pairing module in communication with the semantic signature module and the database to identify candidate pairs of data elements to link using the plurality of semantic signatures;   a similarity determination module in communication with the data element pairing module and the database to generate a similarity index for each candidate pair using a pre-determined measure of similarity; and   a data element linking module in communication with the similarity determination module to link candidate pairs of data elements having a similarity index above a given threshold.   
     
     
         15 . The system of  claim 14 , wherein the database further comprises a semantic annotation determination module to assign semantic annotations to the data values in each data value set. 
     
     
         16 . The system of  claim 14 , wherein the database further comprises:
 a universal data value set comprising all data values from all data value sets; and   an occurrence frequency module to determine a frequency of occurrence for each data value within the universal data value set and to assign a similarity weight to each data value based on the occurrence frequency associated with the data value such that the similarity weight is inversely proportional to the occurrence frequency.   
     
     
         17 . The system of  claim 14 , wherein the candidate pairs comprise a subset of all pairs of data elements. 
     
     
         18 . The system of  claim 14 , wherein the data element pairing module further comprises a signature hashing module comprising at least one signature hash function configured to map each signature into one or more of a plurality of categories to generate a categorical signature for each semantic signature associated with each data element and to compare categorical signatures to identify pairs of semantic signatures mapping to common categories, the identified candidate pairs of data elements associated with the identified pairs of semantic signatures. 
     
     
         19 . The system of  claim 18 , wherein the signature hash function comprises a locality sensitive hash function. 
     
     
         20 . The system of  claim 20 , wherein the pre-determined measure of similarity comprises jaccard similarity, cosine similarity or combinations thereof.

Join the waitlist — get patent alerts

Track US2013332467A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.