US2014379743A1PendingUtilityA1

Finding and disambiguating references to entities on web pages

Assignee: GOOGLE INCPriority: Oct 20, 2006Filed: Aug 12, 2014Published: Dec 25, 2014
Est. expiryOct 20, 2026(~0.2 yrs left)· nominal 20-yr term from priority
G06F 17/30011G06F 16/93G06F 16/955G06N 5/04G06N 20/00
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for disambiguating references to entities in a document. In one embodiment, an iterative process is used to disambiguate references to entities in documents. An initial model is used to identify documents referring to an entity based on features contained in those documents. The occurrence of various features in these documents is measured. From the number occurrences of features in these documents, a second model is constructed. The second model is used to identify documents referring to the entity based on features contained in the documents. The process can be repeated, iteratively identifying documents referring to the entity and improving subsequent models based on those identifications. Additional features of the entity can be extracted from documents identified as referring to the entity.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for identifying texts referring to an entity, the entity being associated with a first set of features, the method comprising:
 at a computer having one or more processors and memory storing programs for execution by the one or more processors:
 identifying a first set of text as associated with the entity in accordance with a first set of features that are sufficient for identifying a document referring to the entity; 
 identifying a second set of text as associated with the entity in accordance with a second set of features that are sufficient for identifying a document referring to the entity, wherein the second set of feature is distinct from the first set of features; 
 identifying a representative feature associated with the entity, in accordance with the first set of features and the second set of features; 
   wherein the first set of text and the second set of text are identified from a same audio file.   
     
     
         2 . The method of  claim 1 , further comprising: identifying, as associated with the entity, a third set of text distinct from the first set of documents and the second set of documents. 
     
     
         3 . The method of  claim 2 , further comprising: extracting facts from the third set of text and identifying the facts as associated with the entity. 
     
     
         4 . The method of  claim 1 , wherein the first set of features is stored as a set of facts in a fact repository in association with a second object that corresponds to the entity. 
     
     
         5 . The method of  claim 1 , wherein the first set of text is identified using a first model; the second set of text is identified using a second model distinct from the first model. 
     
     
         6 . The method of  claim 5 , wherein the second model is selected in accordance with a number of occurrences of the first set of features in a document. 
     
     
         7 . The method of  claim 1 , wherein the second set of features includes at least one feature not included in the first set of features. 
     
     
         8 . The method of  claim 1 , wherein the first set of features includes at least one feature not included in the second set of features. 
     
     
         9 . The method of  claim 1 , further comprising: storing at least one feature of the second set of features as a fact in the fact repository. 
     
     
         10 . The method of  claim 1 , further comprising: estimating importance of the entity. 
     
     
         11 . The method of  claim 1 , further comprising:
 estimating importance of the entity based on an estimated importance of at least one of the documents in the second set of documents.   
     
     
         12 . The method of  claim 1 , further comprising:
 associating at least one of the documents with the entity.   
     
     
         13 . The method of  claim 1 , wherein identifying a second set of documents and the first set of features comprises estimating a probability that a document of the second set of documents refers to the entity. 
     
     
         14 . The method of  claim 1 , wherein the first set of features comprises at least a first feature and a second feature, and wherein a second model specifies that an occurrence of the first feature is sufficient to identify a document referring to the entity. 
     
     
         15 . The method of  claim 14 , wherein the second model specifies that an occurrence of the second feature is not sufficient to identify a document referring to the entity. 
     
     
         16 . A system for identifying texts referring to an entity, the entity being associated with a first set of features, the system comprising one or more instructions for:
 identifying a first set of text as associated with the entity in accordance with a first set of features that are sufficient for identifying a document referring to the entity;   identifying a second set of text as associated with the entity in accordance with a second set of features that are sufficient for identifying a document referring to the entity, wherein the second set of feature is distinct from the first set of features;
 identifying a representative feature associated with the entity, in accordance with the first set of features and the second set of features; 
   wherein the first set of text and the second set of text are identified from a same audio file.   
     
     
         17 . The system of  claim 16 , wherein the one or more programs further comprise instructions for: identifying, as associated with the entity, a third set of text distinct from the first set of documents and the second set of documents. 
     
     
         18 . The system of  claim 16 , wherein the one or more programs further comprise instructions for: extracting facts from the third set of text and identifying the facts as associated with the entity. 
     
     
         19 . A non-transitory computer readable storage medium storing one or more programs configured for execution by a computer, the one or more programs comprising instructions for:
 identifying a first set of text as associated with the entity in accordance with a first set of features that are sufficient for identifying a document referring to the entity;   identifying a second set of text as associated with the entity in accordance with a second set of features that are sufficient for identifying a document referring to the entity, wherein the second set of feature is distinct from the first set of features;   identifying a representative feature associated with the entity, in accordance with the first set of features and the second set of features;   wherein the first set of text and the second set of text are identified from a same audio file.   
     
     
         20 . The non-transitory computer readable storage medium of  claim 19 , wherein the one or more programs further comprise instructions for: identifying, as associated with the entity, a third set of text distinct from the first set of documents and the second set of documents.

Join the waitlist — get patent alerts

Track US2014379743A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.