Disambiguation for Preprocessing Content to Determine Relationships
Abstract
Relationships are determined by disambiguation for preprocessing content. A first canonical identifier associated with a first element that can be represented in content in a plurality of forms is defined. A second canonical identifier associated with a second element that can be represented in content in a plurality of forms is defined. A first content available over a network is retrieved. An entity name element associated with the first content is identified. The entity name element being able to represent the first element and the second element. The entity name element is associated with the first element or the second element based on context associated with the first content.
Claims
exact text as granted — not AI-modified1 . A method of disambiguation for preprocessing content to determine relationships comprising:
defining a first canonical identifier associated with a first element that can be represented in content in a plurality of forms; defining a second canonical identifier associated with a second element that can be represented in content in a plurality of forms; retrieving a first content available over a network; identifying an entity name element associated with the first content, the entity name element being able to represent the first element and the second element; and associating the entity name element with the first element or the second element based on context associated with the first content.
2 . The method of claim 1 wherein the context associated with the first content comprises an overall category of content typically served from a content provider providing the first content.
3 . The method of claim 1 wherein the context associated with the first content comprises an URL associated with the first content.
4 . The method of claim 1 wherein the context associated with the first content comprises localized usage of the entity name element associated with the content provider providing the first content.
5 . The method of claim 1 wherein the context associated with the first content comprises a rule from a rule database defining a chosen association between the entity name element and the first element or the second element.
6 . The method of claim 1 wherein the context associated with the first content comprises:
identifying one or more additional entity name elements associated with the first content; and determining whether the entity name element and the one or more additional entity name elements co-occurred more often with the first element or the second element.
7 . The method of claim 6 further comprising determining co-occurrence based on tables in a database.
8 . The method of claim 6 further comprising determining co-occurrence based on a frequency of two elements occurring with each other.
9 . The method of claim 1 wherein the context associated with the first content comprises:
displaying the first element and the second element to a user; receiving a response indicating an action by the user; and determining if the entity name element is more likely associated with the first element or the second element based on the response.
10 . The method of claim 9 wherein displaying comprises displaying the first element and the second element in a did-you-mean area.
11 . The method of claim 9 wherein displaying comprises displaying the first element and the second element as links.
12 . The method of claim 11 wherein the action by the user comprises selecting one of the links.
13 . The method of claim 1 wherein the context associated with the first content comprises:
identifying one or more first-type elements associated with the first content using a rule-based algorithm, the one or more first-type elements being selected from a plurality of predefined elements associated with a topic, industry, or any combination thereof; assigning a corresponding score to the one or more first-type elements based on relevancy; identifying a top scored first-type element from the one or more first-type elements; and determining if the top scored first-type element is more likely associated with the first element or the second element.
14 . The method of claim 1 , wherein the first content comprises an electronic document associated with the content provider's web site, a syndicated news feed, an electronic document associated with a third-party web site, an electronic document associated with a weblog, or any combination thereof.
15 . A system for disambiguation for preprocessing content to determine relationships comprising one or more computing devices configured to:
define a first canonical identifier associated with a first element that can be represented in content in a plurality of forms; define a second canonical identifier associated with a second element that can be represented in content in a plurality of forms; retrieve a first content available over a network; identify an entity name element associated with the first content, the entity name element being able to represent the first element and the second element; and associate the entity name element with the first element or the second element based on context associated with the first content.
16 . A computer program product, tangibly embodied in an information carrier, the computer program product including instructions being operable to cause a data processing apparatus to:
define a first canonical identifier associated with a first element that can be represented in content in a plurality of forms; define a second canonical identifier associated with a second element that can be represented in content in a plurality of forms; retrieve a first content available over a network; identify an entity name element associated with the first content, the entity name element being able to represent the first element and the second element; and associate the entity name element with the first element or the second element based on context associated with the first content.Join the waitlist — get patent alerts
Track US2007150721A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.