Entity assessment and ranking
Abstract
General entity retrieval and ranking is described. A first set of documents is retrieved from one or more document repositories based on a query formed according to the topic. The first set of documents is characterized based on its first set of metadata values. One or more candidate entities are identified based on the first set of documents and the original query is thereafter augmented according to a candidate entity. The second set of documents resulting from the augmented query is then characterized in a similar manner. For each candidate entity, the first and second document set characterizations are compared to determine their degree of similarity. Increasingly similar document set characterizations indicates that the candidate entity is increasingly relevant to the original query. Repeating this process for each of the one or more candidate entities can give rise to rankings according to the respective degrees of similarity.
Claims
exact text as granted — not AI-modified1 - 18 . (canceled)
19 . An entity determination method for a query, the method comprising:
determining characterizations of a first set of documents returned as search results for the query based on metadata associated with the first set of documents; determining, by a processing device, candidate entities for the first set of documents; determining a similarity of each candidate entity to the query based on the characterizations of the first set of documents and information associated with each candidate entity.
20 . The method of claim 19 , comprising:
determining a ranked order of the candidate entities according to the similarities determined for the candidate entities; and presenting the ranked order.
21 . The method of claim 19 , wherein the determining of the information associated with each candidate entity comprises:
for each candidate entity,
retrieving a second set of documents based on a query for the candidate entity, wherein the second set of documents have associated metadata, and
determining characterizations of the second set of documents based on the second set of documents and the metadata associated with the second set of documents.
22 . The method of claim 21 , wherein the determining of the similarity of each candidate entity to the query comprises:
comparing the characterizations of the second set of documents to the characterizations of the first set of documents to determine the similarity for each candidate entity.
23 . The method of claim 22 , wherein the characterizations of the first and second set of documents comprise attribute vectors describing each of the documents in the first and second sets.
24 . The method of claim 23 , wherein the comparing of the characterizations of the second set of documents to the characterizations of the first set of documents comprises:
determining distances between the attribute vectors of the documents in the first and second sets; and comparing the distances to determine the candidate entities similar to the query.
25 . The method of claim 19 , comprising:
receiving information identifying a topic; generating the query from the information; and executing the query on documents in the document repository.
26 . The method of claim 25 , wherein the determining of the similarity of each candidate entity to the query comprises determining a similarity of each candidate entity to the topic.
27 . The method of claim 25 , wherein the determining of the candidate entities for the first set of documents comprises determining top K candidate entities having a closest similarity to the topic, wherein K is an integer greater than 1.
28 . The method of claim 27 , wherein the determining of the top K candidate entities comprises:
determining occurrence frequencies for each identified candidate entity in the first set of documents; and selecting the top K candidate entities based on the occurrence frequencies.
29 . A search engine system comprising:
a data repository to store documents; a document search engine executed by at least one processor to execute a query associated with a topic on the data repository; a document characterizer to determine characterizations of a first set of documents returned as search results for the query based on metadata associated with the first set of documents; and an entity determination subsystem to determine candidate entities for the first set of documents and determine a similarity of each candidate entity to the topic of the query based on the characterizations of the first set of documents and information associated with each candidate entity.
30 . The search engine system of claim 29 , wherein the entity determination subsystem is to determine a ranked order of the candidate entities according to the similarities determined for the candidate entities.
31 . The search engine system of claim 29 , comprising:
a user interface to present the ranked order to a user.
32 . The search engine system of claim 29 , wherein the entity determination subsystem is to determine the information associated with each candidate entity by:
for each candidate entity,
retrieving a second set of documents based on a query for the candidate entity, wherein the second set of documents have metadata, and
determining characterizations of the second set of documents based on the second set of documents and the metadata for the second set of documents.
33 . The search engine system of claim 29 , wherein to determine the similarity of each candidate entity to the topic, the entity determination subsystem is to compare the characterizations of the second set of documents to the characterizations of the first set of documents.
34 . The search engine system of claim 33 , wherein the characterizations of the first and second set of documents comprise attribute vectors describing each of the documents in the first and second sets.
35 . The search engine system of claim 34 , wherein to compare the characterizations of the second set of documents to the characterizations of the first set of documents, the entity determination subsystem is to determine distances between the attribute vectors of the documents in the first and second sets.
36 . A system to determine entities associated with a topic based on query search results, the system comprising:
an interface to receive information associated with a topic from a user; an entity ranking subsystem executed by at least one processor to determine characterizations of a first set of documents returned as search results for the query based on metadata associated with the first set of documents, determine candidate entities for the first set of documents, determine a similarity of each candidate entity to the topic of the query based on the characterizations of the first set of documents and information associated with each candidate entity, and determine rankings of the candidate entities based on their similarity to the topic, wherein the entity ranking subsystem is to present the rankings via the interface.
37 . The system of claim 36 , wherein the entity ranking subsystem is to, for each candidate entity,
retrieve a second set of documents based on a query for the candidate entity, wherein the second set of documents have metadata, and determine characterizations of the second set of documents based on the second set of documents and the metadata for the second set of documents.
38 . The system of claim 36 , wherein the entity ranking subsystem is to compare the characterizations of the second set of documents to the characterizations of the first set of documents to determine the similarity for each candidate entity.Join the waitlist — get patent alerts
Track US2014095466A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.