US2018046716A1PendingUtilityA1

Domain-based ranking in document search

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Mar 3, 2009Filed: Oct 24, 2017Published: Feb 15, 2018
Est. expiryMar 3, 2029(~2.6 yrs left)· nominal 20-yr term from priority
G06F 16/951G06F 16/9532G06F 17/30864
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In one example, documents that are examined by a search process may be scored in a manner that is specific to a domain. A domain may be a substantive area, such as medicine, sports, etc. Different scoring approaches that take aspects of the domain into account may be applied to the documents, thereby producing different scores than might have been produced by a simple comparison of the terms in the query with the terms in the documents. These domain-based approaches may take a query into account in scoring the documents, or may be query-independent. Each approach may be implemented by a scorer. The combined output of the scorers may be used to generate a score for each document. Documents then may be ranked based on the scores, and search results may be provided.

Claims

exact text as granted — not AI-modified
1 . A computer-readable storage device that stores executable instructions that, when executed by a computer, cause the computer to perform operations comprising:
 receiving a query;   calculating scores for a plurality of documents obtained with respect to the received query by comparing terms in the query with terms in the documents;   calling a same first function implemented by each of a plurality of domain-based scorers of different types, to determine, without utilizing one or more documents of the plurality of documents, which of the domain-based scorers will contribute and which will not contribute to scoring of the documents in response to the calculation of the scores for the plurality of documents, wherein the same first function is used to determine whether the received query is too vague and will not be scored or is not too vague and will be scored, each of the domain-based scorers calculating a domain-based score based on features of the documents or of the query that are specific to a substantive field of knowledge after the calculation of the scores for the plurality of documents, the each of the plurality of domain-based scorers implementing its own version of a same second function to calculate the domain-based score of the documents without obtaining the documents again with respect to the received query, wherein the same second function computes domain-based scores, wherein the same second function of each of the plurality of domain-based scorers utilizes the documents which have already received scores based on the terms in the query to calculate the domain-based scores of the documents;   including, on a list, those domain-based scorers that indicate, through the same first function, that they will contribute to scoring of the documents;   computing adjusted scores of the documents by combining the contributions from all of the domain-based scorers;   creating a set of search results based on the adjusted scores of the documents; and   presenting the search results to a user.   
     
     
         2 . The computer-readable storage device of  claim 1 , wherein determining whether the received query is too vague or not too vague is based upon each domain-based scorer using its own set of first criteria for determining whether the received query is too vague or not too vague. 
     
     
         3 . The computer-readable storage device of  claim 1 , further comprising:
 reducing the domain-based scorers of the documents based on at least one document having an amount of concepts relevant to the query in excess of a predefined number.   
     
     
         4 . The computer-readable storage device of  claim 1 , wherein the operations further comprise:
 identifying a set of concepts associated with the terms in the query;   and either:
 decreasing a domain-based score of one or more of the documents based on how many concepts in the set of concepts are not in the one or more of the documents; or 
 decreasing a domain-based score of the documents based on how many concepts in the one or more of the documents are not in the set of concepts. 
   
     
     
         5 . The computer-readable storage device of  claim 1 , further comprising:
 identifying a first set of concepts in the query;   identifying a second set of concepts in a summary of one or more of the documents;   determining that the first set has a defined level of similarity to the second set; and   based on the first set having the defined level of similarity to the second set, increasing a domain-based score of the one or more of the documents.   
     
     
         6 . The computer-readable storage device of  claim 1 , wherein a first one of the domain-based scorers evaluates the documents without regard to the query based on a determination that the query is vague with respect to domain based scoring of the documents for which the scores have been determined. 
     
     
         7 . The computer-readable storage device of  claim 1 , further comprising:
 determining a number of concepts from a domain that appears in the documents;   determining that the number falls within a range; and   modifying a domain-based score of the documents based on the number falls within the range, wherein the domain-based score is increased on determining that the number falls within a first range, the domain-based score remains same on determining that the number falls within a second range, and the domain-based score is decreased on determining that the number falls within a third range.   
     
     
         8 . The computer-readable storage device of  claim 1 , further comprising:
 determining a number of concepts from a domain that appears in one or more of the documents; and   modifying a domain-based score of the one or more of the documents by an amount that is based on the number.   
     
     
         9 . The computer-readable storage device of  claim 1 , further comprising:
 determining that one or more of the documents have a concept, from a domain, that has a level of popularity across a group of documents in the corpus; and   based on the concept being in the one or more of the documents, modifying a domain-based score of the one or more of the documents.   
     
     
         10 . The computer-readable storage device of  claim 1 , wherein the operations further comprises:
 calling the same first function in each of the scorers to determine which of the scorers will form the set of domain-based scorers to produce a domain-based score for one or more of the documents, wherein each of the scorers implements its own version of the function.   
     
     
         11 . A system that responds to a document search request, the system comprising:
 a memory; and   a processor programmed to:
 receive a query; 
 calculate scores for a plurality of documents obtained with respect to the received query by comparing terms in the query with terms in the documents; 
 call a same first function implemented by each of a plurality of domain-based scorers of different types, to determine, without utilizing one or more documents of the plurality of documents, which of the domain-based scorers will contribute and which will not contribute to scoring of the documents in response to the calculation of the scores for the plurality of documents, wherein the same first function is used to determine whether the received query is too vague and will not be scored or is not too vague and will be scored, each of the domain-based scorers calculating a domain-based score based on features of the documents or of the query that are specific to a substantive field of knowledge after the calculation of the scores for the plurality of documents, the each of the plurality of domain-based scorers implementing its own version of a same second function to calculate the domain-based score of the documents without obtaining the documents again with respect to the received query, wherein the same second function computes domain-based scores, wherein the same second function of each of the plurality of domain-based scorers utilizes the documents which have already received scores based on the terms in the query to calculate the domain-based scores of the documents; 
 include, on a list, those domain-based scorers that indicate, through the same first function, that they will contribute to scoring of the documents; 
 compute adjusted scores of the documents by combining the contributions from all of the domain-based scorers; 
 create a set of search results based on the adjusted scores of the documents; and 
 present the search results to a user. 
   
     
     
         12 . The system of  claim 11 , wherein the same second function includes receiving document identifiers to identify the documents in a database and returning scores for the documents and using the returned scores as input into an aggregation formula, wherein each domain-based scorer uses its own set of second criteria within the aggregation formula. 
     
     
         13 . The system of  claim 11 , wherein a first one of the domain-based scorers evaluates the documents without regard to the query based on a determination that the query is vague with respect to domain based scoring of the documents for which the scores have been determined. 
     
     
         14 . The system of  claim 11 , wherein a first one of the plurality of domain-based scorers determines that a number of concepts from a domain that appear in a given one of the documents falls within a range, and contributes to a modification in score of the given one of the documents based on the number falling within the range. 
     
     
         15 . The system of  claim 11 , wherein a first one of the plurality of domain-based scorers determines a number of concepts of a domain that appear in a given one of the documents and contributes to a modification in score of the given one of the documents, the increase being in an amount that is based on the number. 
     
     
         16 . The system of  claim 11 , wherein a first one of the plurality of domain-based scorers determines that a given one of the documents has a concept that has a level of popularity across the documents, and, based on the concept being in the given one of the documents, contributes to a modification in score of the given one of the documents. 
     
     
         17 . The system of  claim 11 , wherein each of the plurality of domain-based scorers exposes a callable function that, when called, provides an explanation describing a reason for which a final score has been assigned to a given one of the documents. 
     
     
         18 . The system of  claim 11 , wherein a domain comprises a medical domain that includes concepts that describe either conditions of a human body or medical treatments of the human body. 
     
     
         19 . A method of responding to a search query, the method comprising:
 using a processor to perform acts comprising:
 receiving a query; 
 calculating scores for a plurality of documents obtained with respect to the received query by comparing terms in the query with terms in the documents; 
 calling a same first function implemented by each of a plurality of domain-based scorers of different types, to determine, without utilizing one or more documents of the plurality of documents, which of said domain-based scorers will contribute and which will not contribute to scoring of said documents in response to the calculation of said scores for the plurality of documents, wherein the same first function is used to determine whether the received query is too vague and will not be scored or is not too vague and will be scored, each of said domain-based scorers calculating a domain-based score based on features of said documents or of said query that are specific to a substantive field of knowledge after the calculation of the scores for the plurality of documents, said each of said plurality of domain-based scorers implementing its own version of a same second function to calculate the domain-based score of said documents without obtaining said documents again with respect to the received query, wherein the same second function computes domain-based scores, wherein said same second function of each of the plurality of domain-based scorers utilizes said documents which have already received scores based on the terms in said query to calculate the domain-based scores of said documents; 
 including, on a list, those domain-based scorers that indicate, through said same first function, that they will contribute to scoring of said documents; 
 computing adjusted scores of said documents by combining the contributions from all of the said domain-based scorers; 
 creating a set of search results based on the adjusted scores of the documents; and 
 presenting the search results to a user. 
   
     
     
         20 . The method of  claim 19 , further comprising using a configurable parameter selected based on a different scoring scheme by those ones of the domain-based scorers that are on the list to adjust the scores, whereby adjusted scores of the documents are created by combining the contributions from all of the domain-based scorers.

Join the waitlist — get patent alerts

Track US2018046716A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.