Answer retrieval technique
Abstract
A method for analyzing a number of candidate answer texts to determine their respective relevance to a query text, includes the following steps: producing, for respective candidate answer texts being analyzed, respective pluralities of component scores that result from respective comparisons with the query text, the comparisons including a measure of word occurrences, word group occurrences, and word sequences occurrences; determining, for respective candidate answer texts being analyzed, a composite relevance score as a function of the component scores; and outputting at least some of the candidate answer texts having the highest composite relevance scores.
Claims
exact text as granted — not AI-modified1 . A method for analyzing a number of candidate answer texts to determine their respective relevance to a query text, comprising the steps of:
producing, for respective candidate answer texts being analyzed, a word occurrence score that includes a measure of query text words that occur in the candidate answer text; producing, for respective candidate answer texts being analyzed, a word sequence score that includes a measure of query text word sequences that occur in the candidate answer text; and determining, for respective candidate answer texts being analyzed, a composite relevance score as a function of the respective word occurrence score and the respective word sequence score.
2 . The method as defined by claim 1 , further comprising the step of arranging said candidate answer texts in accordance with their composite relevance scores.
3 . The method as defined by claim 1 , wherein said step of producing, for respective candidate texts being analyzed, a word occurrence score includes normalization of the word occurrence score in accordance with the total number of words in the query text.
4 . The method as defined by claim 1 , wherein said query text includes a prime query portion and an explanation portion, and wherein said word occurrence score comprises a weighted sum of prime query portion words that occur in the text and explanation portion words that occur in the text, divided by a weighted sum of the total words in the prime query portion and the total words in the explanation portion.
5 . The method as defined by claim 1 , wherein said query text includes a prime query portion and an explanation portion, and further comprising the step of producing, for said respective answer texts being analyzed, a prime word occurrence score that includes a measure of the number of prime query portion words that occur in the candidate answer text divided by the number of words in the prime query portion; and wherein said composite relevance score, for respective candidate answer texts, is also a function of said prime word occurrence score.
6 . The method as defined by claim 1 , further comprising the steps of: determining the presence at least one of corresponding sequence of a plurality of words in the query text and the respective candidate answer text being analyzed; producing, for the respective candidate answer text being analyzed, a length index score that depends on the respective ratio of minimum to maximum sequence length as between the candidate answer text being analyzed and the query text; and wherein said composite relevance score, for respective candidate answer texts, is also a function of said length index score.
7 . The method as defined by claim 4 , further comprising the steps of: determining the presence at least one of corresponding sequence of a plurality of words in the query text and the respective candidate answer text being analyzed; producing, for the respective candidate answer text being analyzed, a length index score that depends on the respective ratio of minimum to maximum sequence length as between the candidate answer text being analyzed and the query text; and wherein said composite relevance score, for respective candidate answer texts, is also a function of said length index score.
8 . The method as defined by claim 1 , further comprising the steps of: determining the presence at least one of corresponding sequence of a plurality of words in the query text and the respective candidate answer text being analyzed; producing, for the respective candidate answer text being analyzed, a length index that depends on the respective ratio of minimum to maximum sequence length as between the candidate answer text being analyzed and the query text; producing, for the respective candidate answer text being analyzed, an order match index that depends on a summation, over all the corresponding sequences, of the ratio of minimum to maximum distance between words of a sequence; and producing a length and order match score from said length index and said order match index; and wherein said composite relevance score, for respective candidate answer texts, is also a function of said length and order match score.
9 . The method as defined by claim 4 , further comprising the steps of: determining the presence at least one of corresponding sequence of a plurality of words in the query text and the respective candidate answer text being analyzed; producing, for the respective candidate answer text being analyzed, a length index that depends on the respective ratio of minimum to maximum sequence length as between the candidate answer text being analyzed and the query text; producing, for the respective candidate answer text being analyzed, an order match index that depends on a summation, over all the corresponding sequences, of the ratio of minimum to maximum distance between words of a sequence; and producing a length and order match score from said length index and said order match index; and wherein said composite relevance score, for respective candidate answer texts, is also a function of said length and order match score.
10 . The method as defined by claim 8 , wherein said step of producing a length and order match score from said length index and said order match index comprises producing a product of said length index and said order match index.
11 . The method as defined by claim 1 , wherein the components of said composite relevance score are non-linearly processed.
12 . The method as defined by claim 4 , wherein the components of said composite relevance score are non-linearly processed.
13 . The method as defined by claim 10 , wherein the components of said composite relevance score are non-linearly processed.
14 . The method as defined by claim 1 , further comprising the step of outputting at least some of said candidate answer texts having the highest composite relevance scores.
15 . The method as defined by claim 2 , further comprising the step of outputting at least some of said candidate answer texts having the highest composite relevance scores.
16 . The method as defined by claim 4 , further comprising the step of outputting at least some of said candidate answer texts having the highest composite relevance scores
17 . A method for analyzing a number of candidate answer texts to determine their respective relevance to a query text, comprising the steps of:
producing, for respective candidate answer texts being analyzed, a respective pluralities of component scores that result from respective comparisons with said query text, said comparisons including a measure of word occurrences, word group occurrences, and word sequences occurrences; determining, for respective candidate answer texts being analyzed, a composite relevance score as a function of said component scores; and outputting at least some of said candidate answer texts having the highest composite relevance scores.
18 . The method as defined by claim 17 , wherein said composite relevance score is obtained as a weighted sum of non-linear functions of said component scores.
19 . The method as defined by claim 17 , wherein said query text includes a prime query portion and an explanation portion, and wherein at least one of said component scores result from comparison of respective candidate answer texts with the entire query text, and wherein at least one of the said component scores result from comparison of respective candidate answer texts with only the query portion.
20 . The method as defined by claim 18 , wherein said query text includes a prime query portion and an explanation portion, and wherein at least one of said component scores result from comparison of respective candidate answer texts with the entire query text, and wherein at least one of the said component scores result from comparison of respective candidate answer texts with only the query portion.
21 . An answer retrieval method, comprising the steps of:
producing a query text; implementing a search of knowledge sources to obtain a number of candidate answer texts, and determining their respective relevance to the query text, as follows:
producing, for respective candidate answer texts being analyzed, a respective pluralities of component scores that result from respective comparisons with said query text, said comparisons including a measure of word occurrences, word group occurrences, and word sequences occurrences;
determining, for respective candidate answer texts being analyzed, a composite relevance score as a function of said component scores; and
outputting at least some of said candidate answer texts having the highest composite relevance scores.
22 . The method as defined by claim 21 , further comprising the steps of implementing a second search of knowledge sources to obtain different candidate answer texts, and determining the respective relevance of said different candidate answer texts to said query text.
23 . The method as defined by claim 21 , further comprising filtering said query and said candidate answer texts before said determinations or respective relevance.
24 . The method as defined by claim 22 , further comprising filtering said query and said candidate answer texts before said determinations or respective relevance.
25 . The method as defined by claim 21 , wherein said composite relevance score is obtained as a weighted sum of non-linear functions of said component scores.
26 . The method as defined by claim 21 , wherein said query text includes a prime query portion and an explanation portion, and wherein at least one of said component scores result from comparison of respective candidate answer texts with the entire query text, and wherein at least one of the said component scores result from comparison of respective candidate answer texts with only the query portion.Join the waitlist — get patent alerts
Track US2003074353A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.