US2018232623A1PendingUtilityA1

Techniques for answering questions based on semantic distances between subjects

Assignee: IBMPriority: Feb 10, 2017Filed: Feb 10, 2017Published: Aug 16, 2018
Est. expiryFeb 10, 2037(~10.5 yrs left)· nominal 20-yr term from priority
G06N 7/01G06N 5/04G06F 16/3329G06F 40/30G06F 17/18G06N 3/006
49
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A technique for answering questions includes receiving a question directed to a first subject. A mathematical operation is performed between each of one or more first topic vectors (associated with the first subject) and each of one or more second topic vectors (associated with a second subject) to generate respective strength values. Relevant ones of the respective strength values are summed to provide an overall strength value, which is utilized to determine a semantic distance (SD) between the first subject and the second subject. In response to the SD being within a threshold distance value (TDV), information associated with the first subject and the second subject is utilized to answer the question. In response to the SD not being within the TDV, information associated with the first subject is utilized to answer the question.

Claims

exact text as granted — not AI-modified
1 - 8 . (canceled) 
     
     
         9 . A computer program product configured to answer a question, the computer program product comprising:
 a computer-readable storage device; and   computer-readable program code embodied on the computer-readable storage device, wherein the computer-readable program code, when executed by a data processing system, causes the data processing system to:
 receive a question directed to a first subject; 
 perform a mathematical operation between each of one or more first topic vectors and each of one or more second topic vectors to generate respective strength values, wherein the first topic vectors are associated with respective first topics of the first subject, the second topic vectors are associated with respective second topics of a second subject, and the respective strength values indicate a relative closeness between associated ones of the first and second topics; 
 sum relevant ones of the respective strength values to provide an overall strength value between the first subject and the second subject; 
 determine a semantic distance between the first subject and the second subject based on the overall strength value; 
 in response to the semantic distance being within a threshold distance value, utilize information associated with the first subject and the second subject to answer the question; and 
 in response to the semantic distance not being within the threshold distance value, utilizing information associated with the first subject to answer the question. 
   
     
     
         10 . The computer program product of  claim 9 , wherein the computer- readable program code, when executed by the data processing system, further causes the data processing system to:
 generate the first topic vectors and the second topic vectors based on a statistical model analysis.   
     
     
         11 . The computer program product of  claim 10 , wherein the mathematical operation is a dot product operation. 
     
     
         12 . The computer program product of  claim 10 , wherein the statistical model analysis is a latent Dirichlet allocation (LDA) analysis. 
     
     
         13 . The computer program product of  claim 9 , wherein a number of the first topics for the first subject is determined by taking the square root of a number of first documents associated with the first subject divided by two. 
     
     
         14 . The computer program product of  claim 13 , wherein the first documents are located using a Web search. 
     
     
         15 . The computer program product of  claim 9 , wherein the semantic distance between the first subject and the second subject is determined by taking an inverse of the overall strength value. 
     
     
         16 . The computer program product of  claim 9 , wherein each of the first and second topic vectors have an associated word and an associated strength that is normalized to one. 
     
     
         17 . A data processing system, comprising:
 a cache memory; and   a processor coupled to the cache memory, wherein the processor is configured to:
 receive a question directed to a first subject; 
 perform a mathematical operation between each of one or more first topic vectors and each of one or more second topic vectors to generate respective strength values, wherein the first topic vectors are associated with respective first topics of the first subject, the second topic vectors are associated with respective second topics of a second subject, and the respective strength values indicate a relative closeness between associated ones of the first and second topics; 
 sum relevant ones of the respective strength values to provide an overall strength value between the first subject and the second subject; 
 determine a semantic distance between the first subject and the second subject based on the overall strength value; 
 in response to the semantic distance being within a threshold distance value, utilize information associated with the first subject and the second subject to answer the question; and 
 in response to the semantic distance not being within the threshold distance value, utilizing information associated with the first subject to answer the question. 
   
     
     
         18 . The data processing system of  claim 17 , wherein the processor is further configured to:
 generate the first topic vectors and the second topic vectors based on a statistical model analysis.   
     
     
         19 . The data processing system of  claim 18 , wherein the mathematical operation is a dot product operation, the statistical model analysis is a latent Dirichlet allocation (LDA) analysis, and a number of the first topics for the first subject is determined by taking the square root of a number of first documents associated with the first subject divided by two. 
     
     
         20 . The data processing system of  claim 17 , wherein the semantic distance between the first subject and the second subject is determined by taking an inverse of the overall strength value.

Join the waitlist — get patent alerts

Track US2018232623A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.