US10061845B2ActiveUtilityA1

Analysis of unstructured computer text to generate themes and determine sentiment

Assignee: FMR LLCPriority: Feb 18, 2016Filed: Feb 18, 2016Granted: Aug 28, 2018
Est. expiryFeb 18, 2036(~9.6 yrs left)· nominal 20-yr term from priority
G06F 16/3344G06F 16/9535G06F 17/30684G06F 17/30867
50
PatentIndex Score
1
Cited by
16
References
8
Claims

Abstract

Methods and apparatuses are described for analyzing unstructured computer text for theme generation to determine sentiment. A computer store stores unstructured text that is delimited, a searched phrases log, and a phrase click log. A computer server extracts phrases from the unstructured delimited text by splitting each line of the unstructured delimited text into one or more phrases. The computer server generates tokens from the unstructured delimited text, where the tokens comprise segments of the unstructured delimited text. The computer server determines one or more themes present in the unstructured delimited text.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A system used in a computing environment in which unstructured computer text is analyzed for theme generation to determine sentiment, the system comprising:
 a computer store including unstructured text that is delimited, a searched phrases log, and a phrase click log,
 the unstructured text being input via a web page, input directly into the computer store via a first computer file, or any combination thereof, 
 the searched phrases log comprising a) a unique set of phrases from all phrases that are searched on one or more specified websites for a specified duration and b) a search phrase frequency count indicating the number of times of the phrases within the unique set of phrases was searched on the one or more specified websites over the specified duration, the searched phrases log being i) retrieved from the internet and stored in the computer store via a second computer file, ii) input directly into the computer store via a third computer file, or any combination of i) and ii), and 
 the phrase click log comprising a) a unique set of phrases from all phrases that correspond to one or more Uniform Resource Locators (URLs) that are activated for a specified duration and b) a clicked-on frequency count indicating the number of times of the URLs associated with the phrases in the unique set of phrases are activated over the specified duration, the phrase click log is i) retrieved from the internet and stored in the computer store via a fourth computer file, ii) input via a webpage, input directly into the computer store via a fifth computer file, or any combination of i) and ii), and 
 
 a computer server in communication with the computer store and programmed to:
 extract phrases from the unstructured delimited text by splitting each line of the unstructured delimited text into one or more phrases; 
 generate tokens from the unstructured delimited text, wherein the tokens comprise segments of the unstructured delimited text; 
 determine one or more themes present in the unstructured delimited text by:
 a) identifying candidate phrases based on the tokens and the searched phrases log, 
 b) ranking each of the candidate phrases based on frequency of a respective phrase in the unstructured delimited text, the search phrase frequency count corresponding to the respective phrase, and the clicked-on frequency count corresponding to the respective phrase, 
 c) selecting a subset of phrases from the candidate phrases, where the subset of phrases selected is based on the respective rank of each phrase, 
 d) grouping the subset of phrases based on words in the subset of phrases, and 
 e) determining themes for each grouping based on the number of times a phrase appears in the group. 
 
 
 
     
     
       2. The system of  claim 1  wherein grouping the subset of phrases is based on trigrams in the subset of phrases that substantially match. 
     
     
       3. The system of  claim 1  wherein grouping the subset of phrases is based on a co-occurrence of matching words in the subset of phrases. 
     
     
       4. The system of  claim 1  wherein determining themes further comprises picking the phrase that appears in the group a highest number of times. 
     
     
       5. A computerized method for analyzing unstructured computer text for theme generation to determine sentiment, the method comprising:
 storing, by a computer store, unstructured text that is delimited, a searched phrases log, and a phrase click log,
 wherein the unstructured text is input via a web page, input directly into the computer store via a first computer file, or any combination thereof, 
 wherein the searched phrases log comprises a) a unique set of phrases from all phrases that are searched on one or more specified websites for a specified duration and b) a search phrase frequency count indicating the number of times of the phrases within the unique set of phrases was searched on the one or more specified websites over the specified duration, the searched phrases log being i) retrieved from the internet and stored in the computer store via a second computer file, ii) input directly into the computer store via a third computer file, or any combination of i) and ii), and 
 wherein the phrase click log comprises a) a unique set of phrases from all phrases that correspond to one or more Uniform Resource Locators (URLs) that are activated for a specified duration and b) a clicked-on frequency count indicating the number of times of the URLs associated with the phrases in the unique set of phrases are activated over the specified duration, the phrase click log being i) retrieved from the internet and stored in the computer store via a fourth computer file, ii) input via a webpage, input directly into the computer store via a fifth computer file, or any combination of i) and ii); 
 
 extracting, by a computer server in communication with the computer store, phrases from the unstructured delimited text by splitting each line of the unstructured delimited text into one or more phrases; 
 generating, by the computer server, tokens from the unstructured delimited text, wherein the tokens comprise segments of the unstructured delimited text; and 
 determining, by the computer server, one or more themes present in the unstructured delimited text by:
 a) identifying candidate phrases based on the tokens and the searched phrases log, 
 b) ranking each of the candidate phrases based on frequency of a respective phrase in the unstructured delimited text, the search phrase frequency count corresponding to the respective phrase, and the clicked-on frequency count corresponding to the respective phrase, 
 c) selecting a subset of phrases from the candidate phrases, where the subset of phrases selected is based on the respective rank of each phrase, 
 d) grouping the subset of phrases based on words in the subset of phrases, and 
 e) determining themes for each grouping based on the number of times a phrase appears in the group. 
 
 
     
     
       6. The method of  claim 5  wherein grouping the subset of phrases is based on trigrams in the subset of phrases that substantially match. 
     
     
       7. The method of  claim 5  wherein grouping the subset of phrases is based on a co-occurrence of matching words in the subset of phrases. 
     
     
       8. The method of  claim 5  wherein the step of determining themes further comprises picking the phrase that appears in the group a highest number of times.

Join the waitlist — get patent alerts

Track US10061845B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.