US2016232232A1PendingUtilityA1

Mining product aspects from opinion text

Assignee: IBMPriority: Jun 26, 2014Filed: Apr 28, 2016Published: Aug 11, 2016
Est. expiryJun 26, 2034(~7.9 yrs left)· nominal 20-yr term from priority
G06F 40/211G06F 16/335G06F 40/284G06F 16/345G06F 17/271G06F 17/30719G06F 17/30699G06F 17/277
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A text stream having one or more sentences is received, and any number of the one or more sentences are parsed to determine corresponding subject-verb-object (SVO) triples. Each sentence whose corresponding SVO triple contains an identified verb is selected, based on the identified verb, or a lemma of the identified verb, matching a predefined verb. A subject of each selected sentence is identified as an aspect candidate. Each identified aspect candidate is tokenized and normalized. One or more n-grams are generated for each tokenized and normalized aspect candidate. For each generated n-gram, a frequency at which the n-gram is generated is determined. A number of the generated n-grams are selected as aspects based on the frequency with which the number of n-grams are generated.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer implemented method for extracting aspects from a text stream, comprising:
 receiving a text stream having one or more sentences;   parsing one or more of the sentences to determine corresponding subject-verb-object (SVO) triples;   selecting each sentence having a corresponding SVO triple that contains a predefined verb or lemma;   identifying a subject of each selected sentence as an aspect candidate;   tokenizing and normalizing each identified aspect candidate, wherein the normalizing includes lemmatization;   generating one or more n-grams for each tokenized and normalized aspect candidate;   filtering the one or more n-grams to exclude at least one n-gram, based on a filtering criteria;   determining, for each generated n-gram, a frequency at which the n-gram is generated for the tokenized and normalized aspect candidates;   selecting one or more of the generated n-grams as aspects based on the frequency with which the one or more n-grams, respectively, are generated; and   presenting a user with a graphical representation of a sentiment associated with the selected one or more generated n-grams.

Join the waitlist — get patent alerts

Track US2016232232A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.