US2015310099A1PendingUtilityA1

System And Method For Generating Labels To Characterize Message Content

Assignee: PALO ALTO RES CT INCPriority: Nov 6, 2012Filed: Mar 16, 2015Published: Oct 29, 2015
Est. expiryNov 6, 2032(~6.3 yrs left)· nominal 20-yr term from priority
G06F 16/353G06F 40/30G06F 17/30707G06F 17/2785
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for generating labels to characterize message content are provided. At least one component, associated with a document, is extracted from a message. Words regarding the extracted component are extracted from the message as candidate labels. Those candidate labels that are discriminative of the document associated with the extracted component are identified by comparing the candidate labels for the component with other candidate labels extracted from other messages with at least one of a same and a different component. Content of the message is characterized using the discriminative candidate labels.

Claims

exact text as granted — not AI-modified
1 . A system for generating labels to characterize message content, comprising:
 a component extraction module to extract at least one component from a message, wherein the component is associated with a document;   a term extraction module to extract from the message, words regarding the extracted component as candidate labels;   a label determination module to identify those candidate labels that are discriminative of the document associated with the extracted component by comparing the candidate labels for the component with other candidate labels extracted from other messages with at least one of a same and a different component; and   a characterization module to characterize content of the message using the discriminative candidate labels.   
     
     
         2 . A system according to  claim 1 , further comprising:
 an assignment module to assign a relevance value to each of the candidate labels extracted from the message and the other messages, wherein the relevance value comprises a measure of relevance of one such candidate label to the document.   
     
     
         3 . A system according to  claim 2 , further comprising:
 a message vector module to generate for the message and each of the other messages, a vector comprising the candidate labels and the relevance values for the candidate labels.   
     
     
         4 . A system according to  claim 3 , further comprising:
 a similarity module to determine a local similarity between the message and at least one of the other messages that includes the same component by comparing the vector of the message with the vector of the at least one other message.   
     
     
         5 . A system according to  claim 3 , further comprising:
 a component vector module to generate a vector for the document of the component by combining the message with the other messages that share the same component as the message, by identifying the candidate labels within the message and the other messages that share the same component, and determining a relevance value for each of the identified candidate labels.   
     
     
         6 . A system according to  claim 5 , further comprising:
 a similarity module to determine at least one of a global similarity and a global dissimilarity by comparing the document vector with other vectors for other documents referenced by one or more of the other messages.   
     
     
         7 . A system according to  claim 1 , further comprising:
 a label identification module to determine the discriminative candidate labels as those candidate labels that occur in many of the other messages having the same component and fail to occur in many of the other messages with the different components.   
     
     
         8 . A system according to  claim 1 , further comprising:
 a variant module to determine variants for one or more of the candidate labels;   a comparison module to compare the candidate labels and the variants of the message with candidate labels extracted from the document and variants for the candidate labels extracted from the document;   an information identification module to identify at least one candidate label or variant of the message that is not included in the document as new information that is an opinion regarding the document; and   a message classification module to classify the message as an opinion message.   
     
     
         9 . A system according to  claim 10 , further comprising:
 a characterization module to characterize content of the opinion message as one of positive or negative, comprising:
 lists of predetermined positive words and negative words; 
 an application module to apply the lists to the opinion message; 
 a word identification module to identify words in the opinion message as one of positive or negative; and 
 a word classification module to classify the opinion message as one of a positive message, a negative message, and a ratio of positive and negative words. 
   
     
     
         10 . A system according to  claim 1 , further comprising:
 a variant module to determine variants for one or more of the candidate labels;   a comparison module to compare the candidate labels and the variants of the message with candidate labels extracted from the document and variants for the candidate labels extracted from the document;   a match determination module to determine that the candidate labels and the variants of the message match the candidate labels and the variants of the document; and   a message classification module to classify the message as a descriptive message.   
     
     
         11 . A system according to  claim 1 , further comprising:
 a topic module to identify topics based on the discriminatory candidate labels; and   a cluster module to cluster the message with the other messages based on a similarity of the topics.   
     
     
         12 . A method for generating labels to characterize message content, comprising:
 extracting at least one component from a message, wherein the component is associated with a document;   extracting from the message, words regarding the extracted component as candidate labels;   identifying those candidate labels that are discriminative of the document associated with the extracted component by comparing the candidate labels for the component with other candidate labels extracted from other messages with at least one of a same and a different component; and   characterizing content of the message using the discriminative candidate labels.   
     
     
         13 . A method according to  claim 12 , further comprising:
 assigning a relevance value to each of the candidate labels extracted from the message and the other messages, wherein the relevance value comprises a measure of relevance of one such candidate label.   
     
     
         14 . A method according to  claim 13 , further comprising:
 generating for the message and each of the other messages, a vector comprising the candidate labels and the relevance values for the candidate labels.   
     
     
         15 . A method according to  claim 14 , further comprising:
 determining a local similarity between the message and at least one of the other messages that includes the same component by comparing the vector of the message with the vector of the at least one other message.   
     
     
         16 . A method according to  claim 14 , further comprising:
 generating a vector for the document of the component, comprising:
 combining the message with the other messages that share the same component as the message; 
 identifying the candidate labels within the message and the other messages that share the same component; and 
 determining a relevance value for each of the identified candidate labels. 
   
     
     
         17 . A method according to  claim 16 , further comprising:
 determining at least one of a global similarity and a global dissimilarity by comparing the document vector with other vectors for other documents referenced by one or more of the other messages.   
     
     
         18 . A method according to  claim 12 , further comprising:
 determining the discriminative candidate labels as those candidate labels that occur in many of the other messages having the same component and fail to occur in many of the other messages with the different components.   
     
     
         19 . A method according to  claim 12 , further comprising:
 determining variants for one or more of the candidate labels;   comparing the candidate labels and the variants of the message with candidate labels extracted from the document and variants for the candidate labels extracted from the document;   identifying at least one candidate label or variant of the message that is not included in the document as new information that is an opinion regarding the document; and   classifying the message as an opinion message.   
     
     
         20 . A method according to  claim 19 , further comprising:
 characterizing content of the opinion message as one of positive or negative, comprising:
 obtaining lists of predetermined positive words and negative words; 
 applying the lists to the opinion message; 
 identifying words in the opinion message as one of positive or negative; and 
 classifying the opinion message as one of a positive message, a negative message, and a ratio of positive and negative words. 
   
     
     
         21 . A method according to  claim 12 , further comprising:
 determining variants for one or more of the candidate labels;   comparing the candidate labels and the variants of the message with candidate labels extracted from the document and variants for the candidate labels extracted from the document;   determining that the candidate labels and the variants of the message match the candidate labels and the variants of the document; and   classifying the message as a descriptive message.   
     
     
         22 . A method according to  claim 12 , further comprising:
 identifying topics based on the discriminatory candidate labels; and   clustering the message with the other messages based on a similarity of the topics.

Join the waitlist — get patent alerts

Track US2015310099A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.