US2002188681A1PendingUtilityA1

Method and system for informing users of subjects of discussion in on -line chats

Priority: Aug 28, 1998Filed: Mar 18, 2002Published: Dec 12, 2002
Est. expiryAug 28, 2018(expired)· nominal 20-yr term from priority
H04L 12/1831H04M 3/42221G10L 15/1815H04M 3/563H04M 2203/4536H04L 12/1827G06Q 10/107H04M 3/56H04M 2201/40
40
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method for informing a user of topics of discussion in a recorded chat between two or more people is described. The method includes the steps of identifying elements from the chat having similar content, labeling some or the identified elements as topics, and presenting the topics to the user. Identifying elements from the chat having similar content includes decomposing the chat into utterances made by the people involved in the chat and clustering the utterances using document clustering techniques on each utterance to identify elements in the utterances having similar content.

Claims

exact text as granted — not AI-modified
What is claimed is:  
     
         1 . A method for informing a user of topics of discussion in a recorded chat between two or more people, the method comprising: 
 identifying elements from the chat having similar content;    labeling some or all of the identified elements as topics; and    presenting the topics to the user.    
     
     
         2 . The method of  claim 1 , wherein the step of identifying elements from the chat having similar content comprises decomposing the chat into a plurality of utterances made by the people involved in the chat and clustering the utterances to identify elements in the utterances having similar content.  
     
     
         3 . The method of  claim 2 , wherein the step of identifying elements from the chat having similar content comprises parsing each decomposed utterances into one or more tokens and representing each utterance as a vector comprising a combination of some or all of the one or more tokens.  
     
     
         4 . The method of  claim 3 , wherein the step of representing each utterance as a vector comprises removing some of the tokens in the utterance before representing the utterance as a vector.  
     
     
         5 . The method of  claim 4 , wherein the chat is not ongoing, and wherein the step of removing some tokens comprises removing tokens appearing in a percentage of all utterances in the chat which is below a first percentage or above a second percentage.  
     
     
         6 . The method of  claim 3 , wherein the chat is ongoing, and wherein the step of representing each utterance as a vector comprises representing all tokens in the utterance in the vector.  
     
     
         7 . The method of  claim 3 , wherein the step of representing each utterance as a vector comprises weighting each token in the vector.  
     
     
         8 . The method of  claim 7 , wherein the step of weighting each token comprises computing the weight of a each token as the frequency of occurrence of the token in the utterance divided by the largest frequency of occurrence for any token in the utterance.  
     
     
         9 . The method of  claim 7 , wherein the step of weighting each token comprises computing the weight of each token as the frequency.  
     
     
         10 . The method of  claim 7 , comprising normalizing each vector.  
     
     
         11 . The method of  claim 3 , comprising generating a vector space model comprising a matrix having a plurality of rows and a plurality of columns, wherein the number of rows equals the number of utterances represented by vectors and the number of columns equals the number of tokens contained in the vectors.  
     
     
         12 . The method of  claim 1 , wherein the step of presenting the topics comprises hyperlinking the topics to documents containing utterances having the respective identified elements.  
     
     
         13 . The method of  claim 12 , wherein the step of presenting the topics comprises hyperlinking each utterance in the documents to a location in the chat in which the respective utterance appears.  
     
     
         14 . The method of  claim 1 , wherein the step of labeling comprises selecting some of the topics according to a predefined criteria.  
     
     
         15 . The method of  claim 14 , wherein the step of selecting some of the topics comprises identifying topics which are nouns or noun phrases and selecting the topics so identified.  
     
     
         16 . The method of  claim 1 , wherein the chat is ongoing, and wherein the step of identifying elements from the chat having similar content comprises: 
 receiving a first set of ongoing chat data from the ongoing chat;    decomposing the first set of ongoing chat data into a plurality of first utterances;    when a first number of first utterances has been received, clustering the first utterances to generate a plurality of first clusters;    receiving a second set of ongoing chat data from the ongoing chat after the first set of ongoing chat data;    decomposing the second set of ongoing chat data into a plurality of second utterances; and    when a second number of second utterances has been received, clustering the second utterances into the plurality of first clusters.    
     
     
         17 . A method for clustering an ongoing chat, the method comprising: 
 receiving ongoing chat data;    decomposing a first set of ongoing chat data into a plurality of first utterances;    when a first number of first utterances has been received, clustering the first utterances to generate a plurality of first clusters;    decomposing a second set of ongoing chat data received after the first set of ongoing chat data into a plurality of second utterances; and    when a second number of second utterances has been received, clustering the second utterances into the plurality of first clusters.    
     
     
         18 . The method of  claim 17 , comprising parsing each of the first and second utterances into tokens.  
     
     
         19 . The method of  claim 18 , comprising representing each utterance as a vector comprising a combination of the tokens in the utterance.  
     
     
         20 . The method of  claim 19 , wherein the step of representing each utterance as a vector comprises weighting each token in the vector.  
     
     
         21 . The method of  claim 20 , wherein the step of weighting each token comprises computing the weight of each token as the frequency of occurrence of the token in the utterance divided by the largest frequency of occurrence for any token in the utterance.  
     
     
         22 . The method of  claim 17 , comprising generating a new cluster after clustering the second utterances into the first clusters.  
     
     
         23 . The method of  claim 22 , wherein the step of generating the new cluster comprises identifying a given cluster as larger than all other clusters and selectively breaking the largest cluster into two or more smaller clusters.  
     
     
         24 . The method of  claim 23 , identifying a cluster having a centroid which is further from the largest cluster than all other clusters, and wherein the step of breaking the largest cluster into two or more clusters is performed if the largest cluster contains a number of utterances greater than a predefined number and the distance of the largest cluster from the centroid of the furthest cluster exceeds a predefined distance.  
     
     
         25 . The method of  claim 23 , wherein the step of breaking the largest cluster comprising breaking the largest cluster using a k-means clustering technique.  
     
     
         26 . A method for identifying elements from a chat having similar content, the method comprising 
 decomposing the chat into a plurality of utterances made by the people involved in the chat    parsing each decomposed utterance into one or more tokens;    representing each utterance as a vector comprising a combination of some or all of the one or more tokens; and    clustering the utterances using the vectors to identify elements in the utterances having similar content.    
     
     
         27 . An article of manufacture comprising a computer readable medium containing a program which when executed on a computer causes the computer to perform a method for informing a user of topics of discussion in a recorded chat between two or more people, the method comprising: 
 identifying elements from the chat having similar content;    labeling some or all of the identified elements as topics; and    presenting the topics to the user.

Join the waitlist — get patent alerts

Track US2002188681A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.