US2009112588A1PendingUtilityA1

Method for segmenting communication transcripts using unsupervsed and semi-supervised techniques

Assignee: IBMPriority: Oct 31, 2007Filed: Oct 31, 2007Published: Apr 30, 2009
Est. expiryOct 31, 2027(~1.3 yrs left)· nominal 20-yr term from priority
G06F 16/355G10L 15/04
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method is provided for forming discrete segment clusters of one or more sequential sentences from a corpus of communication transcripts of transactional communications that comprises dividing the communication transcripts of the corpus into a first set of sentences spoken by a caller and a second set of sentences spoken by a responder; generating a specified number of sentence clusters by grouping the first and second sets of sentences according to a measure of lexical similarity using an unsupervised partitional clustering method; generating a collection of sequences of sentence types by assigning a distinct sentence type to each sentence cluster and representing each sentence of each communication transcript of the corpus with the sentence type assigned to the sentence cluster into which the sentence is grouped; and generating a specified number of discrete segment clusters by successively merging sentence clusters according to a proximity-based measure between the sentence types assigned to the sentence clusters within sequences of the collection.

Claims

exact text as granted — not AI-modified
1 - 14 . (canceled) 
     
     
         15 . A method for forming discrete segment clusters of one or more sequential sentences from a corpus of communication transcripts of transactional communications, each communication transcript including a sequence of sentences spoken between a caller and a responder, the method comprising:
 dividing the communication transcripts of the corpus into a first set of sentences spoken by the caller and a second set of sentences spoken by the responder;   grouping the first and second sets of sentences into a set of sentence clusters using a K-means algorithm that is performed adaptively until a quality measure is optimized, the quality measure being calculated by first determining a normalized entropy value for each communication transcript in the corpus with respect to the set of sentence clusters, and then determining a cardinality-weighted average of the normalized entropy values for every communication transcript in the corpus with respect to the set of sentence clusters;   assigning a distinct sentence type to each sentence cluster of the set of sentence clusters;   representing each sentence of each communication transcript of the corpus with the sentence type assigned to the sentence cluster into which the sentence is grouped to generate a collection of sequences of sentence types;   performing agglomerative hierarchical clustering to successively merge pairs of sentence clusters according to a proximity-based measure between pairs of sentence clusters to generate a specified number of discrete segment clusters of one or more sequential sentences, the proximity-based measure between any pair of sentence clusters being proportional to a frequency of co-occurrence of the pair of sentence types assigned to the pair of sentence clusters in a certain neighborhood of the sequences of sentence types in the collection;   obtaining a distinct predetermined collection of key phrases for each of one or more segment types;   assigning each discrete segment cluster of the specified number of discrete segment clusters for which most of the one or more sequential sentences of the discrete segment cluster are within the collection of key phrases for one segment type of the one or more segment types to the segment type;   removing each discrete segment cluster from the specified number of discrete segment clusters for which most of the one or more sequential sentences of the discrete segment cluster are not within the collection of key phrases for any of the one or more segment types; and   merging any discrete segment clusters of the specified number of discrete segment clusters that are assigned to the same segment type of the one or more segment types.

Join the waitlist — get patent alerts

Track US2009112588A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.