US2015066506A1PendingUtilityA1

System and Method of Text Zoning

Assignee: VERINT SYSTEMS LTDPriority: Aug 30, 2013Filed: Aug 25, 2014Published: Mar 5, 2015
Est. expiryAug 30, 2033(~7.1 yrs left)· nominal 20-yr term from priority
G10L 15/26G10L 15/04G10L 15/18G10L 15/1822
52
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method of zoning a transcription of audio data includes separating the transcription of audio data into a plurality of utterances. A that each word in an utterances is a meaning unit boundary is calculated. The utterance is split into two new utterances at a work with a maximum calculated probability. At least one of the two new utterances that is shorter than a maximum utterance threshold is identified as a meaning unit.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of zoning a transcription of audio data, the method comprising:
 separating the transcription of audio data into a plurality of utterances;   identifying utterances of the plurality of utterances that are shorter than a predetermined minimum threshold as meaning units;   calculating a probability that each word in an utterance of the plurality of utterances which is longer than the predetermined minimum threshold is a meaning unit boundary;   splitting the utterance longer than the predetermined minimum threshold into two new utterances at a word with a maximum calculated probability; and   identifying at least one of the two utterances that is shorter than a maximum utterance threshold as a meaning unit.   
     
     
         2 . The method of  claim 1 , wherein calculating the probability that each word in the utterance longer than the predetermined minimum threshold is a meaning unit boundary is further based upon at least a linguistic model. 
     
     
         3 . The method of  claim 2 , wherein the linguistic model comprises statistics, distributions, or frequencies of word pairs or word triplets. 
     
     
         4 . The method of  claim 2 , wherein the linguistic model comprises probability of words to form the beginning or end of a meaning unit. 
     
     
         5 . The method of  claim 2  wherein calculating the probability that each word in the utterance longer than the predetermined minimum threshold is a meaning unit boundary is further based upon an acoustic model. 
     
     
         6 . The method of  claim 2 , further comprising receiving audio data and decoding the audio data to create the transcription of audio data. 
     
     
         7 . The method of  claim 5 , wherein at least the linguistic model is used when decoding the audio data to create the transcription of audio data. 
     
     
         8 . The method of  claim 1 , further comprising applying speech analytics to the meaning unit to identify at least one of context or content of the meaning unit. 
     
     
         9 . The method of  claim 1 , further comprising applying speech analytics to identified meaning units to group the meaning units into call segments. 
     
     
         10 . The method of  claim 9 , further comprising applying speech analytics to the identified meaning units to identify dialog acts within the identified meaning units. 
     
     
         11 . The method of  claim 1 , wherein the predetermined minimum threshold is thirty words. 
     
     
         12 . The method of  claim 1 , further comprising:
 selecting utterances of the plurality that are longer than the predetermined minimum threshold for subdivision; and   splitting the selected utterances of the plurality into widows, each window being twice the maximum utterance threshold.   
     
     
         13 . The method of  claim 12 , wherein calculating the probability that each word in an utterance longer than the predetermined minimum threshold is a meaning unit boundary is calculated for each word in each window. 
     
     
         14 . The method of  claim 13 , further comprising applying at least one of a linguistic exception and an acoustic exception to the two new utterances. 
     
     
         15 . The method of  claim 14 , wherein the at least one linguistic exception comprises a minimum meaning unit boundary probability or a minimum meaning unit boundary probability differential. 
     
     
         16 . The method of  claim 14 , wherein the at least one acoustic exception comprises an identification of a pause between adjacent utterances in the transcription of the audio data. 
     
     
         17 . A method of zoning, a transcription of audio data, the method comprising:
 separating the transcription of audio data into a plurality of utterances   identifying utterances of the plurality of utterances that are shorter than a predetermined minimum threshold as meaning units;   selecting utterances of the plurality of utterances that are longer than the predetermined minimum threshold for subdivision;   splitting the selected utterances into widows, each window being twice a maximum utterance threshold;   calculating a probability that each word in the plurality of windows is a meaning unit boundary based upon at least a linguistic model applied to each of the plurality of windows;   splitting the selected utterances which are longer than the predetermined minimum threshold into two new utterances at a word with a maximum calculated probability; and   identifying at least one of the two new utterances that is shorter than a maximum utterance threshold as a meaning unit.   
     
     
         18 . The method of  claim 17 , wherein the linguistic model comprises probability of words to form the beginning or end of a meaning unit. 
     
     
         19 . The method of  claim 18 , further comprising receiving audio data and decoding the audio data with at least the linguistic model to create the transcription of audio data. 
     
     
         20 . The method of  claim 19 , further comprising applying at least one of a linguistic exception and an acoustic exception to the two new utterances.

Join the waitlist — get patent alerts

Track US2015066506A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.