US2003018472A1PendingUtilityA1

Vocoder-based voice recognizer

Priority: Jan 8, 1998Filed: Jan 22, 2002Published: Jan 23, 2003
Est. expiryJan 8, 2018(expired)· nominal 20-yr term from priority
G10L 19/04G10L 15/30G10L 15/02G10L 2025/783G10L 13/00
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A vocoder based voice recognizer recognizes a spoken word using linear prediction coding (LPC) based, vocoder data without completely reconstructing the voice data. The recognizer generates at least one energy estimate per frame of the vocoder data and searches for word boundaries in the vocoder data using the associated energy estimates. If a word is found, the LPC word parameters are extracted from the vocoder data associated with the word and recognition features are calculated from the extracted LPC word parameters. Finally, the recognition features are matched with previously stored recognition features of other words, thereby to recognize the spoken word.

Claims

exact text as granted — not AI-modified
1 . A method for recognizing a spoken word using linear prediction coding (LPC) based, vocoder data without completely reconstructing the voice data, the vocoder data formed into a series of frames, the method comprising the steps of: 
 generating at least one energy estimate per frame of said vocoder data;    searching for word boundaries in said vocoder data using the associated energy estimates;    if a word is found, extracting the LPC word parameters from the vocoder data associated with said word;    calculating recognition features from said extracted LPC word parameters; and    matching said recognition features with previously stored recognition features of other words, thereby to recognize the spoken word.    
     
     
         2 . A method for preparing to recognize a spoken word using linear prediction coding (LPC) based, vocoder data without completely reconstructing the voice data, the vocoder data formed into a series of frames, the method comprising the steps of: 
 generating at least one energy estimate per frame of said vocoder data;    searching for word boundaries in said vocoder data using the associated energy estimates;    if a word is found, extracting the LPC word parameters from the vocoder data associated with said word;    calculating recognition features from said extracted LPC word parameters.    
     
     
         3 . A method according to  claim 2  and wherein said step of generating comprises the step of estimating the energy from residual data found in said vocoder data.  
     
     
         4 . A method according to  claim 3  and wherein said step of estimating comprises the steps of reconstructing residual data from said vocoder data and generating the norm of said residual data.  
     
     
         5 . A method according to  claim 3  and wherein said step of estimating comprises the steps of extracting a pitch-gain value from said vocoder data and using said extracted pitch-gain value as said energy estimate.  
     
     
         6 . A method according to  claim 3  and wherein said step of generating comprises the steps of: 
 extracting pitch-gain values, lag values and remnant data from said vocoder data;  
 reconstructing a remnant signal from said remnant data;  
 generating an energy estimate of said remnant signal;  
 generating an energy estimate of a non-remnant portion of said residual by using said pitch-gain value and a previous energy estimate defined by said lag value; and  
 combining said remnant and non-remnant energy estimates.  
 
     
     
         7 . A method according to  claim 1  and wherein said vocoder data is of the type produced by any of the following vocoders: RPE-LTP full and half rate, QCELP 8 and 13 Kbps, EVRC, LD CELP, VSELP, CS ACELP, Enhanced Full Rate Vocoder and LPC10.  
     
     
         8 . A method according to  claim 2  and wherein said vocoder data is of the type produced by any of the following vocoders: RPE-LTP full and half rate, QCELP 8 and 13 Kbps, EVRC, LD CELP, VSELP, CS ACELP, Enhanced Full Rate Vocoder and LPC10.  
     
     
         9 . The use of LPC-based vocoder data as an input to a voice recognition system.  
     
     
         10 . A digital cellular telephone comprising: 
 a mobile telephone operating system;    a vocoder which compresses a voice signal using at least linear predication coding (LPC) thereby to produce vocoder data; and    a vocoder based voice recognizer comprising: 
 a front end processor which processes said vocoder data to determine when a word was spoken and to generate recognition features of said spoken word; and  
 a recognizer which at least recognizes said spoken word as one of a set of reference words.  
   
     
     
         11 . A digital cellular telephone according to  claim 10  and wherein said front end processor includes: 
 an energy estimator which uses residual information forming part of said vocoder data to estimate the energy of a voice signal;  
 an LPC parameter extractor which extracts the LPC parameters of said vocoder data; and  
 a recognition feature generator which generates said recognition features from said LPC parameters.  
 
     
     
         12 . A cellular telephone according to  claim 10  and wherein said front end processor is selectably operable with multiple vocoder types.  
     
     
         13 . A cellular telephone according to  claim 10  and wherein said vocoder is any of the following vocoders: RPE-LTP full and half rate, QCELP 8 and 13 Kbps, EVRC, LD CELP, VSELP, CS ACELP, Enhanced Full Rate Vocoder and LPC10.  
     
     
         14 . A vocoder based voice recognizer operable with the data produced by an LPC-based vocoder, the voice recognizer comprising: 
 a front end processor which processes said vocoder data to determine when a word was spoken and to generate recognition features of said spoken word; and    a recognizer which at least recognizes said spoken word as one of a set of reference words.    
     
     
         15 . A voice recognizer according to  claim 14  and wherein said front end processor comprises: 
 an energy estimator which uses residual information forming part of said vocoder data to estimate the energy of a voice signal;  
 an LPC parameter extractor which extracts the LPC parameters of said vocoder data; and  
 a recognition feature generator which generates said recognition features from said LPC parameters.  
 
     
     
         16 . A voice recognizer according to  claim 15  and wherein said energy estimator comprises a residual energy estimator which estimates the energy from residual data found in said vocoder data.  
     
     
         17 . A voice recognizer according to  claim 16  and wherein said residual energy estimator comprises a residual reconstructor which reconstructs residual data from said vocoder data and a norm generator which generates the norm of said residual data thereby to produce said energy estimate.  
     
     
         18 . A voice recognizer according to  claim 16  and wherein said residual energy estimator comprises an extractor which extracts a pitch-gain value from said vocoder data thereby to produce said energy estimate.  
     
     
         19 . A voice recognizer according to  claim 16  and wherein said residual energy estimator comprises: 
 an extractor which extracts pitch-gain values, lag values and remnant data from said vocoder data;  
 a reconstructor which reconstructs a remnant signal from said remnant data;  
 a remnant energy estimator which generates an energy estimate of said remnant signal;  
 a non-remnant energy estimator which generates an energy estimate of a non-remnant portion of said residual by using said pitch-gain value and a previous energy estimate defined by said lag value; and  
 a combiner which combines said remnant and non-remnant energy estimates thereby to produce said energy estimate.  
 
     
     
         20 . A voice recognizer according to  claim 14  and wherein said vocoder is any of the following vocoders: RPE-LTP full and half rate, QCELP 8 and 13 Kbps, EVRC, LD CELP, VSELP, CS ACELP, Enhanced Full Rate Vocoder and LPC10.

Join the waitlist — get patent alerts

Track US2003018472A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.