US2005209855A1PendingUtilityA1

Speech signal processing apparatus and method, and storage medium

Assignee: CANON KKPriority: Mar 31, 2000Filed: May 11, 2005Published: Sep 22, 2005
Est. expiryMar 31, 2020(expired)· nominal 20-yr term from priority
G09B 7/02G09B 19/04
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A speech segment search unit searches a speech database for speech segments that satisfy a phonetic environment, and a HMM learning unit computes the HMMs of phonemes on the basis of the search result. A segment recognition unit performs segment recognition of speech segments on the basis of the computed HMMs of the phonemes, and when the phoneme of the segment recognition result is equal to a phoneme of the source speech segment, that speech segment is registered in a segment dictionary.

Claims

exact text as granted — not AI-modified
1 . A speech signal processing apparatus comprising: 
 HMM learning means for computing HMMs of speech segments with information indicating a phonetic environment in a speech database;    segment recognition means for performing segment recognition of the speech segments in the speech database on the basis of the HMMs; and    registration means for registering a speech segment in a segment dictionary, in a case where the recognition result of the speech segment by said segment recognition means corresponds to the information indicating the phonetic environment of the speech segment.    
   
   
       2 - 23 . (canceled)  
   
   
       24 . The apparatus according to  claim 1 , wherein the information indicating the phonetic environment is a diphone label, and said segment recognition means categorizes speech segments into four categories CC, CV, VC, and VV (C: a consonant, V: a vowel), and performs segment recognition in each category.  
   
   
       25 . The apparatus according to  claim 1 , wherein said registration means comprises: 
 pattern storage means which has allowable patterns of information indicating the phonetic environment, and    said registration means checks if information indicating the phonetic environment of the speech segment matches one of the allowable patterns of information indicating the phonetic environment even if the information indicating the phonetic environment is not equal to the recognition result of said segment recognition means.    
   
   
       26 . The apparatus according to  claim 1 , wherein said segment recognition means computes likelihoods of speech segments of identical information indicating the phonetic environment, and 
 said registration means registers, in the segment dictionary, speech segments having maximum likelihoods or having likelihoods not less than a predetermined value.    
   
   
       27 . The apparatus according to  claim 26 , wherein said registration means registers, in the segment dictionary, speech segments having upper values obtained by normalizing the likelihoods by durations of the speech segments or likelihoods having the values not less than a predetermined value.  
   
   
       28 . A speech signal processing method comprising: 
 an HMM learning step of computing HMMs of speech segments with information indicating a phonetic environment in a speech database;    a segment recognition step of performing segment recognition of the speech segments in the speech database on the basis of the HMMs; and    a registration step of registering a speech segment in a segment dictionary, in a case where the recognition result of the speech segment in said segment recognition step corresponds to the information indicating the phonetic environment of the speech segment.    
   
   
       29 . The method according to  claim 28 , wherein the information indicating the phonetic environment is a diphone label, and said segment recognition step categorizes speech segments into four categories CC, CV, VC, and VV (C: a consonant, V: a vowel), and includes the step of performing segment recognition in each category.  
   
   
       30 . The method according to  claim 28 , wherein said registration step comprises: 
 a pattern storage step of registering allowable patterns of information indicating the phonetic environment, and    said registration step includes a step of checking whether the information indicating the phonetic environment of the speech segment matches one of the allowable patterns of information indicating the phonetic environment even if the information indicating the phonetic environment is not equal to the result in said segment recognition step.    
   
   
       31 . The method according to  claim 28 , wherein said segment recognition step includes a step of computing likelihoods of speech segments of identical information indicating the phonetic environment, and 
 said registration step includes a step of registering, in the segment dictionary, speech segments having maximum likelihoods or having likelihoods not less than a predetermined value.    
   
   
       32 . The method according to  claim 31 , wherein said registration step includes a step of registering, in the segment dictionary, speech segments having upper values obtained by normalizing the likelihoods by durations of the speech segments or likelihoods having the values not less than a predetermined value.  
   
   
       33 . A computer readable storage medium storing a program for implementing the method according to  claim 28 .  
   
   
       34 . A speech synthesis apparatus comprising: 
 speech synthesis means for synthesizing speech using the segment dictionary made by the speech signal processing apparatus according to  claim 1 .    
   
   
       35 . A speech synthesis method comprising: 
 a speech synthesis step of synthesizing speech using the segment dictionary made by the speech signal processing method according to  claim 28 .    
   
   
       36 . A computer readable storage medium storing a program for implementing the method according to  claim 35 .  
   
   
       37 . A speech signal processing apparatus comprising: 
 HMM learning means for computing HMMs of speech segments with information indicating a phonetic environment in a speech database;    segment recognition means for performing segment recognition of the speech segments in the speech database on the basis of the HMMs;    judgment means for judging whether the result of the segment recognition corresponds to the information indicating the phonetic environment of a speech segment; and    storage means for storing the result of the judgment judged by said judgment means associated with the speech segment.    
   
   
       38 . A speech signal processing method comprising: 
 an HMM learning step of computing HMMs of speech segments with information indicating a phonetic environment in a speech database;    a segment recognition step of performing segment recognition of the speech segments in the speech database on the basis of the HMMs;    a judgment step of judging whether the result of the segment recognition corresponds to the information indicating the phonetic environment of a speech segment; and    a storage step of storing the result of the judgment judged in said judgment step associated with the speech segment.    
   
   
       39 . A computer readable storage medium storing a program for implementing the method according to  claim 38.

Join the waitlist — get patent alerts

Track US2005209855A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.