Speech synthesis apparatus and method, and storage medium
Abstract
Input text data undergoes language analysis to generate prosody, and a speech database is searched for a synthesis unit on the basis of the prosody. A modification distortion of the found synthesis unit, and concatenation distortions upon connecting that synthesis unit to those in the preceding phoneme are computed, and a distortion determination unit weights the modification and concatenation distortions to determine the total distortion. An Nbest determination unit obtains N best paths that can minimize the distortion using the A* search algorithm, and a registration unit determination unit selects a synthesis unit to be registered in a synthesis unit inventory on the basis of the N best paths in the order of frequencies of occurrence, and registers it in the synthesis unit inventory.
Claims
exact text as granted — not AI-modified1 . A synthesis unit selection apparatus comprising:
n-best obtaining means for obtaining one or more best sequences of synthesis units corresponding to a phonetic string on the basis of a distortion obtained by concatenating synthesis units; obtaining means for obtaining a plurality of sequences by applying said n-best obtaining means to a corpus including a plurality of phonetic strings; and selection means for selecting synthesis units on the basis of the plurality of sequences obtained by said obtaining means.
2 . The apparatus according to claim 1 , wherein the distortion comprises at least one of a concatenation distortion and a modification distortion, and the modification distortion is a distortion between a synthesis unit before and after modification.
3 . The apparatus according to claim 1 , further comprising:
text reception means for receiving text data, wherein the plurality of phonetic strings are included in the text data received by said text reception means.
4 . The apparatus according to claim 1 , further comprising:
registration means for registering the synthesis units selected by said selection means in a synthesis unit inventory in a memory.
5 . The apparatus according to claim 2 , wherein said selection means selects a synthesis units on the basis of a weighted sum of the concatenation and modification distortions.
6 . (canceled)
7 . (canceled)
8 . The apparatus according to claim 2 , wherein said obtaining means determines the modification distortion by looking up a table that stores the modification distortion.
9 . The apparatus according to claim 2 , wherein said obtaining means determines the concatenation distortion by looking up a table that stores the concatenation distortion.
10 . (canceled)
11 . A synthesis unit selection method comprising:
an n-best obtaining step of obtaining one or more best sequences of synthesis units corresponding to a phonetic string on the basis of a distortion obtained by concatenating synthesis units; an obtaining step of obtaining a plurality of sequences by applying said n-best obtaining step to a corpus including a plurality of phonetic strings; and a selection step of selecting synthesis units on the basis of the plurality of sequences obtained in said obtaining step.
12 . The method according to claim 11 , wherein the distortion comprises at least one of a concatenation distortion and a modification distortion, and the modification distortion is a distortion between a synthesis unit before and after modification.
13 . The method according to claim 11 , further comprising the step of:
receiving text data, wherein the plurality of phonetic strings are included in the text data received in said receiving step.
14 . The method according to claim 11 , further comprising the step of:
registering the synthesis units selected in said selection step in a synthesis unit inventory.
15 . The method according to claim 12 , wherein in said selection step, a synthesis unit is selected on the basis of a weighted sum of the concatenation and modification distortions.
16 . (canceled)
17 . (canceled)
18 . The method according to claim 12 , wherein in said obtaining step, the modification distortion is determined by looking up a table that stores the modification distortion.
19 . The method according to claim 12 , wherein in said obtaining step, the concatenation distortion is determined by looking up a table that stores the concatenation distortion.
20 . (canceled)
21 . A computer readable storage medium storing a program that implements the method recited in claim 11 .
22 . A synthesis unit selection apparatus comprising:
an n-best obtaining unit configured to obtain one or more best sequences of synthesis units corresponding to a phonetic string on the basis of a distortion obtained by concatenating synthesis units; an obtaining unit configured to obtain a plurality of sequences by applying said n-best obtaining unit to a corpus including a plurality of phonetic strings; and a selection unit configured to select synthesis units on the basis of the plurality of sequences obtained by said obtaining unit.
23 . A program for implementing a synthesis unit selection method comprising:
an n-best obtaining step module for obtaining one or more best sequences of synthesis units corresponding to a phonetic string on the basis of a distortion obtained by concatenating synthesis units; an obtaining step module for obtaining a plurality of sequences by applying said n-best obtaining step module to a corpus including a plurality of phonetic strings; and a selection step module for selecting synthesis units on the basis of the plurality of sequences obtained by said obtaining step module.Join the waitlist — get patent alerts
Track US2006085194A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.