US11600259B2ActiveUtilityA1

Voice synthesis method, apparatus, device and storage medium

Assignee: Baidu online network technology beijing co ltdPriority: Dec 20, 2018Filed: Sep 10, 2019Granted: Mar 7, 2023
Est. expiryDec 20, 2038(~12.4 yrs left)· nominal 20-yr term from priority
Inventors:Jie Yang
G10L 13/10G10L 13/027G10L 13/0335G10L 13/08G10L 13/033G10L 2013/083
45
PatentIndex Score
0
Cited by
9
References
9
Claims

Abstract

Provided are a voice synthesis method, an apparatus, a device, and a storage medium, involving obtaining text information and determining characters in the text information and a text content of each of the characters; performing a character recognition on the text content of each of the characters, to determine character attribute information of each of the characters; obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, where the speakers are pre-stored pronunciation object having the character attribute information; and generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information. These improve pronunciation diversities of different characters in the synthesized voices, improve an audience's discrimination between different characters in the synthesized voices, and thereby improve experience of a user.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A voice synthesis method, comprising:
 obtaining text information and determining characters in the text information and a text content of each of the characters; 
 performing a character recognition on the text content of each of the characters, to determine character attribute information of the each of the characters; 
 obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, wherein the speakers are pre-stored speakers having the character attribute information; and 
 generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information; 
 wherein the character attribute information comprises a basic attribute, and the basic attribute comprises at least one of a gender attribute and an age attribute; 
 before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the method further comprises: 
 determining the basic attribute corresponding to each of the pre-stored speakers according to voice parameter information of the pre-stored speakers; and 
 correspondingly the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters comprises: 
 for each of the characters, obtaining a speaker having the basic attribute corresponding to the each of the characters, 
 wherein the character attribute information further comprises an additional attribute, and the additional attribute comprises at least one of the following: 
 regional information, timbre information, and pronunciation style information; 
 before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the method further comprises: 
 determining the additional attribute and additional attribute priority corresponding to each of the pre-stored speakers according to the voice parameter information of the pre-stored speakers, and 
 correspondingly the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters further comprises: 
 determining whether the speaker having the basic attribute corresponding to the character is unique such that the speaker having the basic attribute is the only one of the pre-stored speakers having the basic attribute; 
 if yes, using the unique speaker as the speaker in one-to-one correspondence with the character; 
 if no, determining, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute; 
 wherein the determining, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute comprises: 
 obtaining a character voice description class keyword in text contents of the characters, 
 determining the additional attribute corresponding to the characters according to the character voice description class keyword, and 
 in the speakers having the basic attribute corresponding to the characters, using speakers with highest additional attribute priorities as the speakers in one-to-one correspondence with the characters. 
 
     
     
       2. The method according to  claim 1 , wherein the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters comprises:
 obtaining a candidate speaker for each of the characters according to the character attribute information of the each of the characters; 
 displaying description information of the candidate speaker to a user and receiving an indication of the user; and 
 obtaining the speakers in one-to-one correspondence with the characters in the candidate speaker of each of the characters according to the indication of the user. 
 
     
     
       3. The method according to  claim 1 , wherein the generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information comprises:
 processing a corresponding text content in the text information according to the speakers corresponding to the characters, to generate the multi-character synthesized voices. 
 
     
     
       4. A device comprising a sender, a receiver, a memory, and a processor;
 the memory is configured to store computer instructions; the processor is configured to execute the computer instructions stored in the memory to: 
 obtain text information and determining characters in the text information and a text content of each of the characters; 
 perform a character recognition on the text content of each of the characters, to determine character attribute information of the each of the characters; 
 obtain speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, wherein the speakers are pre-stored speakers having the character attribute information; and 
 generate multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information; 
 wherein the character attribute information comprises a basic attribute, and the basic attribute comprises at least one of a gender attribute and an age attribute; 
 before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to: 
 determine the basic attribute corresponding to each of the pre-stored speakers according to voice parameter information of the pre-stored speakers; and 
 correspondingly, in the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to: 
 for each of the characters, obtain a speaker having the basic attribute corresponding to the each of the characters, 
 wherein the character attribute information further comprises an additional attribute, and the additional attribute comprises at least one of the following: 
 regional information, timbre information, and pronunciation style information; 
 before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to: 
 determine the additional attribute and additional attribute priority corresponding to each of the pre-stored speakers according to the voice parameter information of the pre-stored speakers, and 
 correspondingly, in the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to: 
 determine whether the speaker having the basic attribute corresponding to the character is unique such that the speaker having the basic attribute is the only one of the pre-stored speakers having the basic attribute; 
 if yes, using the unique speaker as the speaker in one-to-one correspondence with the character; 
 if no, determine, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute; 
 wherein in determining, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute, the processor is configured to: 
 obtain a character voice description class keyword in text contents of the characters, 
 determine the additional attribute corresponding to the characters according to the character voice description class keyword, and 
 in the speakers having the basic attribute corresponding to the characters, use speakers with highest additional attribute priorities as the speakers in one-to-one correspondence with the characters. 
 
     
     
       5. The device according to  claim 4 , wherein in the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to:
 obtain a candidate speaker for each of the characters according to the character attribute information of the each of the characters; 
 display description information of the candidate speaker to a user and receiving an indication of the user; and 
 obtain the speakers in one-to-one correspondence with the characters in the candidate speaker of each of the characters according to the indication of the user. 
 
     
     
       6. The device according to  claim 4 , wherein in the generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information, the processor is configured to:
 process a corresponding text content in the text information according to the speakers corresponding to the characters, to generate the multi-character synthesized voices. 
 
     
     
       7. A storage medium comprising a non-transitory readable storage medium and computer instructions stored in the non-transitory readable storage medium; the computer instructions are configured to implement the voice synthesis method according to  claim 1 . 
     
     
       8. The method according to  claim 3 , wherein after the processing a corresponding text content in the text information according to the speakers corresponding to the characters, to generate the multi-character synthesized voices, the method further comprises:
 obtaining background audios that are matched with a plurality of consecutive text contents in the text information; and 
 adding the background audio to voices corresponding to the plurality of text contents, in the multi-character synthesized voices. 
 
     
     
       9. The device according to  claim 6 , wherein after the processing a corresponding text content in the text information according to the speakers corresponding to the characters to generate the multi-character synthesized voices, the processor is configured to:
 obtain background audios that are matched with a plurality of consecutive text contents in the text information; and 
 add the background audio to voices corresponding to the plurality of text contents, in the multi-character synthesized voices.

Join the waitlist — get patent alerts

Track US11600259B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.