Voice synthesis method, apparatus, device and storage medium
Abstract
Provided are a voice synthesis method, an apparatus, a device, and a storage medium, involving obtaining text information and determining characters in the text information and a text content of each of the characters; performing a character recognition on the text content of each of the characters, to determine character attribute information of each of the characters; obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, where the speakers are pre-stored pronunciation object having the character attribute information; and generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information. These improve pronunciation diversities of different characters in the synthesized voices, improve an audience's discrimination between different characters in the synthesized voices, and thereby improve experience of a user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A voice synthesis method, comprising:
obtaining text information and determining characters in the text information and a text content of each of the characters;
performing a character recognition on the text content of each of the characters, to determine character attribute information of the each of the characters;
obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, wherein the speakers are pre-stored speakers having the character attribute information; and
generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information;
wherein the character attribute information comprises a basic attribute, and the basic attribute comprises at least one of a gender attribute and an age attribute;
before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the method further comprises:
determining the basic attribute corresponding to each of the pre-stored speakers according to voice parameter information of the pre-stored speakers; and
correspondingly the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters comprises:
for each of the characters, obtaining a speaker having the basic attribute corresponding to the each of the characters,
wherein the character attribute information further comprises an additional attribute, and the additional attribute comprises at least one of the following:
regional information, timbre information, and pronunciation style information;
before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the method further comprises:
determining the additional attribute and additional attribute priority corresponding to each of the pre-stored speakers according to the voice parameter information of the pre-stored speakers, and
correspondingly the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters further comprises:
determining whether the speaker having the basic attribute corresponding to the character is unique such that the speaker having the basic attribute is the only one of the pre-stored speakers having the basic attribute;
if yes, using the unique speaker as the speaker in one-to-one correspondence with the character;
if no, determining, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute;
wherein the determining, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute comprises:
obtaining a character voice description class keyword in text contents of the characters,
determining the additional attribute corresponding to the characters according to the character voice description class keyword, and
in the speakers having the basic attribute corresponding to the characters, using speakers with highest additional attribute priorities as the speakers in one-to-one correspondence with the characters.
2. The method according to claim 1 , wherein the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters comprises:
obtaining a candidate speaker for each of the characters according to the character attribute information of the each of the characters;
displaying description information of the candidate speaker to a user and receiving an indication of the user; and
obtaining the speakers in one-to-one correspondence with the characters in the candidate speaker of each of the characters according to the indication of the user.
3. The method according to claim 1 , wherein the generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information comprises:
processing a corresponding text content in the text information according to the speakers corresponding to the characters, to generate the multi-character synthesized voices.
4. A device comprising a sender, a receiver, a memory, and a processor;
the memory is configured to store computer instructions; the processor is configured to execute the computer instructions stored in the memory to:
obtain text information and determining characters in the text information and a text content of each of the characters;
perform a character recognition on the text content of each of the characters, to determine character attribute information of the each of the characters;
obtain speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, wherein the speakers are pre-stored speakers having the character attribute information; and
generate multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information;
wherein the character attribute information comprises a basic attribute, and the basic attribute comprises at least one of a gender attribute and an age attribute;
before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to:
determine the basic attribute corresponding to each of the pre-stored speakers according to voice parameter information of the pre-stored speakers; and
correspondingly, in the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to:
for each of the characters, obtain a speaker having the basic attribute corresponding to the each of the characters,
wherein the character attribute information further comprises an additional attribute, and the additional attribute comprises at least one of the following:
regional information, timbre information, and pronunciation style information;
before the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to:
determine the additional attribute and additional attribute priority corresponding to each of the pre-stored speakers according to the voice parameter information of the pre-stored speakers, and
correspondingly, in the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to:
determine whether the speaker having the basic attribute corresponding to the character is unique such that the speaker having the basic attribute is the only one of the pre-stored speakers having the basic attribute;
if yes, using the unique speaker as the speaker in one-to-one correspondence with the character;
if no, determine, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute;
wherein in determining, from speakers having the basic attribute corresponding to the characters, the speakers in one-to-one correspondence with the characters according to the additional attribute, the processor is configured to:
obtain a character voice description class keyword in text contents of the characters,
determine the additional attribute corresponding to the characters according to the character voice description class keyword, and
in the speakers having the basic attribute corresponding to the characters, use speakers with highest additional attribute priorities as the speakers in one-to-one correspondence with the characters.
5. The device according to claim 4 , wherein in the obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, the processor is configured to:
obtain a candidate speaker for each of the characters according to the character attribute information of the each of the characters;
display description information of the candidate speaker to a user and receiving an indication of the user; and
obtain the speakers in one-to-one correspondence with the characters in the candidate speaker of each of the characters according to the indication of the user.
6. The device according to claim 4 , wherein in the generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information, the processor is configured to:
process a corresponding text content in the text information according to the speakers corresponding to the characters, to generate the multi-character synthesized voices.
7. A storage medium comprising a non-transitory readable storage medium and computer instructions stored in the non-transitory readable storage medium; the computer instructions are configured to implement the voice synthesis method according to claim 1 .
8. The method according to claim 3 , wherein after the processing a corresponding text content in the text information according to the speakers corresponding to the characters, to generate the multi-character synthesized voices, the method further comprises:
obtaining background audios that are matched with a plurality of consecutive text contents in the text information; and
adding the background audio to voices corresponding to the plurality of text contents, in the multi-character synthesized voices.
9. The device according to claim 6 , wherein after the processing a corresponding text content in the text information according to the speakers corresponding to the characters to generate the multi-character synthesized voices, the processor is configured to:
obtain background audios that are matched with a plurality of consecutive text contents in the text information; and
add the background audio to voices corresponding to the plurality of text contents, in the multi-character synthesized voices.Join the waitlist — get patent alerts
Track US11600259B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.