Configurable neural speech synthesis
Abstract
A discriminator trained on labeled samples of speech can compute probabilities of voice properties. A speech synthesis generative neural network that takes in text and continuous scale values of voice properties is trained to synthesize speech audio that the discriminator will infer as matching the values of the input voice properties. Voice parameters can include speaker voice parameters, accents, and attitudes, among others. Training can be done by transfer learning from an existing neural speech synthesis model or such a model can be trained with a loss function that considers speech and parameter values. A graphical user interface can allow voice designers for products to synthesize speech with a desired voice or generate a speech synthesis engine with frozen voice parameters. A vector of parameters can be used for comparison to previously registered voices in databases such as ones for trademark registration.
Claims
exact text as granted — not AI-modified1 . A computerized process of training a neural speech synthesis model that can generate speech audio conditioned on a value of a voice property, the computerized process comprising:
obtaining source samples of speech audio; labeling the source samples with discrete values of a voice property; training, from the source samples and labels, a discriminator that can compute a probability of the voice property from a sample of speech audio; and training the neural speech synthesis model by: synthesizing a multiplicity of synthesized speech samples using the neural speech synthesis model with a multiplicity of values of the voice property to generate synthesized speech samples, computing corresponding probabilities for the synthesized speech samples using the discriminator, and computing a property-learning weight adjustment to the neural speech synthesis model by back-propagating changes to minimize a loss function that depends on differences between values of the voice property and corresponding probabilities.
2 . A speech synthesis model obtained by the computerized process of claim 1 .
3 . The speech synthesis model obtained of claim 2 , wherein the speech synthesis model is configured to:
receive a string of text and at least one voice property value with a perceptible meaning; synthesize speech audio corresponding to the string of text using a neural speech synthesis model that conditions a sound of speech audio on the at least one voice property value to generate synthesized speech audio; and output the synthesized speech audio, wherein the sound of the synthesized speech audio perceptually relates to the at least one voice property value.
4 . The speech synthesis model of claim 3 , wherein the at least one voice property value includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property.
5 . The speech synthesis model of claim 3 , wherein the speech synthesis model is further configured to:
enable download of the synthesized speech audio.
6 . The speech synthesis model of claim 3 , wherein the speech synthesis model is further configured to:
enable playback of the synthesized speech audio.
7 . The speech synthesis model of claim 3 , wherein the speech synthesis model is further configured to:
provide a graphical user interface that includes one of a text input field or a voice property value input field.
8 . The speech synthesis model of claim 3 , wherein the string of text is associated with at least one text tag.
9 . The speech synthesis model of claim 3 , wherein the string of text indicates dynamically configurable voice parameter values.
10 . The computerized process of claim 1 , wherein the synthesizing uses a transcription of source samples, the computerized process further comprising:
computing a source-matching weight adjustment by back-propagating changes to minimize a loss function that depends on differences between the source samples and the synthesized speech samples.
11 . A speech synthesis model obtained by the computerized process of claim 10 .
12 . The computerized process of claim 1 , wherein the source samples of the speech audio are obtained from one of a person and an audio generation system.
13 . The computerized process of claim 1 , wherein the voice property includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property.
14 . A computer system for training a neural speech synthesis model to generate speech audio conditioned on a value of a voice property, comprising:
at least one processor; and memory including instructions that, when executed by the at least one processor, cause the computer system to:
obtain source samples of speech audio;
label the source samples with discrete values of a voice property;
train, from the source samples and labels, a discriminator that can compute a probability of the voice property from a sample of speech audio; and
train the neural speech synthesis model by:
synthesize a multiplicity of synthesized speech samples using the neural speech synthesis model with a multiplicity of values of the voice property to generate synthesized speech samples,
compute corresponding probabilities for the synthesized speech samples using the discriminator, and
compute a property-learning weight adjustment to the neural speech synthesis model by back-propagating changes to minimize a loss function that depends on differences between values of the voice property and corresponding probabilities.
15 . The computer system of claim 14 , wherein the at least one voice property value includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property.
16 . The computer system of claim 14 , wherein the neural speech synthesis model is further configured to:
enable download of the synthesized speech audio.
17 . The computer system of claim 14 , wherein the neural speech synthesis model is further configured to:
enable playback of the synthesized speech audio.
18 . The computer system of claim 14 , wherein the neural speech synthesis model is further configured to:
provide a graphical user interface that includes one of a text input field or a voice property value input field.
19 . The computer system of claim 14 , wherein the string of text is associated with at least one text tag.
20 . The computer system of claim 14 , wherein the system uses a transcription of source samples, and wherein the instructions when executed further cause the computer system to:
compute a source-matching weight adjustment by back-propagating changes to minimize a loss function that depends on differences between the source samples and the synthesized speech samples.Join the waitlist — get patent alerts
Track US2024021189A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.