US2024021189A1PendingUtilityA1

Configurable neural speech synthesis

Assignee: SOUNDHOUND INCPriority: Jun 12, 2020Filed: Jul 14, 2023Published: Jan 18, 2024
Est. expiryJun 12, 2040(~13.9 yrs left)· nominal 20-yr term from priority
Inventors:Andrew Richards
G06N 3/0895G06N 3/09G06N 3/096G06N 3/0475G10L 13/047G10L 13/08G10L 13/033G10L 15/26G06N 3/084G06N 3/04G06F 3/167G06F 3/04847G06N 3/048G06N 3/045
71
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A discriminator trained on labeled samples of speech can compute probabilities of voice properties. A speech synthesis generative neural network that takes in text and continuous scale values of voice properties is trained to synthesize speech audio that the discriminator will infer as matching the values of the input voice properties. Voice parameters can include speaker voice parameters, accents, and attitudes, among others. Training can be done by transfer learning from an existing neural speech synthesis model or such a model can be trained with a loss function that considers speech and parameter values. A graphical user interface can allow voice designers for products to synthesize speech with a desired voice or generate a speech synthesis engine with frozen voice parameters. A vector of parameters can be used for comparison to previously registered voices in databases such as ones for trademark registration.

Claims

exact text as granted — not AI-modified
1 . A computerized process of training a neural speech synthesis model that can generate speech audio conditioned on a value of a voice property, the computerized process comprising:
 obtaining source samples of speech audio;   labeling the source samples with discrete values of a voice property;   training, from the source samples and labels, a discriminator that can compute a probability of the voice property from a sample of speech audio; and   training the neural speech synthesis model by:   synthesizing a multiplicity of synthesized speech samples using the neural speech synthesis model with a multiplicity of values of the voice property to generate synthesized speech samples,   computing corresponding probabilities for the synthesized speech samples using the discriminator, and   computing a property-learning weight adjustment to the neural speech synthesis model by back-propagating changes to minimize a loss function that depends on differences between values of the voice property and corresponding probabilities.   
     
     
         2 . A speech synthesis model obtained by the computerized process of  claim 1 . 
     
     
         3 . The speech synthesis model obtained of  claim 2 , wherein the speech synthesis model is configured to:
 receive a string of text and at least one voice property value with a perceptible meaning;   synthesize speech audio corresponding to the string of text using a neural speech synthesis model that conditions a sound of speech audio on the at least one voice property value to generate synthesized speech audio; and   output the synthesized speech audio, wherein the sound of the synthesized speech audio perceptually relates to the at least one voice property value.   
     
     
         4 . The speech synthesis model of  claim 3 , wherein the at least one voice property value includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property. 
     
     
         5 . The speech synthesis model of  claim 3 , wherein the speech synthesis model is further configured to:
 enable download of the synthesized speech audio.   
     
     
         6 . The speech synthesis model of  claim 3 , wherein the speech synthesis model is further configured to:
 enable playback of the synthesized speech audio.   
     
     
         7 . The speech synthesis model of  claim 3 , wherein the speech synthesis model is further configured to:
 provide a graphical user interface that includes one of a text input field or a voice property value input field.   
     
     
         8 . The speech synthesis model of  claim 3 , wherein the string of text is associated with at least one text tag. 
     
     
         9 . The speech synthesis model of  claim 3 , wherein the string of text indicates dynamically configurable voice parameter values. 
     
     
         10 . The computerized process of  claim 1 , wherein the synthesizing uses a transcription of source samples, the computerized process further comprising:
 computing a source-matching weight adjustment by back-propagating changes to minimize a loss function that depends on differences between the source samples and the synthesized speech samples.   
     
     
         11 . A speech synthesis model obtained by the computerized process of  claim 10 . 
     
     
         12 . The computerized process of  claim 1 , wherein the source samples of the speech audio are obtained from one of a person and an audio generation system. 
     
     
         13 . The computerized process of  claim 1 , wherein the voice property includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property. 
     
     
         14 . A computer system for training a neural speech synthesis model to generate speech audio conditioned on a value of a voice property, comprising:
 at least one processor; and   memory including instructions that, when executed by the at least one processor, cause the computer system to:   
       obtain source samples of speech audio; 
       label the source samples with discrete values of a voice property; 
       train, from the source samples and labels, a discriminator that can compute a probability of the voice property from a sample of speech audio; and 
       train the neural speech synthesis model by: 
       synthesize a multiplicity of synthesized speech samples using the neural speech synthesis model with a multiplicity of values of the voice property to generate synthesized speech samples, 
       compute corresponding probabilities for the synthesized speech samples using the discriminator, and 
       compute a property-learning weight adjustment to the neural speech synthesis model by back-propagating changes to minimize a loss function that depends on differences between values of the voice property and corresponding probabilities. 
     
     
         15 . The computer system of  claim 14 , wherein the at least one voice property value includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property. 
     
     
         16 . The computer system of  claim 14 , wherein the neural speech synthesis model is further configured to:
 enable download of the synthesized speech audio.   
     
     
         17 . The computer system of  claim 14 , wherein the neural speech synthesis model is further configured to:
 enable playback of the synthesized speech audio.   
     
     
         18 . The computer system of  claim 14 , wherein the neural speech synthesis model is further configured to:
 provide a graphical user interface that includes one of a text input field or a voice property value input field.   
     
     
         19 . The computer system of  claim 14 , wherein the string of text is associated with at least one text tag. 
     
     
         20 . The computer system of  claim 14 , wherein the system uses a transcription of source samples, and wherein the instructions when executed further cause the computer system to:
 compute a source-matching weight adjustment by back-propagating changes to minimize a loss function that depends on differences between the source samples and the synthesized speech samples.

Join the waitlist — get patent alerts

Track US2024021189A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.