US2023230572A1PendingUtilityA1

End-to-end speech conversion

Assignee: GOOGLE LLCPriority: Feb 21, 2019Filed: Mar 23, 2023Published: Jul 20, 2023
Est. expiryFeb 21, 2039(~12.5 yrs left)· nominal 20-yr term from priority
G06N 3/0464G06N 3/0455G06N 3/0442G06N 3/09G10L 13/027G10L 21/003G10L 25/18G10L 21/04G10L 15/063G10L 13/08G10L 13/02G06N 3/08G10L 21/10G10L 25/30H04L 51/02
71
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for end to end speech conversion are disclosed. In one aspect, a method includes the actions of receiving first audio data of a first utterance of one or more first terms spoken by a user. The actions further include providing the first audio data as an input to a model that is configured to receive first given audio data in a first voice and output second given audio data in a synthesized voice without performing speech recognition on the first given audio data. The actions further include receiving second audio data of a second utterance of the one or more first terms spoken in the synthesized voice. The actions further include providing, for output, the second audio data of the second utterance of the one or more first terms spoken in the synthesized voice.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:
 receiving a sequence of source audio frames characterizing an utterance spoken in a first accent;   processing, using an encoder of a voice conversion model, the sequence of source audio frames to generate a sequence of source internal representations characterizing the utterance spoken in the first accent;   processing, using a decoder of the voice conversion model, the sequence of source internal representations to generate a sequence of target audio frames characterizing a synthesized speech representation of the utterance in a second accent different than the first accent; and   providing, for output by a computing device, the synthesized speech representation of the utterance in the second accent.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein processing the sequence of source internal representations to generate a synthesized speech representation of the utterance in a second accent comprises processing the sequence of source internal representations to generate the synthesized speech representation without performing any speech recognition on the sequence of source audio frames. 
     
     
         3 . The computer-implemented method of  claim 1 , wherein the sequence of source audio frames comprises a sequence of input spectrograms. 
     
     
         4 . The computer-implemented method of  claim 1 , wherein the sequence of target audio frames comprises a sequence of output spectrograms. 
     
     
         5 . The computer-implemented method of  claim 1 , wherein a cadence of the utterance spoken in the first accent is different than a cadence of the synthesized speech representation of the utterance in the second accent. 
     
     
         6 . The computer-implemented method of  claim 1 , wherein the encoder comprises a bidirectional long short-term memory (LSTM) layer. 
     
     
         7 . The computer-implemented method of  claim 1 , wherein the decoder comprises a spectrogram decoder with attention. 
     
     
         8 . The computer-implemented method of  claim 1 , wherein the operations further comprise:
 receiving audio data of a collection of utterances;   obtaining a transcription of each utterance in the collection of utterances;   providing the transcriptions of each utterance as an input to a text to speech model;   receiving, for each transcription of each utterance, audio data of an additional collection of utterances in a synthesized voice; and   training the model using the audio data of the collection of utterances and the audio data of an additional collection of utterances in a synthesized voice.   
     
     
         9 . The computer-implemented method of  claim 1 , wherein the operations further comprise bypassing obtaining a transcription of the utterance. 
     
     
         10 . The computer-implemented method of  claim 1 , wherein the speech conversion model is configured to adjust a time period between each term in the utterance spoken in the first accent. 
     
     
         11 . A system comprising:
 data processing hardware; and   memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:
 receiving a sequence of source audio frames characterizing an utterance spoken in a first accent; 
 processing, using an encoder of a voice conversion model, the sequence of source audio frames to generate a sequence of source internal representations characterizing the utterance spoken in the first accent; 
 processing, using a decoder of the voice conversion model, the sequence of source internal representations to generate a sequence of target audio frames characterizing a synthesized speech representation of the utterance in a second accent different than the first accent; and 
 providing, for output by a computing device, the synthesized speech representation of the utterance in the second accent. 
   
     
     
         12 . The system of  claim 11 , wherein processing the sequence of source internal representations to generate a synthesized speech representation of the utterance in a second accent comprises processing the sequence of source internal representations to generate the synthesized speech representation without performing any speech recognition on the sequence of source audio frames. 
     
     
         13 . The system of  claim 11 , wherein the sequence of source audio frames comprises a sequence of input spectrograms. 
     
     
         14 . The system of  claim 11 , wherein the sequence of target audio frames comprises a sequence of output spectrograms. 
     
     
         15 . The system of  claim 11 , wherein a cadence of the utterance spoken in the first accent is different than a cadence of the synthesized speech representation of the utterance in the second accent. 
     
     
         16 . The system of  claim 11 , wherein the encoder comprises a bidirectional long short-term memory (LSTM) layer. 
     
     
         17 . The system of  claim 11 , wherein the decoder comprises a spectrogram decoder with attention. 
     
     
         18 . The system of  claim 11 , wherein the operations further comprise:
 receiving audio data of a collection of utterances;   obtaining a transcription of each utterance in the collection of utterances;   providing the transcriptions of each utterance as an input to a text to speech model;   receiving, for each transcription of each utterance, audio data of an additional collection of utterances in a synthesized voice; and   training the model using the audio data of the collection of utterances and the audio data of an additional collection of utterances in a synthesized voice.   
     
     
         19 . The system of  claim 11 , wherein the operations further comprise bypassing obtaining a transcription of the utterance. 
     
     
         20 . The system of  claim 11 , wherein the speech conversion model is configured to adjust a time period between each term in the utterance spoken in the first accent.

Join the waitlist — get patent alerts

Track US2023230572A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.