End-to-end speech conversion
Abstract
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for end to end speech conversion are disclosed. In one aspect, a method includes the actions of receiving first audio data of a first utterance of one or more first terms spoken by a user. The actions further include providing the first audio data as an input to a model that is configured to receive first given audio data in a first voice and output second given audio data in a synthesized voice without performing speech recognition on the first given audio data. The actions further include receiving second audio data of a second utterance of the one or more first terms spoken in the synthesized voice. The actions further include providing, for output, the second audio data of the second utterance of the one or more first terms spoken in the synthesized voice.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:
receiving a sequence of source audio frames characterizing an utterance spoken in a first accent; processing, using an encoder of a voice conversion model, the sequence of source audio frames to generate a sequence of source internal representations characterizing the utterance spoken in the first accent; processing, using a decoder of the voice conversion model, the sequence of source internal representations to generate a sequence of target audio frames characterizing a synthesized speech representation of the utterance in a second accent different than the first accent; and providing, for output by a computing device, the synthesized speech representation of the utterance in the second accent.
2 . The computer-implemented method of claim 1 , wherein processing the sequence of source internal representations to generate a synthesized speech representation of the utterance in a second accent comprises processing the sequence of source internal representations to generate the synthesized speech representation without performing any speech recognition on the sequence of source audio frames.
3 . The computer-implemented method of claim 1 , wherein the sequence of source audio frames comprises a sequence of input spectrograms.
4 . The computer-implemented method of claim 1 , wherein the sequence of target audio frames comprises a sequence of output spectrograms.
5 . The computer-implemented method of claim 1 , wherein a cadence of the utterance spoken in the first accent is different than a cadence of the synthesized speech representation of the utterance in the second accent.
6 . The computer-implemented method of claim 1 , wherein the encoder comprises a bidirectional long short-term memory (LSTM) layer.
7 . The computer-implemented method of claim 1 , wherein the decoder comprises a spectrogram decoder with attention.
8 . The computer-implemented method of claim 1 , wherein the operations further comprise:
receiving audio data of a collection of utterances; obtaining a transcription of each utterance in the collection of utterances; providing the transcriptions of each utterance as an input to a text to speech model; receiving, for each transcription of each utterance, audio data of an additional collection of utterances in a synthesized voice; and training the model using the audio data of the collection of utterances and the audio data of an additional collection of utterances in a synthesized voice.
9 . The computer-implemented method of claim 1 , wherein the operations further comprise bypassing obtaining a transcription of the utterance.
10 . The computer-implemented method of claim 1 , wherein the speech conversion model is configured to adjust a time period between each term in the utterance spoken in the first accent.
11 . A system comprising:
data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:
receiving a sequence of source audio frames characterizing an utterance spoken in a first accent;
processing, using an encoder of a voice conversion model, the sequence of source audio frames to generate a sequence of source internal representations characterizing the utterance spoken in the first accent;
processing, using a decoder of the voice conversion model, the sequence of source internal representations to generate a sequence of target audio frames characterizing a synthesized speech representation of the utterance in a second accent different than the first accent; and
providing, for output by a computing device, the synthesized speech representation of the utterance in the second accent.
12 . The system of claim 11 , wherein processing the sequence of source internal representations to generate a synthesized speech representation of the utterance in a second accent comprises processing the sequence of source internal representations to generate the synthesized speech representation without performing any speech recognition on the sequence of source audio frames.
13 . The system of claim 11 , wherein the sequence of source audio frames comprises a sequence of input spectrograms.
14 . The system of claim 11 , wherein the sequence of target audio frames comprises a sequence of output spectrograms.
15 . The system of claim 11 , wherein a cadence of the utterance spoken in the first accent is different than a cadence of the synthesized speech representation of the utterance in the second accent.
16 . The system of claim 11 , wherein the encoder comprises a bidirectional long short-term memory (LSTM) layer.
17 . The system of claim 11 , wherein the decoder comprises a spectrogram decoder with attention.
18 . The system of claim 11 , wherein the operations further comprise:
receiving audio data of a collection of utterances; obtaining a transcription of each utterance in the collection of utterances; providing the transcriptions of each utterance as an input to a text to speech model; receiving, for each transcription of each utterance, audio data of an additional collection of utterances in a synthesized voice; and training the model using the audio data of the collection of utterances and the audio data of an additional collection of utterances in a synthesized voice.
19 . The system of claim 11 , wherein the operations further comprise bypassing obtaining a transcription of the utterance.
20 . The system of claim 11 , wherein the speech conversion model is configured to adjust a time period between each term in the utterance spoken in the first accent.Join the waitlist — get patent alerts
Track US2023230572A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.