US2024038213A1PendingUtilityA1

Generating method, generating device, and generating program

Assignee: NIPPON TELEGRAPH & TELEPHONEPriority: Nov 25, 2020Filed: Nov 25, 2020Published: Feb 1, 2024
Est. expiryNov 25, 2040(~14.3 yrs left)· nominal 20-yr term from priority
Inventors:Hiroki Kanagawa
G10L 13/047G10L 13/06G10L 25/30
26
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A generation device ( 100 ) extracts a plurality of integrated speech samples by repeatedly executing processing of integrating a plurality of consecutive speech samples included in speech waveform information into one speech sample, and generates a compressed speech sample by compressing the plurality of integrated speech samples extracted. The generation device ( 100 ) generates a plurality of new integrated speech samples subsequent to the plurality of integrated speech samples by inputting the compressed speech sample and an acoustic feature value calculated from the speech waveform information to a speech waveform generation model, and repeatedly executes processing of inputting a compressed speech sample obtained by compressing the plurality of new integrated speech samples and the acoustic feature value to the speech waveform generation model, to generate a plurality of new integrated speech samples a plurality of times.

Claims

exact text as granted — not AI-modified
1 . A generation method comprising:
 extracting a plurality of integrated speech samples, wherein the extracting the plurality of integrated speech samples comprises iteratively performing:
 integrating a plurality of consecutive speech samples extracted from speech waveform information into one speech sample, and 
   compressing the plurality of integrated speech samples in the one speech sample to generate a compressed speech sample; and   generating a plurality of new integrated speech samples subsequent to the plurality of integrated speech samples, wherein the generating the plurality of new integrated speech samples comprises iteratively performing:
 inputting the compressed speech sample and an acoustic feature value calculated from the speech waveform information to a speech waveform generation model, and 
 compressing the plurality of new integrated speech samples and the acoustic feature value to the speech waveform generation model. 
   
     
     
         2 . The generation method according to  claim 1 , wherein the speech waveform generation model outputs a probability value associated with an amplitude of a speech waveform at each of times based on the compressed speech sample and the acoustic feature value as input to the speech waveform generation model, and the generating further comprises generating the plurality of new integrated speech samples based on the probability value associated with the amplitude of the speech waveform at each of the times. 
     
     
         3 . The generation method according to  claim 2 , wherein the generating further comprises learning the speech waveform generation model based on a loss value between the probability value and the speech waveform information. 
     
     
         4 . The generation method according to  claim 3 , further comprising:
 iteratively processing:
 generating a plurality of new integrated speech samples by inputting a combination including the compressed speech sample and a specified acoustic feature value to a learning model; and 
 combining the plurality of integrated speech samples. 
   
     
     
         5 . The generation method according to  claim 3 , further comprising:
 learning a down-sampling model, wherein the down-sampling model outputs the compressed speech sample based on the loss value according to the plurality of integrated speech samples as input.   
     
     
         6 . The generation method according to  claim 3 , further comprising:
 learning a down-sampling model, wherein the down-sampling model outputs the compressed speech sample and a down-sampled acoustic feature value based on the loss value according to the plurality of integrated speech samples and the acoustic feature value as input.   
     
     
         7 . A generation device comprising a processor configured to execute operations comprising:
 extracting a plurality of integrated speech samples, wherein the extracting the plurality of integrated speech samples comprises iteratively performing:
 integrating a plurality of consecutive speech samples included in speech waveform information into one speech sample, and 
 generating a compressed speech sample by compressing the plurality of integrated speech samples to generate a compressed speech sample; and 
   generating a plurality of new integrated speech samples subsequent to the plurality of integrated speech samples, wherein the generating the plurality of new integrated speech samples comprises iteratively performing:
 inputting the compressed speech sample and an acoustic feature value calculated from the speech waveform information to a speech waveform generation model, and 
 compressing the plurality of new integrated speech samples and the acoustic feature value to the speech waveform generation model. 
   
     
     
         8 . A computer-readable non-transitory recording medium storing computer-executable program instructions that when executed by a processor cause a computer system to execute operations comprising:
 extracting a plurality of integrated speech samples, wherein the extracting the plurality of integrated speech samples comprises iteratively performing:
 integrating a plurality of consecutive speech samples included in speech waveform information into one speech sample, and 
 generating a compressed speech sample by compressing the plurality of integrated speech samples to generate a compressed speech sample; and 
   generating a plurality of new integrated speech samples subsequent to the plurality of integrated speech samples, wherein the generating the plurality of new integrated speech samples comprises iteratively performing:
 inputting the compressed speech sample and an acoustic feature value calculated from the speech waveform information to a speech waveform generation model, and 
 compressing the plurality of new integrated speech samples and the acoustic feature value to the speech waveform generation model. 
   
     
     
         9 . The generation device according to  claim 7 , wherein the speech waveform generation model outputs a probability value associated with an amplitude of a speech waveform at each of times based on the compressed speech sample and the acoustic feature value as input to the speech waveform generation model, and the generating further comprises generating the plurality of new integrated speech samples based on the probability value associated with the amplitude of the speech waveform at each of the times. 
     
     
         10 . The generation device according to  claim 9 , wherein the generating further comprises learning the speech waveform generation model based on a loss value between the probability value and the speech waveform information. 
     
     
         11 . The generation device according to  claim 10 , the processor further configured to execute operations comprising:
 iteratively processing:
 generating a plurality of new integrated speech samples by inputting a combination including the compressed speech sample and a specified acoustic feature value to a learning model; and 
 combining the plurality of integrated speech samples. 
   
     
     
         12 . The generation device according to  claim 10 , the processor further configured to execute operations comprising:
 learning a down-sampling model, wherein the down-sampling model outputs the compressed speech sample based on the loss value according to the plurality of integrated speech samples as input.   
     
     
         13 . The generation device according to  claim 10 , the processor further configured to execute operations comprising:
 learning a down-sampling model, wherein the down-sampling model outputs the compressed speech sample and a down-sampled acoustic feature value based on the loss value according to the plurality of integrated speech samples and the acoustic feature value as input.   
     
     
         14 . The computer-readable non-transitory recording medium according to  claim 8 , wherein the speech waveform generation model outputs a probability value associated with an amplitude of a speech waveform at each of times based on the compressed speech sample and the acoustic feature value as input to the speech waveform generation model, and the generating further comprises generating the plurality of new integrated speech samples based on the probability value associated with the amplitude of the speech waveform at each of the times. 
     
     
         15 . The computer-readable non-transitory recording medium according to  claim 14 , wherein the generating further comprises learning the speech waveform generation model based on a loss value between the probability value and the speech waveform information. 
     
     
         16 . The computer-readable non-transitory recording medium according to  claim 15 , the computer-executable program instructions when executed further causing the computer system to execute operations comprising:
 iteratively processing:
 generating a plurality of new integrated speech samples by inputting a combination including the compressed speech sample and a specified acoustic feature value to a learning model; and 
 combining the plurality of integrated speech samples. 
   
     
     
         17 . The computer-readable non-transitory recording medium according to  claim 15 , the computer-executable program instructions when executed further causing the computer system to execute operations comprising:
 learning a down-sampling model, wherein the down-sampling model outputs the compressed speech sample based on the loss value according to the plurality of integrated speech samples as input.   
     
     
         18 . The computer-readable non-transitory recording medium according to  claim 15 , the computer-executable program instructions when executed further causing the computer system to execute operations comprising:
 learning a down-sampling model, wherein the down-sampling model outputs the compressed speech sample and a down-sampled acoustic feature value based on the loss value according to the plurality of integrated speech samples and the acoustic feature value as input.

Join the waitlist — get patent alerts

Track US2024038213A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.