US2023034572A1PendingUtilityA1

Voice synthesis method, voice synthesis apparatus, and recording medium

Assignee: YAMAHA CORPPriority: Nov 29, 2017Filed: Oct 13, 2022Published: Feb 2, 2023
Est. expiryNov 29, 2037(~11.3 yrs left)· nominal 20-yr term from priority
Inventors:Ryunosuke Daido
G06N 3/09G06N 3/0499G10L 13/033G10H 1/0066G06N 3/045G06N 3/08G10H 2250/455G10L 13/0335G10L 13/10G06N 3/088G10L 13/06
68
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Voice synthesis method and apparatus generate second control data using an intermediate trained model with first input data including first control data designating phonetic identifiers, change the second control data in accordance with a first user instruction provided by a user, generate synthesis data representing frequency characteristics of a voice to be synthesized using a final trained model with final input data including the first control data and the changed second control data, and generate a voice signal based on the generated synthesis data.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented voice synthesis method comprising:
 displaying a first image representing first control data that specifies lyrics along a time axis on a display;   generating second control data representing a series of phonemes according to the first control data;   displaying a second image representing the generated second control data along the time axis on the display;   changing the generated second control data in response to a first user instruction from a user; and   generating a voice signal of a synthesis voice in accordance with the first control data and the changed second control data.   
     
     
         2 . The voice synthesis method according to  claim 1 , wherein the second control data is generated by supplying the first control data as input to a first trained model. 
     
     
         3 . The voice synthesis method according to  claim 1 , further comprising:
 generating third control data representing musical expressions in accordance with the first control data and the changed second control data;   displaying a third image representing the generated third control data along the time axis on the display; and   changing the generated third control data in response to a second user instruction from the user,   wherein the voice signal is generated according to the first control data, the changed second control data, and the changed third control data.   
     
     
         4 . The voice synthesis method according to  claim 3 , wherein:
 the second control data is generated by supplying the first control data as input to a first trained model; and   the third control data is generated by supplying the first control data and the changed second control data as inputs to a second trained model.   
     
     
         5 . The voice synthesis method according to  claim 3 , wherein the voice signal is generated by:
 generating synthesis data representing frequency characteristics of the synthesis voice according to the first control data, the changed second control data, and the changed third control data;   displaying a fourth image representing the generated synthesis data along the time axis on the display;   changing the generated synthesis data in response to a third user instruction from the user; and   generating the voice signal according to the changed synthesis data.   
     
     
         6 . The voice synthesis method according to  claim 5 , wherein:
 the second control data is generated by supplying the first control data as input to a first trained model;   the third control data is generated by supplying the first control data and the changed second control data as inputs to a second trained model; and   the synthesis data is generated by supplying the first control data, the changed second control data, and the changed third control data as inputs to a third trained model.   
     
     
         7 . A voice synthesis system comprising:
 a display;   one or more memories for storing instructions; and   one or more processors communicatively connected to the display and the one or more memories and that execute the instructions to perform a plurality of tasks, including:   a first displaying task that displays a first image representing first control data that specifies lyrics along a time axis on the display;   a first generating task that generates second control data representing a series of phonemes according to the first control data;   a second displaying task that displays a second image representing the generated second control data along the time axis on the display;   a first changing task that changes the generated second control data in response to a first user instruction from a user; and   a second generating task that generates a voice signal of a synthesis voice in accordance with the first control data and the changed second control data.   
     
     
         8 . The voice synthesis system according to  claim 7 , wherein the first generating task generates the second control data by supplying the first control data as input to a first trained model. 
     
     
         9 . The voice synthesis system according to  claim 7 , wherein the plurality of tasks further includes:
 a third generating task that generates third control data representing musical expressions in accordance with the first control data and the changed second control data;   a third displaying task that displays a third image representing the generated third control data along the time axis on the display; and   a second changing task that changes the generated third control data in response to a second user instruction from the user,   wherein the second generating task generates the voice signal according to the first control data, the changed second control data, and the changed third control data.   
     
     
         10 . The voice synthesis system according to  claim 9 , wherein:
 the first generating task generates the second control data by supplying the first control data as input to a first trained model; and   the third generating task generates the third control data by supplying the first control data and the changed second control data as inputs to a second trained model.   
     
     
         11 . The voice synthesis system according to  claim 9 , wherein the second generating task further:
 generates synthesis data representing frequency characteristics of the synthesis voice according to the first control data, the changed second control data, and the changed third control data;   displays a fourth image representing the generated synthesis data along the time axis on the display;   changes the generated synthesis data in response to a third user instruction from the user; and   generates the voice signal according to the changed synthesis data.   
     
     
         12 . The voice synthesis system according to  claim 11 , wherein:
 the first generating task generates the second control data by supplying the first control data as input to a first trained model;   the third generating task generates the third control data by supplying the first control data and the changed second control data as inputs to a second trained model; and   the second generating task further generates the synthesis data by supplying the first control data, the changed second control data, and the changed third control data as inputs to a third trained model.

Join the waitlist — get patent alerts

Track US2023034572A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.