Text-to-Speech Adapted by Machine Learning
Abstract
Machine learned models take in vectors representing desired behaviors and generate voice vectors that provide the parameters for text-to-speech (TTS) synthesis. Models may be trained on behavior vectors that include user profile attributes, situational attributes, or semantic attributes. Situational attributes may include age of people present, music that is playing, location, noise, and mood. Semantic attributes may include presence of proper nouns, number of modifiers, emotional charge, and domain of discourse. TTS voice parameters may apply per utterance and per word as to enable contrastive emphasis.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of speech synthesis, the method comprising:
processing a sensor signal of a vehicle; producing a TTS prosody parameter according to a model; synthesizing digital audio samples of speech, an attribute of which depends upon the TTS prosody parameter; and driving a speaker to produce audio as represented by the digital audio samples, wherein prosody parameter can change at run time for more dynamic effects.
2 . The method of claim 1 wherein:
processing the sensor signal determines a value of a situational attribute; and
producing the TTS prosody parameter is in dependence upon the value of the situational attribute.
3 . The method of claim 2 wherein the dependence upon the value of the situational attribute is programmable using text rules.
4 . A method of speech synthesis, the method comprising:
processing a sensor signal of a vehicle to determine a value of a situational attribute; producing a TTS parameter according to a model in dependence upon a value of the situational attribute; synthesizing digital audio samples of speech based on input text such that an attribute of the digital audio samples depends upon the TTS parameter; and driving a speaker to produce audio as represented by the digital audio samples.
5 . The method of claim 4 wherein the dependence upon the value of the situational attribute is programmable using text rules.
6 . The method of claim 4 wherein the TTS parameter represents a prosody attribute.
7 . The method of claim 6 wherein prosody can be changed at run time for more dynamic effects.
8 . The method of claim 6 , further comprising synthesizing the digital audio samples of speech such that prosody attribute is further based on markup in the input text.
9 . A method of speech synthesis, the method comprising:
processing a sensor signal of a vehicle; producing a TTS parameter according to a model; synthesizing digital audio samples of speech, an attribute of which depends upon the TTS parameter; and driving a speaker to produce audio as represented by the digital audio samples.
10 . The method of claim 9 wherein:
processing the sensor signal determines a value of a situational attribute; and
producing the TTS parameter is in dependence upon the value of the situational attribute.
11 . The method of claim 10 wherein the dependence upon the value of the situational attribute is programmable using text rules.
12 . The method of claim 9 wherein the TTS parameter represents a prosody attribute.
13 . The method of claim 12 wherein prosody can change at run time for more dynamic effects.
14 . A text-to-speech system comprising a computer processor programmed to:
perform machine learned parametric speech synthesis using a TTS voice parameter; and produce the TTS voice parameter by a function that transforms at least one voice attribute and at least one situational attribute according to a model, wherein the model has a specified rule-based algorithm coded with a user text rule.
15 . The text-to-speech system of claim 14 wherein the user text rule depends on situational attributes.
16 . The text-to-speech system of claim 14 wherein a situational attribute is noise level.
17 . The text-to-speech system of claim 14 wherein the user text rule depends on noise level.
18 . The text-to-speech system of claim 14 wherein the TTS voice parameter is related to volume.
19 . The text-to-speech system of claim 14 wherein:
the at least one situational attribute is a noise level;
the TTS voice parameter is related to volume;
the user text rule depends on a value of the at least one situational attribute; and
the value of the at least one situational attribute is received from a sensor in a vehicle.Join the waitlist — get patent alerts
Track US2022148566A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.