US2023352043A1PendingUtilityA1
Systems And Methods For Machine-Generated Avatars
Est. expiryOct 29, 2035(~9.2 yrs left)· nominal 20-yr term from priority
Inventors:Wayne Scholar
G10L 21/10G06T 13/40G10L 13/08G10L 15/02G10L 15/063G10L 15/187G10L 15/25G10L 2015/025G10L 2021/105
64
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and methods are disclosed for creating a machine generated avatar. A machine generated avatar is an avatar generated by processing video and audio information extracted from a recording of a human speaking a reading corpora and enabling the created avatar to be able to say an unlimited number of utterances, i.e., utterances that were not recorded. The video and audio processing consists of the use of machine learning algorithms that may create predictive models based upon pixel, semantic, phonetic, intonation, and wavelets.
Claims
exact text as granted — not AI-modified1 . A method for generating viseme and phoneme sequences, comprising:
receiving recording data comprising video recording data and audio recording data, wherein the recording data is associated with an input utterance; timestamping the recording data; extracting phoneme clips and viseme clips from the timestamped recording data; extracting transition light cones based on the viseme clips; extracting audio clips from the phoneme clips; associating the transition light cones with the audio clips; parsing the audio clips into sentences; and tagging the sentences with parts of speech.
2 . The method of claim 1 , wherein timestamping the recording data comprises:
separating the video recording data and the audio recording data; timestamping frames of the video recording data; timestamping the audio recording data; and processing the audio recording data with corpus associated with the input utterance to generate one or more timestamped phonemes.
3 . The method of claim 1 , wherein the input utterance is indicative of an actor reading a corpus of words.
4 . The method of claim 1 , wherein the phoneme clips comprise an audio file, and the viseme clips comprise an image file from the timestamped recording data.
5 . The method of claim 4 , wherein the audio file comprises at least one .wav file and the image file comprises at least one .jpg files.
6 . The method of claim 1 , further comprising: training a machine learning model based on at least one of: the sentences, the transition light cones, and the audio clips.
7 . The method of claim 6 , further comprising receiving a second input utterance and applying the trained machine learning model to generate an avatar providing audiovisual output corresponding to the second input utterance.
8 . The method of claim 7 , wherein generating the avatar further comprises applying an intonation model and parts-of-speech tagging information to generate the audiovisual output.
9 . A device for generating viseme and phoneme sequences, comprising:
one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the device to:
receive recording data comprising video recording data and audio recording data, wherein the recording data is associated with an input utterance;
timestamp the recording data;
extract phoneme clips and viseme clips from the timestamped recording data;
extract transition light cones based on the viseme clips;
extract audio clips from the phoneme clips;
associate the transition light cones with the audio clips;
parse the audio clips into sentences; and
tag the sentences with parts of speech.
10 . The device of claim 9 , wherein timestamping the recording data comprises:
separating the video recording data and the audio recording data; timestamping frames of the video recording data; timestamping the audio recording data; and processing the audio recording data with corpus associated with the input utterance to generate one or more timestamped phonemes.
11 . The device of claim 9 , wherein the input utterance is indicative of an actor reading a corpus of words.
12 . The device of claim 9 , wherein the phoneme clips comprise an audio file, and the viseme clips comprise an image file from the timestamped recording data.
13 . The device of claim 12 , wherein the audio file comprises at least one .wav file and the image file comprises at least one .jpg files.
14 . The device of claim 9 , further comprising: training a machine learning model based on at least one of: the sentences, the transition light cones, and the audio clips.
15 . A system for generating viseme and phoneme sequences, the system comprising
a recording device; and a computing device configured to:
receive, from the recording device, recording data comprising video recording data and audio recording data, wherein the recording data is associated with an input utterance;
timestamp the recording data;
extract phoneme clips and viseme clips from the timestamped recording data;
extract transition light cones based on the viseme clips;
extract audio clips from the phoneme clips;
associate the transition light cones with the audio clips;
parse the audio clips into sentences; and
tag the sentences with parts of speech.
16 . The system of claim 15 , wherein the instructions to timestamp the recording data further comprise:
separating the video recording data and the audio recording data; timestamping frames of the video recording data; timestamping the audio recording data; and processing the audio recording data with corpus associated with the input utterance to generate one or more timestamped phonemes.
17 . The system of claim 15 , wherein the input utterance is indicative of an actor reading a corpus of words.
18 . The system of claim 15 , wherein the phoneme clips comprise an audio file, and the viseme clips comprise an image file from the timestamped recording data.
19 . The system of claim 18 , wherein the audio file comprises at least one .wav file and the image file comprises at least one .jpg files.
20 . The system of claim 15 , further comprising: training a machine learning model based on at least one of: the sentences, the transition light cones, and the audio clips.Join the waitlist — get patent alerts
Track US2023352043A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.