US2021142818A1PendingUtilityA1

System and method for animated lip synchronization

Assignee: GOVERNING COUNCIL UNIV TORONTOPriority: Mar 3, 2017Filed: Oct 20, 2020Published: May 13, 2021
Est. expiryMar 3, 2037(~10.6 yrs left)· nominal 20-yr term from priority
G10L 25/90G10L 2021/105G10L 2015/025G10L 21/10
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for animated lip synchronization. The method includes: capturing speech input; parsing the speech input into phenomes; aligning the phonemes to the corresponding portions of the speech input; mapping the phonemes to visemes; synchronizing the visemes into viseme action units, the viseme action units comprising jaw and lip contributions for each of the phonemes; and outputting the viseme action units.

Claims

exact text as granted — not AI-modified
1 . A method for animated lip synchronization executed on a processing unit, the method comprising:
 mapping phonemes to visemes;   synchronizing the visemes into viseme action units, the viseme action units comprising jaw and lip contributions for each of the phonemes; and   outputting the viseme action units.   
     
     
         2 . The method of  claim 1 , further comprising capturing speech input; parsing the speech input into the phonemes; and aligning the phonemes to the corresponding portions of the speech input. 
     
     
         3 . The method of  claim 2 , wherein aligning the phonemes comprises one or more of phoneme parsing and forced alignment. 
     
     
         4 . The method of  claim 1 , wherein two or more viseme action units are co-articulated such that the respective two or more visemes are approximately concurrent. 
     
     
         5 . The method of  claim 1 , wherein the jaw contributions and the lip contributions are respectively synchronized to independent visemes, and wherein the viseme action units are a linear combination of the independent visemes. 
     
     
         6 . The method of  claim 1 , wherein the jaw contributions and the lip contributions are each respectively synchronized to activations of one or more facial muscles in a biomechanical muscle model such that the viseme action units represent a dynamic simulation of the biomechanical muscle model. 
     
     
         7 . The method of  claim 1 , wherein mapping the phonemes to visemes comprises at least one of mapping a start time of at least one of the visemes to be prior to an end time of a previous respective viseme and mapping an end time of at least one of the visemes to be after a start time of a subsequent respective viseme. 
     
     
         8 . The method of  claim 1 , wherein a start time of at least one of the visemes is at least 120 ms before the respective phoneme is heard, and an end time of at least one of the visemes is at least 120 ms after the respective phoneme is heard. 
     
     
         9 . The method of  claim 1 , wherein a start time of at least one of the visemes is at least 150 ms before the respective phoneme is heard, and an end time of at least one of the visemes is at least 150 ms after the respective phoneme is heard. 
     
     
         10 . The method of  claim 1 , wherein viseme decay of at least one of the visemes begins between seventy-percent and eighty-percent of the completion of the respective phoneme. 
     
     
         11 . The method of  claim 1 , wherein an amplitude of each viseme is determined by one or more of lexical stress and word prominence. 
     
     
         12 . The method of  claim 1 , wherein the viseme action units further comprise tongue contributions for each of the phonemes. 
     
     
         13 . The method of  claim 1 , wherein the viseme action unit for a neutral pose comprises a viseme mapped to a bilabial phoneme. 
     
     
         14 . The method of  claim 1 , further comprising outputting a phonetic animation curve based on the change of viseme action units over time. 
     
     
         15 . A system for animated lip synchronization, the system having one or more processors and a data storage device, the one or more processors in communication with the data storage device, the one or more processors configured to execute:
 a correspondence module for mapping phonemes to visemes;   a synchronization module for synchronizing the visemes into viseme action units, the viseme action units comprising jaw and lip contributions for each of the phonemes; and   an output module for outputting the viseme action units to an output device.   
     
     
         16 . The system of  claim 15  further comprising an input module for capturing speech input received from an input device, the input module parsing the speech input into the phenomes; and an alignment module for aligning the phonemes to the corresponding portions of the speech input. 
     
     
         17 . The system of  claim 15  further comprising a speech analyzer module for analyzing one or more of pitch and intensity of the speech input. 
     
     
         18 . The system of  claim 15 , wherein the alignment module aligns the phonemes by at least one of phoneme parsing and forced alignment. 
     
     
         19 . The system of  claim 15 , wherein the output module further outputs a phonetic animation curve based on the change of viseme action units over time. 
     
     
         20 . A facial model for animation on a computing device, the computing device having one or more processors, the facial model comprising: a neutral face position; an overlay of skeletal jaw deformation, lip deformation and tongue deformation; and a displacement of the skeletal jaw deformation, the lip deformation and the tongue deformation by a linear blend of weighted blend-shape action units.

Join the waitlist — get patent alerts

Track US2021142818A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.