US2018336716A1PendingUtilityA1

Voice effects based on facial expressions

Assignee: APPLE INCPriority: May 16, 2017Filed: Feb 28, 2018Published: Nov 22, 2018
Est. expiryMay 16, 2037(~10.8 yrs left)· nominal 20-yr term from priority
G06F 18/00H04N 23/611H04N 23/63H04M 2250/52G06T 13/40H04L 51/04G06F 3/012G06F 3/0484H04L 51/10G06F 3/04886G06F 3/04842G10L 15/02G06K 9/00315G06T 13/80G06F 3/167H04L 51/52G06V 40/175H04L 51/08G06V 40/176G06V 20/20H04L 51/58G06F 3/0304H04M 1/72439H04M 1/72436H04W 4/12
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Embodiments of the present disclosure can provide systems, methods, and computer-readable medium for adjusting audio and/or video information of a video clip based at least in part on facial feature and/or voice feature characteristics extracted from hardware components. For example, in response to detecting a request to generate an avatar video clip of a virtual avatar, a video signal associated with a face in a field of view of a camera and an audio signal may be captured. Voice feature characteristics and facial feature characteristics may be extracted from the audio signal and the video signal, respectively. In some examples, in response to detecting a request to preview the avatar video clip, an adjusted audio signal may be generated based at least in part on the facial feature characteristics and the voice feature characteristics, and a preview of the video clip of the virtual avatar using the adjusted audio signal may be displayed.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 at an electronic device having at least a camera and a microphone:
 displaying a virtual avatar generation interface; 
 displaying first preview content of a virtual avatar in the virtual avatar generation interface, the first preview content of the virtual avatar corresponding to realtime preview video frames of a user headshot in a field of view of the camera and associated headshot changes in an appearance; 
 while displaying the first preview content of the virtual avatar, detecting an input in the virtual avatar generation interface; 
 in response to detecting the input in the virtual avatar generation interface:
 capturing, via the camera, a video signal associated with the user headshot during a recording session; 
 capturing, via the microphone, a user audio signal during the recording session; 
 extracting audio feature characteristics from the captured user audio signal; and 
 extracting facial feature characteristics associated with the face from the captured video signal; and 
 
 in response to detecting expiration of the recording session:
 generating an adjusted audio signal from the captured audio signal based at least in part on the facial feature characteristics and the audio feature characteristics; 
 generating second preview content of the virtual avatar in the virtual avatar generation interface according to the facial feature characteristics and the adjusted audio signal; and 
 presenting the second preview content in the virtual avatar generation interface. 
 
   
     
     
         2 . The method of  claim 1 , further comprising storing facial feature metadata associated with the facial feature characteristics extracted from the video signal and strong audio metadata associated with the audio feature characteristics extracted from the audio signal. 
     
     
         3 . The method of  claim 2 , further comprising generating adjusted facial feature metadata from the facial feature metadata based at least in part on the facial feature characteristics and the audio feature characteristics. 
     
     
         4 . The method of  claim 3 , wherein the second preview of the virtual avatar is displayed further according to the adjusted facial metadata. 
     
     
         5 . An electronic device, comprising:
 a camera;   a microphone; and   one or more processors in communication with the camera and the microphone, the one or more processors configured to:
 while displaying a first preview of a virtual avatar, detecting an input in a virtual avatar generation interface; 
 in response to detecting the input in the virtual avatar generation interface, initiating a capture session including:
 capturing, via the camera, a video signal associated with a face in a field of view of the camera; 
 capturing, via the microphone, an audio signal associated with the captured video signal; 
 extracting audio feature characteristics from the captured audio signal; and 
 extracting facial feature characteristics associated with the face from the captured video signal; and 
 
 in response to detecting expiration of the capture session:
 generating an adjusted audio signal based at least in part on the audio feature characteristics and the facial feature characteristics; and 
 displaying a second preview of the virtual avatar in the virtual avatar generation interface according to the facial feature characteristics and the adjusted audio signal. 
 
   
     
     
         6 . The electronic device of  claim 5 , wherein the audio signal is further adjusted based at least in part on a type of the virtual avatar. 
     
     
         7 . The electronic device of  claim 6 , wherein the type of the virtual avatar is received based at least in part on an avatar type selection affordance presented in the virtual avatar generation interface. 
     
     
         8 . The electronic device of  claim 6 , wherein the type of the virtual avatar includes an animal type, and wherein the adjusted audio signal is generated based at least in part on a predetermined sound associated with the animal type. 
     
     
         9 . The electronic device of  claim 5 , wherein the one or more processors are further configured to determine whether a portion of the audio signal corresponds to the face in the field of view. 
     
     
         10 . The electronic device of  claim 9 , wherein the one or more processors are further configured to, in accordance with a determination that the portion of the audio signal corresponds to the face, store the portion of the audio signal for use in generating the adjusted audio signal. 
     
     
         11 . The electronic device of  claim 9 , wherein the one or more processors are further configured to, in accordance with a determination that the portion of the audio signal does not correspond to the face, discard at least the portion of the audio signal. 
     
     
         12 . The electronic device of  claim 5 , wherein the audio feature characteristics comprise features of a voice associated with the face in the field of view. 
     
     
         13 . The electronic device of  claim 5 , wherein the one or more processors are further configured to store facial feature metadata associated with the facial feature characteristics extracted from the video signal. 
     
     
         14 . The electronic device of  claim 13 , wherein the one or more processors are further configured to generate adjusted facial metadata based at least in part on the facial feature characteristics and the audio feature characteristics. 
     
     
         15 . The electronic device of  claim 14 , wherein the second preview of the virtual avatar is generated according to the adjusted facial metadata and the adjusted audio signal. 
     
     
         16 . A computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, configure the one or more processors to perform operations comprising:
 in response to detecting a request to generate an avatar video clip of a virtual avatar:
 capturing, via a camera of an electronic device, a video signal associated with a face in a field of view of the camera; 
 capturing, via a microphone of the electronic device, an audio signal; 
 extracting voice feature characteristics from the captured audio signal; and 
 extracting facial feature characteristics associated with the face from the captured video signal; and 
   in response to detecting a request to preview the avatar video clip:
 generating an adjusted audio signal based at least in part on the facial feature characteristics and the voice feature characteristics; and 
 displaying a preview of the video clip of the virtual avatar using the adjusted audio signal. 
   
     
     
         17 . The computer-readable storage medium of  claim 16 , wherein the audio signal is adjusted based at least in part on a facial expression identified in the facial feature characteristics associated with the face. 
     
     
         18 . The computer-readable storage medium of  claim 16 , wherein the adjusted audio signal is further adjusted by inserting one or more pre-stored audio samples. 
     
     
         19 . The computer-readable storage medium of  claim 16 , wherein the audio signal is adjusted based at least in part on a level, pitch, duration, variable playback speed, speech spectral-format positions, speech spectral-format-levels, instantaneous playback speed, or change in a voice associated with the face. 
     
     
         20 . The computer-readable storage medium of  claim 16 , wherein the one or more processors are further configured to perform the operations comprising transmitting the video clip of the virtual avatar to another electronic device.

Join the waitlist — get patent alerts

Track US2018336716A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.