US11659193B2ActiveUtilityA1

Framework for video conferencing based on face restoration

Assignee: Tencent America LLCPriority: Jan 6, 2021Filed: Sep 30, 2021Granted: May 23, 2023
Est. expiryJan 6, 2041(~14.5 yrs left)· nominal 20-yr term from priority
G06N 3/0475G06N 3/0464G06N 3/094G06N 3/0455G06N 3/09G06N 3/0495H04N 7/15H04N 19/85G06V 40/161H04N 19/20G06T 2207/10016G06T 2207/30201H04N 19/593G06N 3/047G06T 7/73H04N 19/14G06V 10/993H04N 19/59G06N 3/088G06V 40/165G06N 3/08H04N 19/132G06N 3/045H04N 19/423G06T 7/62G06T 9/002G06T 2207/20084H04N 19/70H04N 19/27G06V 40/171H04N 19/17H04N 19/29H04N 19/30H04N 19/184G06V 40/168G06T 3/40G06V 10/82
58
PatentIndex Score
0
Cited by
9
References
8
Claims

Abstract

There is included a method and apparatus comprising computer code configured to cause a processor or processors to perform obtaining video data, detecting at least one face from at least one frame of the video data, determining a set of facial landmark features of the at least one face from the at least one frame of the video data, and coding the video data at least partly by a neural network based on the determined set of facial landmark features.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A method for video coding performed by at least one processor, the method comprising:
 obtaining video data; 
 detecting at least one face from at least one frame of the video data; 
 determining a set of facial landmark features of the at least one face from the at least one frame of the video data; 
 determining an extended face area (EFA) which comprises a boundary area extended from an area of the detected at least one face from the at least one frame of the video data; 
 determining a set of EFA features from the EFA; and 
 coding the video data at least partly by a neural network based on the determined set of facial landmark features and on aggregating the set of facial landmark features, reconstructed EFA features, and an up-sampled sequence that is up-sampled from at least one down-sampled sequence, 
 wherein the video data comprises an encoded bitstream of the video data, 
 wherein determining the set of facial landmark features comprises up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, 
 wherein determining the EFA and determining the set of EFA features comprise up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, and 
 wherein determining the EFA and determining the set of EFA features further comprise reconstructing the EFA features, into the reconstructed EFA features, each respective to ones of the facial landmark features of the set of facial landmark features by a generative adversarial network. 
 
     
     
       2. The method according to  claim 1 ,
 wherein the at least one face from the at least one frame of the video data is determined to be a largest face among a plurality of faces in the at least one frame of the video data. 
 
     
     
       3. The method according  claim 1 , further comprising:
 determining a plurality of sets of facial landmark features, other than the set of facial landmark features of the at least one face from the at least one frame of the video data, respect to each of the plurality of faces in the at least one frame of the video data; and 
 coding the video data at least partly by the neural network based on the determined set of facial landmark features and the determined plurality of sets of facial landmark features. 
 
     
     
       4. The method according to  claim 1 ,
 wherein the neural network comprises a deep neural network (DNN). 
 
     
     
       5. An apparatus for video coding, the apparatus comprising:
 at least one memory configured to store computer program code; 
 at least one processor configured to execute the computer program code to implement:
 obtaining video data; 
 detecting at least one face from at least one frame of the video data; 
 determining a set of facial landmark features of the at least one face from the at least one frame of the video data; 
 determining an extended face area (EFA) which comprises a boundary area extended from an area of the detected at least one face from the at least one frame of the video data; 
 determining a set of EFA features from the EFA; and 
 coding the video data at least partly by a neural network based on the determined set of facial landmark features and on aggregating the set of facial landmark features, reconstructed EFA features, and an up-sampled sequence that is up-sampled from at least one down-sampled sequence, 
 
 wherein the video data comprises an encoded bitstream of the video data, 
 wherein determining the set of facial landmark features comprises up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, 
 wherein determining the EFA and determining the set of EFA features comprise up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, and 
 wherein determining the EFA and determining the set of EFA features further comprise reconstructing the EFA features, into the reconstructed EFA features, each respective to ones of the facial landmark features of the set of facial landmark features by a generative adversarial network. 
 
     
     
       6. The apparatus according to  claim 5 ,
 wherein the at least one face from the at least one frame of the video data is determined to be a largest face among a plurality of faces in the at least one frame of the video data. 
 
     
     
       7. The apparatus according to  claim 5 ,
 wherein the at least one hardware processor is further configured to execute the computer program code to implement:
 determining a plurality of sets of facial landmark features, other than the set of facial landmark features of the at least one face from the at least one frame of the video data, respect to each of the plurality of faces in the at least one frame of the video data; and 
 coding the video data at least partly by the neural network based on the determined set of facial landmark features and the determined plurality of sets of facial landmark features. 
 
 
     
     
       8. A non-transitory computer readable medium storing a program causing a computer to execute a process, the process comprising:
 obtaining video data; 
 detecting at least one face from at least one frame of the video data; 
 determining a set of facial landmark features of the at least one face from the at least one frame of the video data; 
 determining an extended face area (EFA) which comprises a boundary area extended from an area of the detected at least one face from the at least one frame of the video data; 
 determining a set of EFA features from the EFA; and 
 coding the video data at least partly by a neural network based on the determined set of facial landmark features and on aggregating the set of facial landmarks, reconstructed EFA features, and an up-sampled sequence that is up-sampled from at least one down-sampled sequence, 
 wherein the video data comprises an encoded bitstream of the video data, 
 wherein determining the set of facial landmark features comprises up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, 
 wherein determining the EFA and determining the set of EFA features comprise up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, and 
 wherein determining the EFA and determining the set of EFA features further comprise reconstructing the EFA features, into the reconstructed EFA features, each respective to ones of the facial landmark features of the set of facial landmark features by a generative adversarial network.

Join the waitlist — get patent alerts

Track US11659193B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.