US11659193B2ActiveUtilityA1
Framework for video conferencing based on face restoration
Est. expiryJan 6, 2041(~14.5 yrs left)· nominal 20-yr term from priority
G06N 3/0475G06N 3/0464G06N 3/094G06N 3/0455G06N 3/09G06N 3/0495H04N 7/15H04N 19/85G06V 40/161H04N 19/20G06T 2207/10016G06T 2207/30201H04N 19/593G06N 3/047G06T 7/73H04N 19/14G06V 10/993H04N 19/59G06N 3/088G06V 40/165G06N 3/08H04N 19/132G06N 3/045H04N 19/423G06T 7/62G06T 9/002G06T 2207/20084H04N 19/70H04N 19/27G06V 40/171H04N 19/17H04N 19/29H04N 19/30H04N 19/184G06V 40/168G06T 3/40G06V 10/82
58
PatentIndex Score
0
Cited by
9
References
8
Claims
Abstract
There is included a method and apparatus comprising computer code configured to cause a processor or processors to perform obtaining video data, detecting at least one face from at least one frame of the video data, determining a set of facial landmark features of the at least one face from the at least one frame of the video data, and coding the video data at least partly by a neural network based on the determined set of facial landmark features.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A method for video coding performed by at least one processor, the method comprising:
obtaining video data;
detecting at least one face from at least one frame of the video data;
determining a set of facial landmark features of the at least one face from the at least one frame of the video data;
determining an extended face area (EFA) which comprises a boundary area extended from an area of the detected at least one face from the at least one frame of the video data;
determining a set of EFA features from the EFA; and
coding the video data at least partly by a neural network based on the determined set of facial landmark features and on aggregating the set of facial landmark features, reconstructed EFA features, and an up-sampled sequence that is up-sampled from at least one down-sampled sequence,
wherein the video data comprises an encoded bitstream of the video data,
wherein determining the set of facial landmark features comprises up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream,
wherein determining the EFA and determining the set of EFA features comprise up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, and
wherein determining the EFA and determining the set of EFA features further comprise reconstructing the EFA features, into the reconstructed EFA features, each respective to ones of the facial landmark features of the set of facial landmark features by a generative adversarial network.
2. The method according to claim 1 ,
wherein the at least one face from the at least one frame of the video data is determined to be a largest face among a plurality of faces in the at least one frame of the video data.
3. The method according claim 1 , further comprising:
determining a plurality of sets of facial landmark features, other than the set of facial landmark features of the at least one face from the at least one frame of the video data, respect to each of the plurality of faces in the at least one frame of the video data; and
coding the video data at least partly by the neural network based on the determined set of facial landmark features and the determined plurality of sets of facial landmark features.
4. The method according to claim 1 ,
wherein the neural network comprises a deep neural network (DNN).
5. An apparatus for video coding, the apparatus comprising:
at least one memory configured to store computer program code;
at least one processor configured to execute the computer program code to implement:
obtaining video data;
detecting at least one face from at least one frame of the video data;
determining a set of facial landmark features of the at least one face from the at least one frame of the video data;
determining an extended face area (EFA) which comprises a boundary area extended from an area of the detected at least one face from the at least one frame of the video data;
determining a set of EFA features from the EFA; and
coding the video data at least partly by a neural network based on the determined set of facial landmark features and on aggregating the set of facial landmark features, reconstructed EFA features, and an up-sampled sequence that is up-sampled from at least one down-sampled sequence,
wherein the video data comprises an encoded bitstream of the video data,
wherein determining the set of facial landmark features comprises up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream,
wherein determining the EFA and determining the set of EFA features comprise up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, and
wherein determining the EFA and determining the set of EFA features further comprise reconstructing the EFA features, into the reconstructed EFA features, each respective to ones of the facial landmark features of the set of facial landmark features by a generative adversarial network.
6. The apparatus according to claim 5 ,
wherein the at least one face from the at least one frame of the video data is determined to be a largest face among a plurality of faces in the at least one frame of the video data.
7. The apparatus according to claim 5 ,
wherein the at least one hardware processor is further configured to execute the computer program code to implement:
determining a plurality of sets of facial landmark features, other than the set of facial landmark features of the at least one face from the at least one frame of the video data, respect to each of the plurality of faces in the at least one frame of the video data; and
coding the video data at least partly by the neural network based on the determined set of facial landmark features and the determined plurality of sets of facial landmark features.
8. A non-transitory computer readable medium storing a program causing a computer to execute a process, the process comprising:
obtaining video data;
detecting at least one face from at least one frame of the video data;
determining a set of facial landmark features of the at least one face from the at least one frame of the video data;
determining an extended face area (EFA) which comprises a boundary area extended from an area of the detected at least one face from the at least one frame of the video data;
determining a set of EFA features from the EFA; and
coding the video data at least partly by a neural network based on the determined set of facial landmark features and on aggregating the set of facial landmarks, reconstructed EFA features, and an up-sampled sequence that is up-sampled from at least one down-sampled sequence,
wherein the video data comprises an encoded bitstream of the video data,
wherein determining the set of facial landmark features comprises up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream,
wherein determining the EFA and determining the set of EFA features comprise up-sampling the at least one down-sampled sequence obtained by decompressing the encoded bitstream, and
wherein determining the EFA and determining the set of EFA features further comprise reconstructing the EFA features, into the reconstructed EFA features, each respective to ones of the facial landmark features of the set of facial landmark features by a generative adversarial network.Join the waitlist — get patent alerts
Track US11659193B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.