US2025200701A1PendingUtilityA1
Method and apparatus with high-resolution image restoration
Assignee: SAMSUNG ELECTRONICS CO LTDPriority: Dec 15, 2023Filed: Dec 10, 2024Published: Jun 19, 2025
Est. expiryDec 15, 2043(~17.4 yrs left)· nominal 20-yr term from priority
G06T 2207/20056G06T 2207/20084G06T 2207/20081G06T 5/60G06F 40/20G06V 10/42G06T 5/70G06T 3/4076G06T 3/4046G06T 5/73G06V 10/771G06T 3/4053G06T 7/40G06T 5/10G06T 3/40
53
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A processor-implemented method with high-resolution (HR) image restoration includes mapping an input image with low resolution (LR) to a feature map of a latent domain, generating, using a diffusion model, an HR feature in which a certain frequency component corresponding to the feature map is restored, and based on the HR feature and coordinate information and pixel information of an HR image at a target magnification, restoring the HR image at the target magnification corresponding to the feature map.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor-implemented method with high-resolution (HR) image restoration, the method comprising:
mapping an input image with low resolution (LR) to a feature map of a latent domain; generating, using a diffusion model, an HR feature in which a certain frequency component corresponding to the feature map is restored; and based on the HR feature and coordinate information and pixel information of an HR image at a target magnification, restoring the HR image at the target magnification corresponding to the feature map.
2 . The method of claim 1 , wherein the mapping of the input image to the feature map of the latent domain comprises mapping the input image to the feature map of the latent domain using an encoder network.
3 . The method of claim 1 , wherein the diffusion model is configured to perform:
a reverse diffusion process that gradually subtracts noise values generated by a learned normal distribution from pixel values of the feature map; and a diffusion process that gradually adds noise values according to a fixed normal distribution to the pixel values of the feature map.
4 . The method of claim 3 , wherein the generating of the HR feature comprises restoring the certain frequency component corresponding to the feature map by repeatedly performing the reverse diffusion process of the diffusion model by a predetermined number of times.
5 . The method of claim 1 , wherein the generating of the HR feature comprises generating, using the diffusion model, the HR feature in which the certain frequency component corresponding to the feature map is restored, by concatenating or adding gradually changing input noise with or to the feature map.
6 . The method of claim 1 , wherein the generating of the HR feature comprises:
extracting a scaling factor and a bias factor from the feature map using a normalization technique; and generating the HR feature in which the certain frequency component corresponding to the feature map is restored, by applying the scaling factor and the bias factor to the feature map.
7 . The method of claim 6 , wherein the normalization technique comprises one of a group normalization technique, an adaptive group normalization (AdaGN) technique, and an instance normalization technique.
8 . The method of claim 1 , wherein the generating of the HR feature comprises generating the HR feature in which the certain frequency component corresponding to the feature map is restored, by applying the feature map to a cross-attention technique.
9 . The method of claim 1 , wherein the generating of the HR feature comprises:
receiving image quality-related keywords corresponding to the input image; extracting a text feature corresponding to the image quality-related keywords; and generating the HR feature in which the certain frequency component corresponding to the feature map is restored, by mapping the text feature to the feature map.
10 . The method of claim 9 , wherein the extracting of the text feature comprises extracting the text feature corresponding to the image quality-related keywords using a text encoder based on a pre-trained vision language model (VLM).
11 . The method of claim 1 , wherein
the mapping of the input image to the feature map of the latent domain further comprises estimating frequency information from the feature map, and the generating of the HR feature comprises generating the HR feature in which the certain frequency component corresponding to the feature map is restored, by applying the frequency information to the diffusion model.
12 . The method of claim 11 , wherein the estimating of the frequency information comprises estimating the frequency information from the feature map using either one or both of a fast Fourier transform (FFT) technique and a local texture estimation (LTE) technique for an implicit expression function.
13 . The method of claim 1 , wherein the restoring of the HR image comprises:
estimating pixel values of the HR image at the target magnification comprising the HR feature, based on the HR feature, and the coordinate information and the pixel information of the HR image at the target magnification; and restoring the HR image at the target magnification by the estimated pixel values.
14 . The method of claim 13 , wherein the estimating of the pixel values of the HR image at the target magnification comprises estimating the pixel values of the HR image at the target magnification comprising the HR feature, from the HR feature, and the coordinate information and the pixel information of the HR image at the target magnification, using a decoder network.
15 . The method of claim 14 , wherein the decoder network comprises a network based on multi-layer perceptron.
16 . A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1 .
17 . An apparatus with high-resolution (HR) image restoration, the apparatus comprising:
one or more processors configured to:
map, using an encoder network, an input image with low resolution (LR) to a feature map of a latent domain;
generate, using a diffusion model, an HR feature in which a certain frequency component corresponding to the feature map is restored; and
based on the HR feature and coordinate information and pixel information of an HR image at a target magnification, restore, using a decoder network, the HR image at the target magnification corresponding to the feature map.
18 . The apparatus of claim 17 , wherein, for the generating of the HR feature, the one or more processors are configured to generate, using the diffusion model, the HR feature in which the certain frequency component corresponding to the feature map is restored, by concatenating gradually changing input noise with the feature map.
19 . The apparatus of claim 17 , wherein
the one or more processors are configured to: extract, using a text encoder based on a pre-trained vision language model (VLM), extract a text feature corresponding to image quality-related keywords corresponding to the input image; and for the generating of the HR feature, generate, using the diffusion model, the HR feature in which the certain frequency component corresponding to the feature map is restored, by mapping the text feature to the feature map.
20 . The apparatus of claim 17 , wherein the apparatus is any one or any combination of any two or more of a smartphone, a camera, a closed-circuit television (CCTV), medical image equipment, semiconductor measurement equipment, an autonomous vehicle camera, a mixed reality (MR) device, and an augmented reality (AR) device.Join the waitlist — get patent alerts
Track US2025200701A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.