US2024135515A1PendingUtilityA1
Method of and apparatus for processing digital image data
Est. expiryOct 17, 2042(~16.2 yrs left)· nominal 20-yr term from priority
G06T 5/60G06T 2207/20084G06T 2207/20081G06T 2207/10132G06T 2207/10024G06T 2207/10016G06N 3/094G06N 3/0475G06N 3/0455G01S 17/89G06T 7/00G06T 5/005G06T 7/0002G06V 10/7715G06V 10/774G06V 10/82G06T 2207/20021G06T 2207/30168G06T 11/00G06T 5/77G06N 3/045G06N 3/0464G06N 3/088G06N 3/096G06T 5/70
57
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A computer-implemented method of processing digital image data. The method includes: determining, by an encoder configured to map a first digital image to an extended latent space associated with a generator of a generative adversarial network, GAN, system, a noise prediction associated with the first digital image, determining, by the generator of the GAN system, at least one further digital image based on the noise prediction associated with the first digital image and a plurality of latent variables associated with the extended latent space.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method of processing digital image data, comprising the following steps:
determining, by an encoder configured to map a first digital image to an extended latent space associated with a generator of a generative adversarial network (GAN) system, a noise prediction associated with the first digital image; and determining, by the generator of the GAN system, at least one further digital image based on the noise prediction associated with the first digital image and a plurality of latent variables associated with the extended latent space.
2 . The method according to claim 1 , further comprising:
determining the plurality of latent variables based on at least one of:
a) a second digital image, which is different from the first digital image using the encoder,
b) a plurality of probability distributions.
3 . The method according to claim 1 , wherein at least some of the plurality of latent variables associated with the extended latent space characterize at least one of the following aspects of the first digital image:
a) a style including a non-semantic appearance, b) a texture, c) a color.
4 . The method according to claim 1 , further comprising at least one of:
a) determining a plurality of hierarchical feature maps based on the first digital image, b) determining a plurality of latent variables associated with the extended latent space for the first digital image based on the plurality of hierarchical feature maps, c) determining an additive noise map based on at least one of the plurality of hierarchical feature maps.
5 . The method according to claim 1 , further comprising:
randomly and/or pseudo-randomly masking at least a portion of the noise prediction associated with the first digital image.
6 . The method according to claim 4 , further comprising: masking of the noise map in a random and/or pseudo-random fashion.
7 . The method according to claim 6 , further comprising:
spatially dividing the noise map into a plurality of, non-overlapping patches; selecting, in a random and/or pseudo-random fashion, a subset of the patches; replacing the subset of the patches by patches of unit Gaussian random variables of the same size.
8 . The method according to claim 1 , further comprising:
combining the noise prediction associated with the first digital image with a style prediction of a second digital image; and generating a further digital image using the generator based on the combined noise prediction associated with the first digital image and the style prediction of the second digital image.
9 . The method according to claim 1 , further comprising:
providing the noise prediction associated with the first digital image; providing different sets of latent variables characterizing different styles to be applied to a semantic content of the first digital image; and generating a plurality of digital images with the different styles using the generator based on the noise prediction associated with the first digital image and the different sets of latent variables characterizing the different styles.
10 . The method according to claim 1 , further comprising:
providing image data, including one or more digital images, associated with a first domain; providing image data, including one or more digital images, associated with a second domain; and applying a style of the second domain to the image data associated with the first domain.
11 . The method according to claim 10 , wherein the image data associated with the first domain include labels, wherein the applying of the style of the second domain to the image data associated with the first domain includes preserving the labels.
12 . The method according to claim 1 , further comprising:
providing first image data having first content information; providing second image data, wherein the second image data includes second content information different from the first content information; extracting style information of the second image data; and applying at least a part of the style information of the second image data to the first image data.
13 . The method according to claim 1 , further comprising:
generating training data for training at least one neural network system, wherein the generating is based on image data of a source domain and based on modified image data of the source domain, wherein the modified image data is and/or has been modified with respect to an image style, based on a style of further image data; and training the at least one neural network system based on the training data.
14 . An apparatus configured to process digital image data, the apparatus configured to:
determine, by an encoder configured to map a first digital image to an extended latent space associated with a generator of a generative adversarial network (GAN) system, a noise prediction associated with the first digital image; and determine, by the generator of the GAN system, at least one further digital image based on the noise prediction associated with the first digital image and a plurality of latent variables associated with the extended latent space.
15 . A non-transitory computer-readable storage medium on which is stored a computer program including instructions processing digital image data, the instructions, when executed by a computer, causing the computer to perform the following steps:
determining, by an encoder configured to map a first digital image to an extended latent space associated with a generator of a generative adversarial network (GAN) system, a noise prediction associated with the first digital image; and determining, by the generator of the GAN system, at least one further digital image based on the noise prediction associated with the first digital image and a plurality of latent variables associated with the extended latent space.
16 . The method according to claim 1 , further comprising using the method for at least one of the following:
a) determining at least one further digital image based on the noise prediction associated with the first digital image and the plurality of latent variables associated with the extended latent space, at least some of the plurality of latent variables being associated with another image and/or other data than the first digital image, b) transferring a style from a second digital image to the first digital image, while preserving a content of the first digital image, c) disentangling style and content of at least one digital image. d) creating different stylized digital images with unchanged content, based on the first digital image and a style of at least one further second digital image, e) using or re-using labelled annotations for stylized images, f) avoiding annotation work when changing a style of at least one digital image, g) generating perceptually realistic, digital images with different styles, h) providing proxy validation sets for testing out-of-distribution generalization of a neural network system, i) training a machine learning system, j) testing a machine learning system, k) verifying a machine learning system, l) validating a machine learning system, m) generating training data for a machine learning system, n) data augmentation of existing image data, o) improving a generalization performance of a machine learning system, p) flexibly manipulating image styles without a training associated with multiple data sets, q) utilizing an encoder GAN pipeline to manipulate image styles, r) embedding, by the encoder, information associated with an image style into intermediate latent variables, s) mixing styles of digital images for generating at least one further digital image including a style based on the mixing.Join the waitlist — get patent alerts
Track US2024135515A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.