Image processing system, image processing method, and program
Abstract
Techniques include acquiring 1st to Nth input frames (N is a natural number equal to or greater than 2) having a prescribed input pixel number. The techniques further include acquiring, based on each of the input frames, 1st to Nth intermediate frames by generating an intermediate frame for each input frame which corresponds to the input frame and which includes an intermediate pixel number equal to or greater than the input pixel number. The techniques further include inputting each of the intermediate frames to a machine learning model. The techniques further include acquiring 1st to Nth estimation frames including an estimated pixel number equal to or greater than the intermediate pixel number which is greater than the input pixel number.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An image processing system comprising:
one or more storage media storing instructions; and one or more processors configured to execute the instructions to cause the image processing system to: acquire 1st to Nth input frames (N is a natural number equal to or greater than 2) having a prescribed input pixel number; acquire, based on each of the input frames, 1st to Nth intermediate frames by generating an intermediate frame for each input frame which corresponds to the input frame and which includes an intermediate pixel number equal to or greater than the input pixel number; and input each of the intermediate frames to a machine learning model; and acquire 1st to Nth estimation frames including an estimated pixel number equal to or greater than the intermediate pixel number which is greater than the input pixel number, wherein the machine learning model includes: a cumulative feature information output layer to which the nth intermediate frame (n=2, 3, . . . , N) and (n−1)th auxiliary information based on (n−1)th cumulative feature information indicating the features of the 1st to (n−1)th intermediate frames are inputted, wherein the cumulative feature information output layer outputs nth cumulative feature information indicating the features of the 1st to nth intermediate frames; and an estimation frame output layer to which the nth cumulative feature information is inputted, wherein the estimation frame output layer outputs the nth estimation frame, wherein the machine learning model was trained using a plurality of training data including: a learning intermediate frame including the intermediate pixel number generated based on a learning input frame having the input pixel number; and a learning estimation frame including the estimated pixel number.
2 . The image processing system of claim 1 , wherein each of the input frames includes an image obtained by executing rendering of three-dimensional data indicating one or more objects as seen from a prescribed viewpoint.
3 . The image processing system of claim 2 , wherein each of the input frames includes an image obtained by executing the rendering so that the viewpoint changes for each of the input frames, and wherein the instructions further cause the image processing system to:
acquire change information including information relating to a change of the viewpoint for each of the input frames in the rendering; and obtain a pixel value of a position corresponding to each pixel before the change by interpolation in the input frame, based on the change information and each pixel of each of the input frames; and generate each of the intermediate frames.
4 . The image processing system of claim 2 , wherein the instructions further cause the image processing system to:
acquire the (n−1)th motion information including information indicating an amount and a direction of motion from the (n−1)th input frame towards the nth input frame; and acquire the (n−1)th auxiliary information by applying motion compensation on the (n−1)th cumulative feature information, based on the (n−1)th motion information.
5 . The image processing system of claim 4 , wherein the instructions further cause the image processing system to:
acquire (n−1)th depth information indicating each pixel depth of the (n−1)th input frame, and nth depth information indicating each pixel depth of the nth input frame; specify, amongst the nth intermediate frame pixels, an nth appearance pixel as a fully or partially displayed pixel of the object which is not displayed in the (n−1)th intermediate frame, based on the (n−1)th depth information and the nth depth information; and acquire the (n−1)th auxiliary information by converting a pixel value of the nth appearance pixel in the (n−1)th cumulative feature information to a prescribed value.
6 . The image processing system of claim 1 , wherein the 1st intermediate frame and a given auxiliary information are input to the cumulative feature information output layer which outputs the 1st cumulative feature information.
7 . The image processing system of claim 1 , wherein the cumulative feature information includes image information having the same pixel number as the intermediate pixel number.
8 . A method comprising:
acquiring 1st to Nth input frames (N is a natural number equal to or greater than 2) having a prescribed input pixel number; acquiring, based on each of the input frames, 1st to Nth intermediate frames by generating an intermediate frame for each input frame which corresponds to the input frame and which includes an intermediate pixel number equal to or greater than the input pixel number; and inputting each of the intermediate frames to a machine learning model; and acquiring 1st to Nth estimation frames including an estimated pixel number equal to or greater than the intermediate pixel number which is greater than the input pixel number, wherein the machine learning model includes: a cumulative feature information output layer to which the nth intermediate frame (n=2, 3, . . . , N) and (n−1)th auxiliary information based on (n−1)th cumulative feature information indicating the features of the 1st to (n−1)th intermediate frames are inputted, wherein the cumulative feature information output layer outputs nth cumulative feature information indicating the features of the 1st to nth intermediate frames; and an estimation frame output layer to which the nth cumulative feature information is inputted, wherein the estimation frame output layer outputs the nth estimation frame, wherein the machine learning model was trained using a plurality of training data including: a learning intermediate frame including the intermediate pixel number generated based on a learning input frame having the input pixel number; and a learning estimation frame including the estimated pixel number.
9 . The method of claim 8 , wherein each of the input frames includes an image obtained by executing rendering of three-dimensional data indicating one or more objects as seen from a prescribed viewpoint.
10 . The method of claim 9 , wherein each of the input frames includes an image obtained by executing the rendering so that the viewpoint changes for each of the input frames, and wherein the method further comprises:
acquiring change information including information relating to a change of the viewpoint for each of the input frames in the rendering; and obtaining a pixel value of a position corresponding to each pixel before the change by interpolation in the input frame, based on the change information and each pixel of each of the input frames; and generating each of the intermediate frames.
11 . The method of claim 9 , further comprising:
acquiring the (n−1)th motion information including information indicating an amount and a direction of motion from the (n−1)th input frame towards the nth input frame; and acquiring the (n−1)th auxiliary information by applying motion compensation on the (n−1)th cumulative feature information, based on the (n−1)th motion information.
12 . The method of claim 11 , further comprising:
acquiring (n−1)th depth information indicating each pixel depth of the (n−1)th input frame, and nth depth information indicating each pixel depth of the nth input frame; specifying, amongst the nth intermediate frame pixels, an nth appearance pixel as a fully or partially displayed pixel of the object which is not displayed in the (n−1)th intermediate frame, based on the (n−1)th depth information and the nth depth information; and acquiring the (n−1)th auxiliary information by converting a pixel value of the nth appearance pixel in the (n−1)th cumulative feature information to a prescribed value.
13 . The method of claim 8 , wherein the 1st intermediate frame and a given auxiliary information are input to the cumulative feature information output layer which outputs the 1st cumulative feature information.
14 . The method of claim 8 , wherein the cumulative feature information includes image information having the same pixel number as the intermediate pixel number.
15 . One or more non-transitory computer-readable storage media storing instructions that, upon execution by one or more processors of a system, cause the system to perform operations comprising:
acquiring 1st to Nth input frames (N is a natural number equal to or greater than 2) having a prescribed input pixel number; acquiring, based on each of the input frames, 1st to Nth intermediate frames by generating an intermediate frame for each input frame which corresponds to the input frame and which includes an intermediate pixel number equal to or greater than the input pixel number; and inputting each of the intermediate frames to a machine learning model; and acquiring 1st to Nth estimation frames including an estimated pixel number equal to or greater than the intermediate pixel number which is greater than the input pixel number, wherein the machine learning model includes: a cumulative feature information output layer to which the nth intermediate frame (n=2, 3, . . . , N) and (n−1)th auxiliary information based on (n−1)th cumulative feature information indicating the features of the 1st to (n−1)th intermediate frames are inputted, wherein the cumulative feature information output layer outputs nth cumulative feature information indicating the features of the 1st to nth intermediate frames; and an estimation frame output layer to which the nth cumulative feature information is inputted, wherein the estimation frame output layer outputs the nth estimation frame, wherein the machine learning model was trained using a plurality of training data including: a learning intermediate frame including the intermediate pixel number generated based on a learning input frame having the input pixel number; and a learning estimation frame including the estimated pixel number.
16 . The computer-readable storage media of claim 15 , wherein each of the input frames includes an image obtained by executing rendering of three-dimensional data indicating one or more objects as seen from a prescribed viewpoint.
17 . The computer-readable storage media of claim 16 , wherein each of the input frames includes an image obtained by executing the rendering so that the viewpoint changes for each of the input frames, and wherein the operations further comprise:
acquiring change information including information relating to a change of the viewpoint for each of the input frames in the rendering; and obtaining a pixel value of a position corresponding to each pixel before the change by interpolation in the input frame, based on the change information and each pixel of each of the input frames; and generating each of the intermediate frames.
18 . The computer-readable storage media of claim 16 , wherein the operations further comprise:
acquiring the (n−1)th motion information including information indicating an amount and a direction of motion from the (n−1)th input frame towards the nth input frame; and acquiring the (n−1)th auxiliary information by applying motion compensation on the (n−1)th cumulative feature information, based on the (n−1)th motion information.
19 . The computer-readable storage media of claim 15 , wherein the 1st intermediate frame and a given auxiliary information are input to the cumulative feature information output layer which outputs the 1st cumulative feature information.
20 . The computer-readable storage media of claim 15 , wherein the cumulative feature information includes image information having the same pixel number as the intermediate pixel number.Join the waitlist — get patent alerts
Track US2026094433A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.