Interactive selfie panorama capture and multi-perspective undistorted selfie generation
Abstract
An electronic device includes at least one imaging sensor configured to obtain an input set of images of a scene. The electronic device also includes at least one processing device configured to generate a differentiable 3D model of the scene based on an iterative process using the input set of images and project the differentiable 3D model into an image space to generate an estimated burst of selfie images. To obtain the input set of images, the at least one processing device may be configured to obtain an initial burst of images, generate a map indicating subjects in the scene based on the initial burst of images, provide a prompt for a user to move the electronic device to capture at least one additional burst of images, and obtain the input set of images based on the initial burst of images and the at least one additional burst of images.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
entering, using at least one processing device of an electronic device, a selfie-capture mode; obtaining, using at least one imaging sensor of the electronic device, an input set of images of a scene; generating, using the at least one processing device, a differentiable three-dimensional (3D) model of the scene based on an iterative process using the input set of images; and projecting, using the at least one processing device, the differentiable 3D model into an image space to generate an estimated burst of selfie images.
2 . The method of claim 1 , wherein obtaining the input set of images comprises:
obtaining an initial burst of images; generating a map of the scene based on the initial burst of images, wherein the map indicates subjects in the scene; providing a prompt for a user to move the electronic device to capture at least one additional burst of images; and obtaining the input set of images based on the initial burst of images and the at least one additional burst of images.
3 . The method of claim 1 , wherein generating the differentiable 3D model comprises:
determining a differentiable loss between measured and estimated burst images; and updating parameters of the differentiable 3D model to reduce the differentiable loss.
4 . The method of claim 3 , wherein generating the differentiable 3D model further comprises:
generating a depth map based on an estimated depth for each image of the estimated burst of selfie images; and initializing the differentiable 3D model with the depth map from each image of the estimated burst of selfie images.
5 . The method of claim 3 , further comprising:
classifying pixels for each image of the estimated burst of selfie images into semantic classes; and updating the parameters of the differentiable 3D model to reduce the differentiable loss by increasing weighting factors of some semantic classes compared to other semantic classes for each image of the estimated burst of selfie images.
6 . The method of claim 1 , further comprising:
obtaining perspective information for each image of the estimated burst of selfie images; and performing a rendering for each image of the estimated burst of selfie images based on the perspective information.
7 . The method of claim 6 , further comprising:
training, using the at least one processing device, a machine learning model to predict the rendering for each image based on the perspective information prior to performing the rendering.
8 . The method of claim 7 , further comprising:
generating a metric for each image of the estimated burst of selfie images; and obtaining a final image based on a comparison of the metrics.
9 . The method of claim 7 , further comprising:
providing a prompt for a user to select a desired final image from the rendering for each image based on the perspective information.
10 . The method of claim 6 , further comprising:
generating a metric for each image of the estimated burst of selfie images; and obtaining a final image based on a comparison of the metrics.
11 . The method of claim 6 , further comprising:
training, using the at least one processing device, a machine learning model to predict a differential 3D model of the scene that is optimized further at inference time to produce renderings at different perspectives.
12 . The method of claim 11 , further comprising:
generating a metric for each image of the estimated burst of selfie images; and obtaining a final image based on a comparison of the metrics.
13 . The method of claim 11 , further comprising:
providing a prompt for a user to select a desired final image from the rendering for each image based on the perspective information.
14 . An electronic device comprising:
at least one imaging sensor configured to obtain an input set of images of a scene; and at least one processing device configured to:
generate a differentiable three-dimensional (3D) model of the scene based on an iterative process using the input set of images; and
project the differentiable 3D model into an image space to generate an estimated burst of selfie images.
15 . The electronic device of claim 14 , wherein, to obtain the input set of images, the at least one processing device is configured to:
obtain an initial burst of images; generate a map of the scene based on the initial burst of images, wherein the map indicates subjects in the scene; provide a prompt for a user to move the electronic device to capture at least one additional burst of images; and obtain the input set of images based on the initial burst of images and the at least one additional burst of images.
16 . The electronic device of claim 14 , wherein, to generate the differentiable 3D model, the at least one processing device is configured to:
determine a differentiable loss between measured and estimated burst images; and update parameters of the differentiable 3D model to reduce the differentiable loss.
17 . The electronic device of claim 16 , wherein, to generate the differentiable 3D model, the at least one processing device is further configured to:
generate a depth map based on an estimated depth for each image of the estimated burst of selfie images; and initialize the differentiable 3D model with the depth map from each image of the estimated burst of selfie images.
18 . The electronic device of claim 16 , wherein the at least one processing device is further configured to:
classify pixels for each image of the estimated burst of selfie images into semantic classes; and update the parameters of the differentiable 3D model to reduce the differentiable loss by increasing weighting factors of some semantic classes compared to other semantic classes for each image of the estimated burst of selfie images.
19 . The electronic device of claim 14 , wherein the at least one processing device is further configured to:
obtain perspective information for each image of the estimated burst of selfie images; and perform a rendering for each image of the estimated burst of selfie images based on the perspective information.
20 . The electronic device of claim 19 , wherein the at least one processing device is further configured to train a machine learning model to predict the rendering for each image based on the perspective information prior to performing the rendering.Join the waitlist — get patent alerts
Track US2026038195A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.