US2023123005A1PendingUtilityA1

Real-time video dimensional transformations of video for presentation in mixed reality-based virtual spaces

Assignee: KICKBACK SPACE INCPriority: Feb 26, 2021Filed: Dec 21, 2022Published: Apr 20, 2023
Est. expiryFeb 26, 2041(~14.6 yrs left)· nominal 20-yr term from priority
Inventors:Rocco Haro
G06V 40/174G06V 40/23G06V 20/46G06T 19/006G06T 19/003H04N 7/157
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A non-immersive virtual reality (NIVR) method includes receiving sets of images of a first user and a second user, each image from the sets of images being an image of the associated user taken at a different angle from a set of angles. Video of the first user and the second user is received and processed. A first location and a first field of view are determined for a first virtual representation of the first user, and a second location and a second field of view are determined for a second virtual representation of the second user. Frames are generated for video planes of each of the first virtual representation of the first user and the second virtual representation of the second user based on the processed video, the sets of images, the first and second locations, and the first and second fields of view.

Claims

exact text as granted — not AI-modified
1 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:
 receive, from a first user compute device of a virtual reality system, the first user compute device associated with a first user, a first plurality of images of the first user, each image from the first plurality of images being an image of the first user taken at an associated angle from a plurality of different angles;   receive, from a second user compute device of the virtual reality system, the second user compute device associated with a second user, a second plurality of images of the second user, each image from the second plurality of images being an image of the second user taken at an associated angle from the plurality of different angles;   process a first video of the first user to generate a first processed video;   process a second video of the second user to generate a second processed video;   determine, for a first virtual representation of the first user, (1) a first location of the first virtual representation in a virtual environment, and (2) a first field of view of the first virtual representation in the virtual environment;   determine, for a second virtual representation of the second user, (1) a second location of the second virtual representation in the virtual environment, and (2) a second field of view of the second virtual representation in the virtual environment;   generate at least one first frame for a first video plane of the first virtual representation based on the first processed video, at least one image from the first plurality of images, the first location, the second location, the first field of view, and the second field of view, the at least one first frame including at least one perspective view of the first virtual representation of the first user;   generate at least one second frame for a second video plane of the second virtual representation based on the second processed video, at least one image from the second plurality of images, the first location, the second location, the first field of view, and the second field of view;   cause transmission of at least one first signal representing the at least one first frame for the first video plane to at least one engine, to cause display, at the second user compute device, of the at least one first frame for the first video plane in the virtual environment to the second user;   cause transmission of at least one second signal representing the at least one second frame for the second video plane to the at least one engine, to cause display, at the first user compute device, of the at least one second frame for the second video plane in the virtual environment to the first user; and   cause transmission of at least one third signal to the first user compute device requesting the first plurality of images in response to determining that the first plurality of images has not been received.   
     
     
         2 . The non-transitory, processor-readable medium of  claim 1 , further storing instructions to cause the processor to:
 dynamically update the first video plane with at least one third frame, substantially in real-time, to include a representation of at least one of a facial expression of the first user, a voice of the first user, or a torso movement of the first user, and   dynamically update the second video plane with at least one fourth frame, substantially in real-time, to include a representation of at least one of a facial expression of the second user, a voice of the second user, or a torso movement of the second user.   
     
     
         3 . The non-transitory, processor-readable medium of  claim 1 , wherein:
 the instructions include instructions to:   generate the at least one first frame for the first video plane and the at least one second frame for the second video plane substantially in parallel, and   cause transmission of the first signal and the second signal substantially in parallel.   
     
     
         4 . The non-transitory, processor-readable medium of  claim 1 , wherein the virtual environment is an emulation of a virtual three-dimensional space. 
     
     
         5 . The non-transitory, processor-readable medium of  claim 1 , wherein each frame from the first video has a first common background, each frame from the second video has a second common background different than the first common background, each frame from the first processed video has a third common background, and each frame from the second processed video has the third common background. 
     
     
         6 . The non-transitory, processor-readable medium of  claim 1 , further storing instructions to cause the processor to:
 receive, from the first user compute device, a first request to join the virtual environment;   receive, from the second user compute device, a second request to join the virtual environment; and   cause transmission of at least one fourth signal to the second user compute device requesting the second plurality of images in response to determining that the second plurality of images has not been received.   
     
     
         7 . The non-transitory, processor-readable medium of  claim 1 , wherein:
 the instructions to process the first video include instructions to: 
 decode each frame of the first video to generate a first plurality of decoded frames, and 
 edit, for each background portion of a frame from the first plurality of decoded frames, that background portion to a standard format; and 
   the processing of the second video includes:
 decode each frame of the second video to generate a second plurality of decoded frames, and 
 edit, for each background portion of a frame from the second plurality of decoded frames, that background portion to the standard format. 
   
     
     
         8 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:
 receive first state information indicating (1) a first location of a first virtual representation of a first user in a virtual environment, (2) a second location of a second virtual representation of a second user in the virtual environment, (3) a first field of view of the first virtual representation of the first user in the virtual environment, and (4) a second field of view of the second virtual representation of the second user in the virtual environment;   receive, from a first user compute device associated with the first user, a plurality of images of the first user, each image from the plurality of images being an image of the first user taken at an associated angle from a plurality of different angles;   generate a first set of frames for a video plane of the first virtual representation based on a first set of video frames of the first user, at least one image from the plurality of images, the first location, the second location, the first field of view, and the second field of view;   cause transmission of a first signal representing the first set of frames to at least one engine to cause a second user compute device associated with the second user to display the first set of frames in the virtual environment to the second user;   cause transmission of second state information indicating (1) a third location of the first virtual representation in the virtual environment different than the first location, (2) the second location of the second virtual representation in the virtual environment, (3) a third field of view of the first virtual representation in the virtual environment different than the first field of view, and (4) the second field of view of the second virtual representation in the virtual environment;   receive, from the first user compute device, a second set of video frames of the first user;   generate a second set of frames for the video plane of the first virtual representation (1) different than the first set of frames and (2) based on the second set of video frames, at least one image from the plurality of images, the third location, the second location, the third field of view, and the second field of view;   cause transmission of a second signal representing the second set of frames to the at least one engine;   receive third state information indicating (1) a fourth location of the first virtual representation in the virtual environment different than the first location and the third location, (2) the second location of the second virtual representation in the virtual environment, (3) a fourth field of view of the first virtual representation in the virtual environment different than the first field of view and the third field of view, and (4) the second field of view of the second virtual representation in the virtual environment;   generate a third set of frames for the video plane of the first virtual representation based on a third set of video frames of the first user, at least one image from the plurality of images, the fourth location, the second location, the fourth field of view, and the second field of view; and   cause transmission of a third signal representing the third set of frames to the at least one engine.   
     
     
         9 . The non-transitory, processor-readable medium of  claim 8 , wherein the first set of frames shows at least one first perspective view of the first virtual representation of the first user, and the second set of frames shows at least one second perspective view of the first virtual representation of the first user different than the at least one first perspective view. 
     
     
         10 . The non-transitory, processor-readable medium of  claim 8 , further storing instructions to cause the processor to:
 receive, from the first user compute device, the third set of video frames of the first user.   
     
     
         11 . The non-transitory, processor-readable medium of  claim 8 , further storing instructions to cause the processor to:
 receive fourth state information indicating (1) a fifth location of the first virtual representation in the virtual environment different than the first location and the third location, (2) a sixth location of the second virtual representation in the virtual environment different than the second location, (3) a fifth field of view of the first virtual representation in the virtual environment different than the first field of view and the third field of view, and (4) a sixth field of view of the second virtual representation in the virtual environment different than the second field of view;   receive, from the first user compute device, a fourth set of video frames of the first user;   generate a fourth set of frames for the video plane of the first virtual representation based on the fourth set of video frames, at least one image from the plurality of images, the fifth location, the sixth location, the fifth field of view, and the sixth field of view; and   cause transmission of a fourth signal representing the fourth set of frames to the at least one engine.   
     
     
         12 . The non-transitory, processor-readable medium of  claim 8 , further storing instructions to cause the processor to:
 receive fourth state information indicating (1) a fifth location of the first virtual representation in the virtual environment different than the first location and the third location, (2) the second location of the second virtual representation in the virtual environment, (3) a fifth field of view of the first virtual representation in the virtual environment different than the first field of view and the third field of view, and (4) the second field of view of the second virtual representation in the virtual environment;   determine that the first virtual representation is not in the second field of view of the second virtual representation based on the fifth location, the second location, and the fifth field of view; and   refrain from generating a fourth set of frames of the first virtual representation.   
     
     
         13 . The non-transitory, processor-readable medium of  claim 8 , further storing instructions to cause the processor to:
 dynamically update the video plane, in real-time, to include a representation of at least one of a facial expression of the first user, a voice of the first user, or a torso movement of the first user.   
     
     
         14 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:
 receive first state information indicating (1) a first location of a first virtual representation of a first user in a virtual environment, (2) a second location of a second virtual representation of a second user in the virtual environment, (3) a first field of view of the first virtual representation of the first user in the virtual environment, and (4) a second field of view of the second virtual representation of the second user in the virtual environment;   receive, from a first user compute device associated with the first user, a plurality of images of the first user, each image from the plurality of images being an image of the first user taken at an associated angle from a plurality of different angles;   generate a first set of frames for a video plane of the first virtual representation based on a first set of video frames of the first user, at least one image from the plurality of images, the first location, the second location, the first field of view, and the second field of view;   cause transmission of a first signal representing the first set of frames to at least one engine to cause a second user compute device associated with the second user to display the first set of frames in the virtual environment to the second user;   cause transmission of second state information indicating (1) a third location of the first virtual representation in the virtual environment different than the first location, (2) the second location of the second virtual representation in the virtual environment, (3) a third field of view of the first virtual representation in the virtual environment different than the first field of view, and (4) the second field of view of the second virtual representation in the virtual environment;   receive, from the first user compute device, a second set of video frames of the first user;   generate a second set of frames for the video plane of the first virtual representation (1) different than the first set of frames and (2) based on the second set of video frames, at least one image from the plurality of images, the third location, the second location, the third field of view, and the second field of view;   cause transmission of a second signal representing the second set of frames to the at least one engine;   receive third state information indicating (1) a fourth location of the first virtual representation in the virtual environment different than the first location and the third location, (2) a fifth location of the second virtual representation in the virtual environment different than the second location, (3) a fourth field of view of the first virtual representation in the virtual environment different than the first field of view and the third field of view, and (4) a fifth field of view of the second virtual representation in the virtual environment different than the second field of view;   generate a third set of frames for the video plane of the first virtual representation based on a third set of video frames of the first user, at least one image from the plurality of images, the fourth location, the fifth location, the fourth field of view, and the fifth field of view; and   cause transmission of a third signal representing the third set of frames to the at least one engine.   
     
     
         15 . The non-transitory, processor-readable medium of  claim 14 , further storing instructions to cause the processor to:
 receive, from the first user compute device, the third set of video frames of the first user before generating the third set of frames for the video plane of the first virtual representation.   
     
     
         16 . The non-transitory, processor-readable medium of  claim 14 , further storing instructions to cause the processor to:
 dynamically update the video plane, in real-time, to include a representation of at least one of a facial expression of the first user, a voice of the first user, or a torso movement of the first user.   
     
     
         17 . The non-transitory, processor-readable medium of  claim 14 , wherein the virtual environment is a non-immersive virtual environment. 
     
     
         18 . The non-transitory, processor-readable medium of  claim 14 , wherein the first set of frames shows at least one first perspective view of the first virtual representation of the first user, and the second set of frames shows at least one second perspective view of the first virtual representation of the first user different than the at least one first perspective view. 
     
     
         19 . The non-transitory, processor-readable medium of  claim 14 , further storing instructions to cause the processor to:
 receive fourth state information indicating (1) a sixth location of the first virtual representation in the virtual environment different than the first location and the third location, (2) the second location of the second virtual representation in the virtual environment, (3) a sixth field of view of the first virtual representation in the virtual environment different than the first field of view and the third field of view, and (4) the second field of view of the second virtual representation in the virtual environment;   determine that the first virtual representation is not in the second field of view of the second virtual representation based on the sixth location, the second location, and the sixth field of view; and   refrain from generating a fourth set of frames of the first virtual representation.   
     
     
         20 . The non-transitory, processor-readable medium of  claim 14 , wherein the first set of frames is generated using a generative adversarial network (GAN) and the second set of frames is generated using the GAN.

Join the waitlist — get patent alerts

Track US2023123005A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.