US2023245685A1PendingUtilityA1

Removing Visual Content Representing a Reflection of a Screen

Assignee: YAU SAMUEL CHI HONGPriority: Jun 16, 2020Filed: Mar 31, 2023Published: Aug 3, 2023
Est. expiryJun 16, 2040(~13.9 yrs left)· nominal 20-yr term from priority
G11B 27/02G06V 20/46H04N 5/2621
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A screen is manipulated to display content whose reflection is not captured by a sensor. In an embodiment, the screen inserts a black frame between screen content frames, displaying the black frame during a particular time. A sensor captures a video frame during the particular time. The video frame does not include a screen reflection and is considered a clean frame. In another embodiment, the screen displays screen content with a particular polarization during a particular time. A sensor captures a video frame with another polarization during the same time. The polarizations are selected such that the sensor is unable to capture screen reflections. The video frame is considered a clean frame. The clean frame is used to generate a masking frame, which is applied to target video frames to remove screen reflection. A modified target video, including the reflection-removed target video frame, and/or the clean frame itself, is generated.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 obtaining a target video captured by a first sensor, wherein the target video includes a target video frame, and the target video frame includes visual content representing a reflection of screen content displayed on a screen;   determining a first location of the visual content representing the reflection of the screen content, within the target video frame, with respect to a second location of visual content representing a reference object, within the target video frame;   obtaining a reference video captured by a second sensor, wherein the reference video includes a first reference video frame;   identifying a third location of visual content representing the reference object, within the first reference video frame;   transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a fourth location, within the first reference video frame, based on (a) the second location of the visual content representing the reference object within the target video frame and (b) the third location of the visual content representing the reference object within the first reference video frame;   identifying a first visual content at the fourth location within the first reference video frame;   generating replacement visual content based on the first visual content at the fourth location within the first reference video frame;   replacing the visual content representing the reflection of the screen content, within the target video frame, with the replacement visual content to generate a modified target video frame;   including the modified target video frame into a modified target video;   wherein the method is performed by at least one hardware device comprising a hardware processor.   
     
     
         2 . The method of  claim 1 , further comprising:
 determining whether the first visual content at the fourth location within the first reference video represents any portion of the reflection of the screen content;   wherein generating the replacement visual content based on the first visual content at the fourth location within the first reference video frame is responsive to determining that the first visual content at the fourth location within the first reference video does not represent any portion of the reflection of the screen content.   
     
     
         3 . The method of  claim 2 , further comprising:
 obtaining a second reference video frame;   identifying a fifth location of visual content representing the reference object, within the second reference video frame;   transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a sixth location, within the second reference video frame, based on (a) the second location of the visual content representing the reference object within the target video frame and (b) the fifth location of the visual content representing the reference object within the first reference video frame;   identifying a second visual content at the sixth location within the second reference video frame;   determining whether the second visual content at the sixth location within the second reference video frame represents any portion of the reflection of the screen content;   responsive to determining that the second visual content at the sixth location within the second reference video frame represents at least a portion of the reflection of the screen content: refraining from selecting the second visual content at the sixth location within the second reference video frame for use in generating the replacement visual content;   wherein generating the replacement visual content is not based on the second visual content at the sixth location within the second reference video frame.   
     
     
         4 . The method of  claim 1 , further comprising:
 obtaining a second reference video frame;   identifying a fifth location of visual content representing the reference object, within the second reference video frame;   transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a sixth location, within the second reference video frame, based on (a) the second location of the visual content representing the reference object within the target video frame and (b) the fifth location of the visual content representing the reference object within the first reference video frame;   identifying a second visual content at the sixth location within the second reference video frame;   selecting the second visual content at the sixth location within the second reference video frame for use in generating the replacement visual content;   wherein generating the replacement visual content is based on the first visual content at the fourth location within the first reference video frame and the second visual content at the sixth location within the second reference video frame.   
     
     
         5 . The method of  claim 4 , further comprising:
 aggregating the first visual content at the fourth location within the first reference video frame and the second visual content at the sixth location within the second reference video frame to generate the replacement visual content.   
     
     
         6 . The method of  claim 1 , wherein the first sensor and the second sensor are different sensors located at different physical locations. 
     
     
         7 . The method of  claim 1 , wherein the first sensor and the second sensor are same, and the target video and reference video are same, and the target video frame and the reference video frame are captured at different times. 
     
     
         8 . The method of  claim 1 , wherein the reference video frame is captured during a time period in which a black frame is displayed on the screen. 
     
     
         9 . The method of  claim 1 , further comprising:
 determining one or more of resizing, tilt, or positioning adjustment between the target video frame and the first reference video frame;   applying the one or more of resizing, tilt, or positioning adjustment in transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a fourth location, within the first reference video frame.   
     
     
         10 . The method of  claim 1 , further comprising:
 replacing another visual content representing reflection of the screen content, within a second target video frame within the target video, with the replacement visual content to generate a modified second target video frame;   including the modified second target video frame into the modified target video.   
     
     
         11 . The method  claim 1 , further comprising:
 determining that the target video frame includes the visual content representing the reflection of the screen content.   
     
     
         12 . The method of  claim 11 , wherein determining that the target video frame includes the visual content representing the reflection of the screen content comprises:
 performing a comparison between (a) at least a portion of visual content within the target video frame and (b) the screen content;   based on the comparison, determining a match between (a) at least a portion of visual content within the target video frame and (b) the screen content.   
     
     
         13 . The method of  claim 12 , wherein determining the match is based on anchors displayed in the screen content. 
     
     
         14 . The method of  claim 11 , wherein determining that the target video frame includes the visual content representing the reflection of the screen content comprises:
 determining that unexpected visual content appears within the target video frame.   
     
     
         15 . A system, comprising:
 at least one hardware device comprising a hardware processor;   the system being configured to perform:   obtaining a target video captured by a first sensor, wherein the target video includes a target video frame, and the target video frame includes visual content representing a reflection of screen content displayed on a screen;   determining a first location of the visual content representing the reflection of the screen content, within the target video frame, with respect to a second location of visual content representing a reference object, within the target video frame;   obtaining a reference video captured by a second sensor, wherein the reference video includes a first reference video frame;   identifying a third location of visual content representing the reference object, within the first reference video frame;   transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a fourth location, within the first reference video frame, based on (a) the second location of the visual content representing the reference object within the target video frame and (b) the third location of the visual content representing the reference object within the first reference video frame;   identifying a first visual content at the fourth location within the first reference video frame;   generating replacement visual content based on the first visual content at the fourth location within the first reference video frame;   replacing the visual content representing the reflection of the screen content, within the target video frame, with the replacement visual content to generate a modified target video frame;   including the modified target video frame into a modified target video.   
     
     
         16 . The system of  claim 15 , wherein the system is further configured to perform:
 obtaining a second reference video frame;   identifying a fifth location of visual content representing the reference object, within the second reference video frame;   transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a sixth location, within the second reference video frame, based on (a) the second location of the visual content representing the reference object within the target video frame and (b) the fifth location of the visual content representing the reference object within the first reference video frame;   identifying a second visual content at the sixth location within the second reference video frame;   selecting the second visual content at the sixth location within the second reference video frame for use in generating the replacement visual content;   wherein generating the replacement visual content is based on the first visual content at the fourth location within the first reference video frame and the second visual content at the sixth location within the second reference video frame.   
     
     
         17 . The system of  claim 15 , wherein the first sensor and the second sensor are different sensors located at different physical locations. 
     
     
         18 . The system of  claim 15 , wherein the first sensor and the second sensor are same, and the target video and reference video are same, and the target video frame and the reference video frame are captured at different times. 
     
     
         19 . The system of  claim 15 , wherein the system is further configured to perform: 
 determining one or more of resizing, tilt, or positioning adjustment between the target video frame and the first reference video frame;   applying the one or more of resizing, tilt, or positioning adjustment in transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a fourth location, within the first reference video frame.   
     
     
         20 . One or more non-transitory machine-readable media storing instructions which, when executed by one or more processors, cause:
 obtaining a target video captured by a first sensor, wherein the target video includes a target video frame, and the target video frame includes visual content representing a reflection of screen content displayed on a screen;   determining a first location of the visual content representing the reflection of the screen content, within the target video frame, with respect to a second location of visual content representing a reference object, within the target video frame;   obtaining a reference video captured by a second sensor, wherein the reference video includes a first reference video frame;   identifying a third location of visual content representing the reference object, within the first reference video frame;   transposing the first location of the visual content representing the reflection of the screen content, within the target video frame, to a fourth location, within the first reference video frame, based on (a) the second location of the visual content representing the reference object within the target video frame and (b) the third location of the visual content representing the reference object within the first reference video frame;   identifying a first visual content at the fourth location within the first reference video frame;   generating replacement visual content based on the first visual content at the fourth location within the first reference video frame;   replacing the visual content representing the reflection of the screen content, within the target video frame, with the replacement visual content to generate a modified target video frame;   including the modified target video frame into a modified target video.

Join the waitlist — get patent alerts

Track US2023245685A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.