US2022223153A1PendingUtilityA1

Voice controlled camera with ai scene detection for precise focusing

Assignee: INTEL CORPPriority: Jun 28, 2019Filed: Mar 28, 2022Published: Jul 14, 2022
Est. expiryJun 28, 2039(~12.9 yrs left)· nominal 20-yr term from priority
Inventors:Daniel Pohl
H04N 23/64H04N 23/62H04N 23/67G06N 3/045H04N 23/63G06N 3/044G06N 3/0464G06N 3/0442G06F 40/56G06N 3/086G06F 3/167G06N 3/08G10L 15/22H04N 5/23222
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An apparatus, method and computer readable medium for a voice-controlled camera with artificial intelligence (AI) for precise focusing. The method includes receiving, by the camera, natural language instructions from a user for focusing the camera to achieve a desired photograph. The natural language instructions are processed using natural language processing techniques to enable the camera to understand the instructions. A preview image of a user desired scene is captured by the camera. Artificial Intelligence (AI) is applied to the preview image to obtain context and to detect objects within the preview image. A depth map of the preview image is generated to obtain distances from the detected objects in the preview image to the camera. It is determined whether the detected objects in the image match the natural language instructions from the user.

Claims

exact text as granted — not AI-modified
1 - 25 . (canceled) 
     
     
         26 . An apparatus comprising:
 a voice-controlled camera, the voice-controlled camera to continuously listen for voice commands from a user, the voice commands comprising instructions for focusing the voice-controlled camera on one or more objects to achieve a desired image, the voice-controlled camera including Artificial Intelligence (AI) scene detection based on natural language processing to convert the instructions into keywords that allow the voice-controlled camera to adjust camera settings to perform precise focusing of the desired image.   
     
     
         27 . The apparatus of  claim 26 , wherein the voice-controlled camera is triggered using a wake word, wherein the instructions immediately follow the wake word. 
     
     
         28 . The apparatus of  claim 26 , wherein the voice-controlled camera, after receiving the instructions, to simultaneously capture a preview image and to perform the natural language processing on the instructions using deep learning neural networks based on dense vector representations, the deep learning neural networks including one or more of a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), and a Recursive Neural Network. 
     
     
         29 . The apparatus of  claim 28 , wherein the voice-controlled camera to apply Artificial Intelligence (AI) on the preview image to perform object detection within the preview image and to provide context as to what is in the preview image. 
     
     
         30 . The apparatus of  claim 29 , wherein to perform the object detection and to provide the context uses one or more of Semantic Segmentation in real-time using Fully Convolutional Networks (FCN), R-CNN (Regional-based Convolutional Neural Networks (R-CNN), Fast R-CNN, Faster R-CNN, YOLO (You Only Look Once), and Mask R-CNN Using TensorRT. 
     
     
         31 . The apparatus of  claim 29 , wherein after the one or more objects have been identified in the preview image, distances from each object to the location of the voice-controlled camera are determined. 
     
     
         32 . The apparatus of  claim 31 , wherein the distances of the one or more objects to the voice-controlled camera are determined using a depth map, wherein the depth map is obtained using monocular SLAM (Simultaneous Localization and Mapping). 
     
     
         33 . The apparatus of  claim 31 , wherein the distances of the one or more objects to the voice-controlled camera are determined using depth sensors. 
     
     
         34 . The apparatus of  claim 31 , wherein when the one or more objects identified in the instructions are found in the preview image, a focus point and the camera settings are determined, the voice-controlled camera to adjust the focus point and camera settings to achieve the desired image of the user, and the desired image is taken. 
     
     
         35 . The apparatus of  claim 34 , wherein the focus point and the camera settings are determined by calculating optical formulas for cameras based on one or more detected objects to be photographed, positions of the one or more detected objects in the preview image, and estimated depth or distance of the one or more detected objects to the voice-controlled camera. 
     
     
         36 . The apparatus of  claim 34 , wherein the camera focus point and the camera settings are determined through experimentation by selecting a camera parameter and viewing an image of that selection using depth of field preview, wherein when the image is not good, continuously changing the camera parameter and viewing the image until the image is the desired image. 
     
     
         37 . The apparatus of  claim 34 , wherein the desired image is taken automatically by the voice-controlled camera. 
     
     
         38 . The apparatus of  claim 34 , wherein the user is prompted by the voice-controlled camera to take the desired image. 
     
     
         39 . A method comprising:
 receiving, through a microphone of a voice-controlled camera, voice commands from a user, the voice commands comprising instructions to focus the voice-controlled camera on one or more objects to achieve a desired image of the user; and   converting, by the voice-controlled camera using Artificial Intelligence (AI) scene detection based on natural language processing (NLP), the instructions into keywords that allow the voice-controlled camera to recognize the instructions and to adjust the camera settings to perform the precise focusing of the desired image.   
     
     
         40 . The method of  claim 39 , wherein after receiving the voice commands, the method further comprises:
 simultaneously capturing a preview image and performing the NLP on the instructions using deep learning neural networks based on dense vector representations by the voice-controlled camera, the deep learning neural networks including one or more of a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), and a Recursive Neural Network;   applying, by the voice-controlled camera, the AI scene detection on the preview image to perform object detection within the preview image and to provide context as to what is in the preview image;   determining distances from each of one or more objects found in the preview image to the location of the voice-controlled camera;   wherein when the one or more objects identified in the preview image are found in the instructions,
 determining a focus point and camera settings; 
 adjusting the focus point and camera settings of the voice-controlled camera to achieve the desired image of the user; and 
 taking the image. 
   
     
     
         41 . The method of  claim 39 , wherein the voice-controlled camera is triggered by the user using a wake word, and the wake word is immediately followed by the instructions. 
     
     
         42 . The method of  claim 40 , wherein to perform the object detection and to provide the context uses one or more of Semantic Segmentation in real-time using Fully Convolutional Networks (FCN), R-CNN (Regional-based Convolutional Neural Networks), Fast R-CNN, Faster R-CNN, YOLO (You Only Look Once), and Mask R-CNN Using TensorRT. 
     
     
         43 . The method of  claim 40 , wherein the distances of the one or more objects to the voice-controlled camera are determined using a depth map, wherein the depth map is obtained using monocular SLAM (Simultaneous Localization and Mapping). 
     
     
         44 . The method of  claim 40 , wherein the focus point and the camera settings are determined by calculating optical formulas for cameras based on one or more detected objects to be photographed, positions of the one or more detected objects in the preview image, and estimated depth or distance of the one or more detected objects to the voice-controlled camera. 
     
     
         45 . The method of  claim 40 , wherein the camera focus point and the camera settings are determined through experimentation by selecting a camera parameter and viewing an image of that selection using depth of field preview, wherein when the image is not good, continuously changing the camera parameter and viewing the image until the image is the desired image. 
     
     
         46 . At least one computer readable medium, comprising a set of instructions, which when executed by one or more computing devices, cause the one or more computing devices to:
 receive, through a microphone of the voice-controlled camera, voice commands from a user, the voice commands comprising instructions for focusing the voice-controlled camera on one or more objects to achieve a desired image of the user; and   convert, by the voice-controlled camera using Artificial Intelligence (AI) scene detection based on natural language processing (NLP), the instructions into keywords that allow the voice-controlled camera to recognize the instructions and to adjust the camera settings to perform the precise focusing of the desired image.   
     
     
         47 . The least one computer readable medium of  claim 46 , wherein after instructions to receive the voice commands, which when executed by the one or more computing devices, further cause the one or more computing devices to:
 simultaneously capture a preview image and perform the NLP on the instructions using deep learning neural networks based on dense vector representations by the voice-controlled camera, the deep learning neural networks including one or more of a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), and a Recursive Neural Network;   apply, by the voice-controlled camera, the AI scene detection on the preview image to perform object detection within the preview image and to provide context as to what is in the preview image;   determine distances from each of the one or more objects found in the preview image to the location of the voice-controlled camera;   wherein when the one or more objects identified in the preview image are found in the instructions,
 determine a focus point and camera settings; 
 adjust the focus point and camera settings of the voice-controlled camera to achieve the desired image of the user; and 
 taking the image. 
   
     
     
         48 . The at least one computer readable medium of  claim 46 , wherein the voice-controlled camera is triggered by the user using a wake word, and the wake word is immediately followed by the instructions. 
     
     
         49 . The at least one computer readable medium of  claim 47 , wherein to perform the object detection and to provide the context uses one or more of Semantic Segmentation in real-time using Fully Convolutional Networks (FCN), R-CNN (Regional-based Convolutional Neural Networks), Fast R-CNN, Faster R-CNN, YOLO (You Only Look Once), and Mask R-CNN Using TensorRT. 
     
     
         50 . The at least one computer readable medium of  claim 47 , wherein the distances of the one or more objects to the voice-controlled camera are determined using a depth map, wherein the depth map is obtained using monocular SLAM (Simultaneous Localization and Mapping).

Join the waitlist — get patent alerts

Track US2022223153A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.