US12182982B1ActiveUtility

Optical and other sensory processing of complex objects

Assignee: CALODAR LTDPriority: Jan 8, 2024Filed: Apr 24, 2024Granted: Dec 31, 2024
Est. expiryJan 8, 2044(~17.5 yrs left)· nominal 20-yr term from priority
Inventors:Joseph Swed
G06T 7/62G06T 7/11G06V 10/7753G06V 10/82G06T 2207/10016G06T 2207/10028G06T 2207/20081G06T 2207/20084G06T 2207/30128G06T 2207/10024G06V 20/68G06T 7/0002
56
PatentIndex Score
0
Cited by
37
References
20
Claims

Abstract

Systems and methods for optical and other sensory analysis of nutritional and other complex objects are disclosed. For example, techniques may include capturing an RGB-D image of a food using an integrated camera; inputting the RGB-D image into an instance detection network configured to detect food items; segmenting a plurality of food items from the RGB-D image into a plurality of masks, the plurality of masks representing individual food items; classifying a particular food item among the individual food items using a multimodal large language model; estimating a volume of the particular food item by overlaying an RGB image associated with the RGB-D image with a depth-map to create a point cloud; and estimating the calories of the particular food item using the estimated volume and a nutritional database.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A non-transitory computer-readable medium configured to use an integrated camera to scan food, the non-transitory computer-readable medium storing executable code configured to perform operations comprising:
 capturing an RGB image of the food using the integrated camera; 
 inputting the RGB image into an instance detection network configured to detect food items, the instance detection network having been trained on a plurality of reference food items using a neural network; 
 segmenting a plurality of food items from the RGB image into a plurality of masks, the plurality of masks representing individual food items; 
 classifying a particular food item among the individual food items using a multimodal large language model; 
 estimating a volume of the particular food item by overlaying the RGB image with a depth-map to create a point cloud, wherein the depth map is created by monocular depth estimation; and 
 estimating the calories of the particular food item using the estimated volume and a nutritional database. 
 
     
     
       2. The non-transitory computer-readable medium of  claim 1 , wherein the operations further comprise comparing a first RGB image before intake to a second RGB image after intake, and generating an intake estimate associated with the particular food item. 
     
     
       3. The non-transitory computer-readable medium of  claim 1 , wherein the operations further comprise comparing a first RGB image after intake to a synthetic image. 
     
     
       4. The non-transitory computer-readable medium of  claim 1 , wherein the operations further comprise pairing an unlabeled image with a verified mask from the plurality of masks representing an individual food item to optimize a trained classifier. 
     
     
       5. The non-transitory computer-readable medium of  claim 1 , wherein inputting the RGB image into the instance detection network comprises creating a square around the individual food item. 
     
     
       6. The non-transitory computer-readable medium of  claim 1 , wherein inputting the RGB image into the instance detection network comprises creating a mask representing multiple food items. 
     
     
       7. The non-transitory computer-readable medium of  claim 1 , wherein the point cloud is a synthetic point cloud capable of being captured and stored for analysis. 
     
     
       8. The non-transitory computer-readable medium of  claim 1 , wherein the operations further comprise:
 capturing a video segment using the integrated camera; and 
 extracting a plurality of image frames from the video segment. 
 
     
     
       9. The non-transitory computer-readable medium of  claim 8 , wherein the extracted plurality of image frames includes the captured RGB image. 
     
     
       10. A computer-implemented method for scanning food, comprising:
 capturing an RGB image of the food using an integrated camera; 
 inputting the RGB image into an instance detection network configured to detect food items; 
 segmenting a plurality of food items from the RGB image into a plurality of masks, the plurality of masks representing individual food items; 
 classifying a particular food item among the individual food items using a multimodal large language model; 
 estimating a volume of the particular food item by overlaying the RGB image with a depth-map to create a point cloud, wherein the depth map is created by monocular depth estimation; and 
 estimating the calories of the particular food item using the estimated volume and a nutritional database. 
 
     
     
       11. The computer-implemented method of  claim 10 , further comprising:
 comparing a first RGB image before food intake to a second RGB image after food intake, and 
 generating, based on the comparison, a food intake estimate associated with the particular food item. 
 
     
     
       12. The computer-implemented method of  claim 10 , further comprising:
 comparing a first RGB image after food intake to a synthetic image. 
 
     
     
       13. The computer-implemented method of  claim 10 , further comprising:
 pairing an unlabeled image with a verified mask from the plurality of masks representing an individual food item to update a trained classifier. 
 
     
     
       14. The computer-implemented method of  claim 10 , wherein inputting the RGB image into the instance detection network includes at least one of:
 creating a square around the individual food item; or 
 creating a mask representing multiple food items. 
 
     
     
       15. The computer-implemented method of  claim 10 , further comprising:
 creating training data for synthetic data generation through automatic dataset generation via an assets database; and 
 generating segmented outputs from the synthetic data generation across various combinations of food items, lighting conditions, and other physical environments. 
 
     
     
       16. The computer-implemented method of  claim 15 , wherein the segmented outputs are used for volume and calorie estimation. 
     
     
       17. The computer-implemented method of  claim 10 , further comprising:
 capturing a video segment using the integrated camera; and 
 extracting a plurality of image frames from the video segment. 
 
     
     
       18. The computer-implemented method of  claim 17 , wherein the extracted plurality of image frames includes the captured RGB image. 
     
     
       19. The computer-implemented method of  claim 10 , wherein the monocular depth estimation comprises estimating a distance of a plurality of pixels in the RGB image from the integrated camera. 
     
     
       20. The computer-implemented method of  claim 10 , wherein the monocular depth estimation comprises predicting depth information of the particular food item based on the RGB image.

Join the waitlist — get patent alerts

Track US12182982B1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.