Optical and other sensory processing of complex objects
Abstract
Systems and methods for optical and other sensory analysis of nutritional and other complex objects are disclosed. For example, techniques may include capturing an RGB-D image of a food using an integrated camera; inputting the RGB-D image into an instance detection network configured to detect food items; segmenting a plurality of food items from the RGB-D image into a plurality of masks, the plurality of masks representing individual food items; classifying a particular food item among the individual food items using a multimodal large language model; estimating a volume of the particular food item by overlaying an RGB image associated with the RGB-D image with a depth-map to create a point cloud; and estimating the calories of the particular food item using the estimated volume and a nutritional database.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A non-transitory computer-readable medium configured to use an integrated camera to scan food, the non-transitory computer-readable medium storing executable code configured to perform operations comprising:
capturing an RGB image of the food using the integrated camera;
inputting the RGB image into an instance detection network configured to detect food items, the instance detection network having been trained on a plurality of reference food items using a neural network;
segmenting a plurality of food items from the RGB image into a plurality of masks, the plurality of masks representing individual food items;
classifying a particular food item among the individual food items using a multimodal large language model;
estimating a volume of the particular food item by overlaying the RGB image with a depth-map to create a point cloud, wherein the depth map is created by monocular depth estimation; and
estimating the calories of the particular food item using the estimated volume and a nutritional database.
2. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise comparing a first RGB image before intake to a second RGB image after intake, and generating an intake estimate associated with the particular food item.
3. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise comparing a first RGB image after intake to a synthetic image.
4. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise pairing an unlabeled image with a verified mask from the plurality of masks representing an individual food item to optimize a trained classifier.
5. The non-transitory computer-readable medium of claim 1 , wherein inputting the RGB image into the instance detection network comprises creating a square around the individual food item.
6. The non-transitory computer-readable medium of claim 1 , wherein inputting the RGB image into the instance detection network comprises creating a mask representing multiple food items.
7. The non-transitory computer-readable medium of claim 1 , wherein the point cloud is a synthetic point cloud capable of being captured and stored for analysis.
8. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:
capturing a video segment using the integrated camera; and
extracting a plurality of image frames from the video segment.
9. The non-transitory computer-readable medium of claim 8 , wherein the extracted plurality of image frames includes the captured RGB image.
10. A computer-implemented method for scanning food, comprising:
capturing an RGB image of the food using an integrated camera;
inputting the RGB image into an instance detection network configured to detect food items;
segmenting a plurality of food items from the RGB image into a plurality of masks, the plurality of masks representing individual food items;
classifying a particular food item among the individual food items using a multimodal large language model;
estimating a volume of the particular food item by overlaying the RGB image with a depth-map to create a point cloud, wherein the depth map is created by monocular depth estimation; and
estimating the calories of the particular food item using the estimated volume and a nutritional database.
11. The computer-implemented method of claim 10 , further comprising:
comparing a first RGB image before food intake to a second RGB image after food intake, and
generating, based on the comparison, a food intake estimate associated with the particular food item.
12. The computer-implemented method of claim 10 , further comprising:
comparing a first RGB image after food intake to a synthetic image.
13. The computer-implemented method of claim 10 , further comprising:
pairing an unlabeled image with a verified mask from the plurality of masks representing an individual food item to update a trained classifier.
14. The computer-implemented method of claim 10 , wherein inputting the RGB image into the instance detection network includes at least one of:
creating a square around the individual food item; or
creating a mask representing multiple food items.
15. The computer-implemented method of claim 10 , further comprising:
creating training data for synthetic data generation through automatic dataset generation via an assets database; and
generating segmented outputs from the synthetic data generation across various combinations of food items, lighting conditions, and other physical environments.
16. The computer-implemented method of claim 15 , wherein the segmented outputs are used for volume and calorie estimation.
17. The computer-implemented method of claim 10 , further comprising:
capturing a video segment using the integrated camera; and
extracting a plurality of image frames from the video segment.
18. The computer-implemented method of claim 17 , wherein the extracted plurality of image frames includes the captured RGB image.
19. The computer-implemented method of claim 10 , wherein the monocular depth estimation comprises estimating a distance of a plurality of pixels in the RGB image from the integrated camera.
20. The computer-implemented method of claim 10 , wherein the monocular depth estimation comprises predicting depth information of the particular food item based on the RGB image.Join the waitlist — get patent alerts
Track US12182982B1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.