US2018075774A1PendingUtilityA1

Food recognition using visual analysis and speech recognition

Assignee: STANFORD RES INST INTPriority: Jan 7, 2009Filed: Nov 20, 2017Published: Mar 15, 2018
Est. expiryJan 7, 2029(~2.4 yrs left)· nominal 20-yr term from priority
G09B 19/0092
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and system for analyzing at least one food item on a food plate is disclosed. A plurality of images of the food plate is received by an image capturing device. A description of the at least one food item on the food plate is received by a recognition device. The description is at least one of a voice description and a text description. At least one processor extracts a list of food items from the description; classifies and segments the at least one food item from the list using color and texture features derived from the plurality of images; and estimates the volume of the classified and segmented at least one food item. The processor is also configured to estimate the caloric content of the at least one food item.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A system for analyzing food items on a food plate, comprising:
 an image processing module configured to receive one or more visual images of food items on the food plate from one or more image capturing devices;   a meal content determination module configured to classify and segment one or more of the food items on the food plate using one or more features of that food item, wherein one or more features of that food item are derived from one or more images of the food items on the food plate, where the meal content determination module is configured to make use of the visual images of the food items provided by the image processing module, where the meal content determination module is configured to apply an online feature-based segmentation and classification method using one or more food type recognition classifiers trained during offline feature-based learning; and   a food-portion estimation module configured to estimate a volume of the classified and segmented food items, where the food-portion estimation module is configured to cooperate with an object of know dimensions captured in one or more of the visual images received by the image processing module to estimate the volume of the classified and segmented food items, where any portions of the system implemented in software are stored in an executable format in one or more memories to be executed by one or more processors in the system.   
     
     
         22 . The system of  claim 21 , further comprising:
 a nutritional value module is configured to cooperate with the meal content determination module and the food-portion estimation module to determine a caloric content of one or more of the segmented and classified food items and then convey the determined caloric content to a user of the system.   
     
     
         23 . The system of  claim 21 , further comprising:
 a description device configured to understand at least one of i) a voice recognition description and ii) a text recognition description that describes the food items on the food plate and outputs the understanding to the meal content determination module.   
     
     
         24 . The system of  claim 21 , wherein the image capturing device is one of a cell phone or smart phone equipped with a camera, a laptop or desktop computer or workstation equipped with a webcam, or a camera operating in conjunction with a computing platform. 
     
     
         25 . The system of  claim 21 , wherein a first processor is integrated into a voice processing computer, which is one of: directly connected to the image capturing device and remotely over a cell network and/or the Internet. 
     
     
         26 . The system of  claim 21 , further comprising:
 a software model configured to cooperate with the meal content determination module, where the software model is configured to apply an offline feature-based learning method further comprises the steps of:   selecting three or more additional images which capture a same scene of the food plate by receiving a plurality of images of the food plate in addition to the one or more images of the food plate;   color normalizing one or more of the three or more additional images;   
       employing an annotation tool used to identify each food type to generate annotated images of the three or more additional images; and
 processing the color normalized image to extract color and texture features of each of the food items. 
 
     
     
         27 . The system of  claim 21 , wherein the meal content determination module is further configured to process a color normalized image to extract color and texture features of each of the food items by:
 transforming color features to a CIE L*A*B color space;   determining 2D texture features by applying a histogram of orientation gradient (HOG) method; and   placing the color features and 2D texture features into bins of histograms in a higher dimensional space.   
     
     
         28 . The system of  claim 21 , wherein the meal content determination module is further configured to process a color normalized image to extract color and texture features by:
 extracting color and texture features using Texton histograms;   training a set of one-versus-one classifiers between each pair of foods; and   combining color and texture information from the Texton histograms using an Adaboost-based feature selection classifier.   
     
     
         29 . The system of  claim 21 , wherein the one or more features of that food item are a color and texture feature of a first food item, where the color and texture feature of the first food item are selected to apply a multi-class classifier to every patch of three or more images by employing a bootstrap procedure to sample training data and select features simultaneously. 
     
     
         30 . The system of  claim 21 , wherein the food-portion estimation module is further configured to estimate the volume of the classified and segmented food items by:
 capturing a set three or more 2D images from a plurality of images received in addition to the one or more images, where the additional images were taken at different positions above the food plate with a calibrated image capturing device using the object of known dimensions for a 3D scale determination;   extracting and matching multiple feature points in each image frame;   
       estimating camera poses among the three or more images using the matched feature points;
 selecting two or more images from the three or more images; 
 determining correspondences between the two or more images selected from at least the three or more images; 
 performing a 3D reconstruction on the correspondences and determining a 3D scale based on the object of known dimensions to generate a reconstructed 3D point cloud; 
 estimating one or more surfaces of the one or more food items above the food plate based on at least the reconstructed 30 point cloud; and 
 estimating the volume of the one or more food items based on the one or more surfaces. 
 
     
     
         31 . A method for analyzing food items on a food plate, comprising:
 receiving one or more visual images of food items on the food plate by an image processing module;   classifying and segmenting, one or more of the food items on the food plate, with a meal content determination module, by using one or more features of that food item, wherein one or more features of that food item are derived from one or more images of the food items on the food plate, where the meal content determination module is configured to make use of the visual images of the food items provided by the image processing module, where the meal content determination module is configured to apply an online feature-based segmentation and classification method using one or more food type recognition classifiers trained during offline feature-based learning; and   estimating a volume of the classified and segmented food items with a food-portion estimation module, where the food-portion estimation module is configured to cooperate with an object of know dimensions captured in one or more of the visual images received by the image processing module to estimate the volume of the classified and segmented food items, where any portions of the system implemented in software are stored in an executable format in one or more memories to be executed by one or more processors in the system.   
     
     
         32 . The system of  claim 31 , further comprising:
 determining a caloric content of one or more of the segmented and classified food items with a nutritional value module and then conveying the determined caloric content to a user of the system.   
     
     
         33 . The system of  claim 31 , further comprising:
 understanding at least one of a voice recognition device and a text recognition device a description with a description device that describes the food items on the food plate and outputs the understanding to the meal content determination module.   
     
     
         34 . The system of  claim 31 , further comprising:
 applying an offline feature-based learning method further comprises the steps of selecting three or more additional images which capture a same scene of the food plate by receiving a plurality of images of the food plate in addition to the one or more images of the food plate;   color normalizing one or more of the three or more additional images;   employing an annotation tool used to identify each food type to generate annotated images of the three or more additional images; and   processing the color normalized image to extract color and texture features of each of the food items.   
     
     
         35 . The system of  claim 34 , wherein the step of color normalizing comprises detecting a color pattern in the scene. 
     
     
         36 . The system of  claim 31 , further comprising:
 representing one or more food types by a cluster of color and texture features in a high-dimensional space using an incremental K-means clustering method;   representing another one or more food types by Texton histograms; and   classifying the one or more food type using an ensemble of boosted SVM classifiers.   
     
     
         37 . The system of  claim 31 , wherein the meal content determination module is further configured to process a color normalized image to extract color and texture features of each of the food items by:
 transforming color features to a CIE L*A*B color space;   determining 2D texture features by applying a histogram of orientation gradient (HOG) method; and   placing the color features and 2D texture features into bins of histograms in a higher dimensional space.   
     
     
         38 . The system of  claim 31 , wherein the meal content determination module is further configured to process a color normalized image to extract color and texture features by:
 extracting color and texture features using Texton histograms;   training a set of one-versus-one classifiers between each pair of foods; and   combining color and texture information from the Texton histograms using an Adaboost-based feature selection classifier.   
     
     
         39 . The system of  claim 31 , wherein the one or more features of that food item are a color and texture feature of a first food item, where the color and texture feature of the first food item are selected to apply a multi-class classifier to every patch of three or more images by employing a bootstrap procedure to sample training data and select features simultaneously. 
     
     
         40 . The system of  claim 31 , wherein the food-portion estimation module is further configured to estimate the volume of the classified and segmented food items by:
 capturing a set three or more 2D images from a plurality of images received in addition to the one or more images, where the additional images were taken at different positions above the food plate with a calibrated image capturing device using the object of known dimensions for a 3D scale determination;   extracting and matching multiple feature points in each image frame;   estimating camera poses among the three or more images using the matched feature points;   selecting two or more images from the three or more images;   determining correspondences between the two or more images selected from at least the three or more images;   performing a 3D reconstruction on the correspondences and determining a 3D scale based on the object of known dimensions to generate a reconstructed 3D point cloud;   estimating one or more surfaces of the one or more food items above the food plate based on at least the reconstructed 30 point cloud; and   estimating the volume of the one or more food items based on the one or more surfaces.

Join the waitlist — get patent alerts

Track US2018075774A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.