US2013260345A1PendingUtilityA1

Food recognition using visual analysis and speech recognition

Assignee: STANFORD RES INST INTPriority: Jan 7, 2009Filed: Mar 22, 2013Published: Oct 3, 2013
Est. expiryJan 7, 2029(~2.4 yrs left)· nominal 20-yr term from priority
G09B 19/0092
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and system for analyzing at least one food item on a food plate is disclosed. A plurality of images of the food plate is received by an image capturing device. A description of the at least one food item on the food plate is received by a recognition device. The description is at least one of a voice description and a text description. At least one processor extracts a list of food items from the description; classifies and segments the at least one food item from the list using color and texture features derived from the plurality of images; and estimates the volume of the classified and segmented at least one food item. The processor is also configured to estimate the caloric content of the at least one food item.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for analyzing one or more food items on a food plate, the method being executed by at least one processor, comprising the steps of:
 extracting a list of food items from a description of the one or more food items on the foot plate; and   classifying and segmenting the one or more food item from the list to identify the one or more food items using color and texture features derived from one or more images of the food plate by: and   applying an online feature-based segmentation and classification method using one or more food type recognition classifiers trained by offline feature-based learning.   
     
     
         2 . The method of  claim 1 , further comprising the step of estimating the caloric content of the one or more food item. 
     
     
         3 . The method of  claim 1 , wherein the description is one or more of a voice description and a text description. 
     
     
         4 . The method of  claim 1 , wherein applying an offline feature-based learning method further comprises the steps of:
 selecting three or more images which capture the same scene of the food plate by receiving a plurality of images of the food plate in addition to the one or more images;   color normalizing one or more of the three or more images;   employing an annotation tool used to identify each food type; and   processing the color normalized image to extract color and texture features of each of the food items.   
     
     
         5 . The method of  claim 4 , wherein the step of color normalizing comprises detecting a color pattern in the scene. 
     
     
         6 . The method of  claim 4 , wherein the step of processing the color normalized image to extract color and texture features of each of the food items further comprises the steps of:
 transforming color features to a CIE L*A*B color space;   determining 2D texture features by applying a histogram of orientation gradient (HOG) method; and   placing the color features and 2D texture features into bins of histograms in a higher dimensional space.   
     
     
         7 . The method of  claim 4 , wherein the step of processing the color normalized image to extract color and texture features further comprises the steps of:
 extracting color and texture features using Texton histograms;   training a set of one-versus-one classifiers between each pair of foods; and   combining color and texture information from the Texton histograms using an Adaboost-based feature selection classifier.   
     
     
         8 . The method of  claim 1 , wherein the step of applying an online feature-based segmentation and classification method further comprises the steps of:
 selecting three or more images of a plurality of images received in addition to the one or more images of the food plate, the three or more images capturing the same scene of the food plate;   color normalizing the three or more images;   locating the food plate using a contour based circle detection method; and   processing the color normalized image to extract color and texture features.   
     
     
         9 . The method of  claim 8 , wherein the step of color normalizing comprises detecting a color pattern in the scene. 
     
     
         10 . The method of  claim 8 , further comprising the steps of:
 representing one or more food types by a cluster of color and texture features in a high-dimensional space using an incremental K-means clustering method;   representing another one or more food types by Texton histograms; and   classifying the one or more food type using an ensemble of boosted SVM classifiers.   
     
     
         11 . The method of  claim 8 , wherein the step of applying an online feature-based segmentation and classification method further comprises the steps of:
 applying a k-nearest neighbors (k-NN) classification method to the extracted color and texture features to each pixel of the color normalized image and assigning at least one label to each pixel;   applying a dynamic assembled multi-class classifier to an extracted color and texture feature for each patch of the color normalized image and assigning one label to each patch; and   applying an image segmentation technique to obtain a final segmentation of the plate into its constituent food labels.   
     
     
         12 . The method of  claim 8 , wherein the step of applying an online feature-based segmentation and classification method further comprises the steps of:
 applying a multi-class classifier to every patch of the three input images to generate a segmentation map; and   dynamically assembling a multi-class classifier from a subset of the offline trained pair-wise classifiers to assign a small set of labels to each pixel of the three images.   
     
     
         13 . The method of  claim 12 , wherein features are selected for applying a multi-class classifier to every patch of the three input images by employing a bootstrap procedure to sample training data and select features simultaneously. 
     
     
         14 . The method of  claim 13 , wherein the bootstrap procedure comprises the steps of: randomly sampling a set of training data and computing all features in feature pool;
 training individual SVM classifiers;   applying a 2-fold validation process to evaluate the expected normalized margin for each feature to update the strong classifier;   applying a current strong classifier to densely sampled patches in the annotated images, wherein wrongly classified patches are added as new samples, and weights of all training samples are updated; and   stopping the training if the number of wrongly classified patches in the training images falls below a predetermined threshold.   
     
     
         15 . The method of  claim 1 , wherein the step of estimating volume of the classified and segmented one or more food items further comprises the steps of:
 capturing a set three or more 2D images from a plurality of images received in addition to the one or more images taken at different positions above the food plate with a calibrated image capturing device using an object of known size for 3D scale determination;   extracting and matching multiple feature points in each image frame;   estimating camera poses among the three or more images using the matched feature points;   selecting two or more images from the three or more images;   determining correspondences between the two or more images selected from at least the three or more images;   performing a 3D reconstruction on the correspondences and determining a 3D scale based on the object of known size to generate 3D point cloud;   estimating one or more surfaces of the one or more food items above the food plate based on at least the reconstructed 3D point cloud: and   estimating the volume of the one or more food items based on the one or more surfaces.   
     
     
         16 . A system for analyzing one or more food items on a food plate, comprising:
 a processor for:   classifying and segmenting the one or more food items using one or more features of the one or more food items derived from one or more images, by applying an online feature-based segmentation and classification method using one or more food type recognition classifiers trained during offline feature-based learning; and   estimating the volume of the classified and segmented one or more food items based on determining correspondences between images of the one or more images containing the one or more food items.   
     
     
         17 . The system of  claim 16 , wherein the processor further:
 receives a description from a device describing the one or more food items on the food plate,   wherein the description device is at least one of a voice recognition device and a text recognition device and also supplies the one or more images.   
     
     
         18 . The system of  claim 17 , wherein the image capturing device is one of a cell phone or smart phone equipped with a camera, a laptop or desktop computer or workstation equipped with a webcam, or a camera operating in conjunction with a computing platform. 
     
     
         19 . The system of  claim 17 , wherein the processor is integrated into a voice processing computer, which is one of: directly connected to the image capturing device and remotely over a cell network and/or the Internet. 
     
     
         20 . The system of  claim 17 , wherein the processor estimates the caloric content of the one or more food items.

Join the waitlist — get patent alerts

Track US2013260345A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.