US2023067026A1PendingUtilityA1

Automated data analytics methods for non-tabular data, and related systems and apparatus

Assignee: DATAROBOT INCPriority: Feb 17, 2020Filed: Feb 17, 2021Published: Mar 2, 2023
Est. expiryFeb 17, 2040(~13.6 yrs left)· nominal 20-yr term from priority
G06V 10/771G06V 10/82G06F 18/2113G06Q 30/0206G06F 18/214G06V 10/454G06V 20/00G06V 30/422G06F 18/24133G06Q 50/16G06F 18/24143G06F 18/40G06Q 10/04
37
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Automated data analytics techniques for non-tabular data sets may include methods and systems for (1) automatically developing models that perform tasks in the domains of computer vision, audio processing, speech processing, text processing, or natural language processing; (2) automatically developing models that analyze heterogeneous data sets containing image data and non-image data, and/or heterogeneous data sets containing tabular data and non-tabular data; (3) determining the importance of an image feature with respect to a modeling task, (4) explaining the value of a modeling target based at least in part on an image feature, and (5) detecting drift in image data. In some cases, multi-stage models may be developed, wherein a pre-trained feature extraction model extracts low-, mid-, high-, and/or highest-level features of non-tabular data, and a data analytics models uses those features (or features derived therefrom) to perform a data analytics task.

Claims

exact text as granted — not AI-modified
1 . A method for determining an importance of an aggregate image feature, the method comprising:
 obtaining a plurality of data samples, wherein each of the plurality of data samples is associated with respective values for a set of features and with a respective value for a target, wherein the set of features includes a feature having an aggregate image data type, and wherein the feature having the aggregate image data type comprises a plurality of features each having a constituent image data type;   for each of the plurality of constituent image features, determining a feature importance score indicating an expected utility of the constituent image feature for predicting the values of the target; and   determining a feature importance score for the aggregate image feature based on the feature importance scores of the constituent image features, wherein the feature importance score for the aggregate image feature indicates an expected utility of the aggregate image feature for predicting the values of the target.   
     
     
         2 . The method of  claim 1 , wherein the aggregate image feature comprises an image feature vector. 
     
     
         3 . The method of  claim 1 , wherein the feature importance score comprises a univariate feature importance score, a feature impact score, or a Shapley value. 
     
     
         4 . The method of  claim 1 , further comprising:
 prior to determining the feature importance score for the aggregate image feature based on the feature importance scores of the constituent image features, normalizing and/or standardizing the feature importance scores for the constituent image features.   
     
     
         5 . The method of  claim 1 , further comprising, for each data sample of the plurality of data samples:
 extracting respective values for the plurality of constituent image features from a first plurality of images using a pre-trained image processing model.   
     
     
         6 . The method of  claim 5 , wherein the pre-trained image processing model comprises a pre-trained image feature extraction model or a pre-trained, fine-tunable image processing model. 
     
     
         7 . The method of  claim 5 , wherein the pre-trained image processing model comprises a convolutional neural network model previously trained on a training data set comprising a second plurality of images. 
     
     
         8 . The method of  claim 1 , wherein determining the feature importance score for the aggregate image feature comprises selecting a highest feature importance score among the feature importance scores for the constituent image features, and using the selected highest feature importance score as the feature importance score for the aggregate image feature. 
     
     
         9 . The method of  claim 1 , wherein the set of features further includes a feature having a non-image data type, and wherein the method further comprises:
 quantitatively comparing a feature importance score of the feature having the non-image data type with the feature importance score of the aggregate image feature; and   determining, based on the quantitative comparison, whether the non-image feature or the aggregate image feature has greater expected utility for predicting the values of the target.   
     
     
         10 . An image-based data analytics method, comprising:
 obtaining inference data, wherein the inference data include image data;   extracting, by an image feature extraction model, respective values of a plurality of constituent image features derived from the image data; and   determining a value of a data analytics target based on the values of the plurality of constituent image features, wherein the determining is performed by a trained machine learning model.   
     
     
         11 . The method of  claim 10 , wherein the image feature extraction model is pre-trained. 
     
     
         12 . The method of  claim 10 , wherein the image feature extraction model comprises a convolutional neural network. 
     
     
         13 . The method of  claim 10 , wherein the plurality of constituent image features include one or more low-level image features, one or more mid-level image features, one or more high-level image features, and/or one or more highest-level image features. 
     
     
         14 . The method of  claim 10 , wherein the inference data further include non-image data. 
     
     
         15 . The method of  claim 14 , wherein the determining of the value of the data analytics target is also based on values of one or more features derived from the non-image data. 
     
     
         16 . The method of  claim 15 , further comprising arranging the values of the constituent image features and the values of the features derived from the non-image data in a table, wherein the determining of the value of the data analytics target is performed by applying the trained machine learning model to the table. 
     
     
         17 . The method of  claim 15 , wherein the image feature extraction model is not fitted to the values of the plurality of constituent image features derived from the image data. 
     
     
         18 . The method of  claim 17 , wherein the trained machine learning model includes a gradient boosting machine. 
     
     
         19 . The method of  claim 15 , wherein the value of the data analytics target includes a prediction based on the inference data, a description of the inference data, a classification associated with the inference data, and/or a label associated with the inference data. 
     
     
         20 . A model development system comprising:
 an image feature extraction module operable to extract values of one or more image feature candidates from image data;   a data preparation and feature engineering module operable to obtain values of one or more of a plurality of features based, at least in part, on the values of the image feature candidates; and   a model creation and evaluation module operable to generate and evaluate one or more machine learning models trained to determine a value of a data analytics target based on the values of the plurality of features.   
     
     
         21 - 47 . (canceled)

Join the waitlist — get patent alerts

Track US2023067026A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.