Machine learning modeling of candidate models for software installation usage
Abstract
This document discloses methods and systems for modeling product usage. In one practical application, the systems and methods may be utilized to model product usage based on large volume, machine generated product usage data to optimize product pricing and operations. Specifically, the systems and methods described herein may utilize methods with key components to select the maximum number of dimensions that can be modeled based on the number of data points, use a logarithm kernel function to normalize machine data with long-tailed statistical distributions on different numerical scales, compare a large number of candidate models with different candidate dimensions and different structures, and quantify the amount of change and drift in models over time.
Claims
exact text as granted — not AI-modifiedThe invention claimed is:
1. A computing device, comprising:
one or more hardware processors; and
a non-transitory computer-readable medium having stored thereon instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations including:
generating a first number (N) of normalized data points from N data points, each data point of the N data points having data for a second number (C) of dimensions, by taking a logarithm of the data for each dimension of the C dimensions for each data point of the N data points, and wherein the N data points relate to product usage data;
determining, based on a predetermined threshold and the N normalized data points, a third number (D) of dimensions to use for modeling the N normalized data points;
selecting a plurality of dimension sets, wherein each dimension set includes a different combination of a plurality of dimensions and includes no greater than D dimensions;
for each dimension set in the plurality of dimension sets, generating, via a machine learning model, a candidate model based on the different combination of the plurality of dimensions in the dimension set, wherein each candidate model models product usage;
evaluating each candidate model and generating respective model testing results for each candidate model;
selecting a model from the candidate models based on a plurality of model quality measures from model testing results associated with the candidate models.
2. The computing device of claim 1 , wherein the operations further comprise:
receiving a request to identify a predicted value for an additional data point using the selected model, the additional data point including data on dimensions corresponding to the dimension set corresponding to the selected model; and
based on the additional data point and the generated model, responding to the request with the predicted value.
3. The computing device of claim 1 , wherein the operations further comprise:
determining, for a candidate number of dimensions representing the D dimensions, a number of bits of information (B) per dimension based on the candidate number of dimensions and the first number; and
comparing B to the predetermined threshold.
4. The computing device of claim 3 , wherein the determining of B uses
B
=
1
D
log
2
(
N
)
.
5. The computing device of claim 3 , wherein the operations further comprise:
based on a result of the comparing, rejecting the candidate number of dimensions.
6. The computing device of claim 1 , wherein the generating of the candidate model for each dimension set comprises generating a polynomial model.
7. The computing device of claim 1 , wherein the operations further comprise:
generating the N data points by linking customer relational information from a customer relationship management (CRM) system and the product usage data using shared identifiers.
8. The computing device of claim 7 , wherein the operations further comprise:
accessing CRM data that indicates a parent-subsidiary relationship between a parent account and a subsidiary account; and
accessing first product usage data that is linked to both the parent account and the subsidiary account; wherein
the generating of the N data points comprises generating a data point that links the first product usage data to the subsidiary account.
9. The computing device of claim 8 , wherein the generating of the N data points excludes generating a second data point that links the first product usage data to the parent account.
10. A computer-implemented method, comprising:
generating a first number (N) of normalized data points from N data points, each data point of the N data points having data for a second number (C) of dimensions, by taking a logarithm of the data for each dimension of the C dimensions for each data point of the N data points, and wherein the N data points relate to product usage data;
determining, based on a predetermined threshold and the N normalized data points, a third number (D) of dimensions to use for modeling the N normalized data points;
selecting a plurality of dimension sets, wherein each dimension set includes a different combination of a plurality of dimensions and includes no greater than D dimensions;
for each dimension set in the plurality of dimension sets, generating, via a machine learning model, a candidate model based on the different combination of the plurality of dimensions in the dimension set, wherein each candidate model models product usage;
evaluating each candidate model and generating respective model testing results for each candidate model;
selecting a model from the candidate models based on a plurality of model quality measures from model testing results associated with the candidate models, wherein the selected model models product usage.
11. The computer-implemented method of claim 10 , further comprising:
receiving a request to identify a predicted value for an additional data point using the selected model, the additional data point including data on dimensions corresponding to the dimension set corresponding to the selected model; and
based on the additional data point and the generated model, responding to the request with the predicted value.
12. The computer-implemented method of claim 10 , further comprising:
determining, for a candidate number of dimensions representing the D dimensions, a number of bits of information (B) per dimension based on the candidate number of dimensions and the first number; and
comparing B to the predetermined threshold.
13. The computer-implemented method of claim 12 , wherein the determining of B uses
B
=
1
D
log
2
(
N
)
.
14. The computer-implemented method of claim 12 , further comprising:
based on a result of the comparing, rejecting the candidate number of dimensions.
15. The computer-implemented method of claim 10 , wherein the generating of the candidate model for each dimension set comprises generating a polynomial model.
16. The computer-implemented method of claim 10 , further comprising:
generating the N data points by linking customer relational information from a customer relationship management (CRM) system and the product usage data using shared identifiers.
17. The computer-implemented method of claim 16 , further comprising:
accessing CRM data that indicates a parent-subsidiary relationship between a parent account and a subsidiary account; and
accessing first product usage data that is linked to both the parent account and the subsidiary account; wherein
the generating of the N data points comprises generating a data point that links the first product usage data to the subsidiary account.
18. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform operations including:
generating a first number (N) of normalized data points from N data points, each data point of the N data points having data for a second number (C) of dimensions, by taking a logarithm of the data for each dimension of the C dimensions for each data point of the N data points, and wherein the N data points relate to product usage data;
determining, based on a predetermined threshold and the N normalized data points, a third number (D) of dimensions to use for modeling the N normalized data points;
selecting a plurality of dimension sets, wherein each dimension set includes a different combination of a plurality of dimensions and includes no greater than D dimensions;
for each dimension set in the plurality of dimension sets, generating, via a machine learning model, a candidate model based on the different combination of the plurality of dimensions in the dimension set, wherein each candidate model models product usage;
evaluating each candidate model and generating respective model testing results for each candidate model;
selecting a model from the candidate models based on a plurality of model quality measures from model testing results associated with the candidate models, wherein the selected model models product usage.
19. The non-transitory computer-readable medium of claim 18 , wherein the operations further comprise:
receiving a request to identify a predicted value for an additional data point using the selected model, the additional data point including data on dimensions corresponding to the dimension set corresponding to the selected model; and
based on the additional data point and the generated model, responding to the request with the predicted value.
20. The non-transitory computer-readable medium of claim 18 , wherein the operations further comprise:
determining, for a candidate number of dimensions representing the D dimensions, a number of bits of information (B) per dimension based on the candidate number of dimensions and the first number; and
comparing B to the predetermined threshold.Join the waitlist — get patent alerts
Track US12181999B1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.