Forecasting model generation for sample biased data set
Abstract
A method, apparatus, system, and computer program product for creating a forecasting model for payroll records. Payroll records are received for a group of employers. The payroll records comprise granular data parameters about employees of the group of employers. A forecasting model is created the that aligns the payroll records to high-level employment data. Creating the forecasting model includes identifying predictor variables from the granular data parameters of the payroll records. Creating the forecasting model includes generating a set of basis functions from the predictor variables. Creating the forecasting model includes combining the set of basis functions to create the forecasting model.
Claims
exact text as granted — not AI-modified1 .- 21 . (canceled)
22 . A method comprising:
receiving, by a data processing system comprising a processor coupled with memory, a first data set having a first granularity that is greater than a second granularity of a second data set, comprising data parameters of the first data set; identifying, by the data processing system, variables of the data parameters to predict granularity of the second data set; generating, by the data processing system, a set of basis functions from the variables to build a forecasting model to map nonlinearities of the variables; selecting, by the data processing system, a plurality of basis functions from the set of basis functions to control a residual error of the forecasting model based on a threshold; generating, by the data processing system, a spline of the forecasting model based on the plurality of basis functions; and predicting, by the data processing system using the spline of the forecasting model, data parameters of the second data set that is less granular than the first data set.
23 . The method of claim 22 , wherein the first data set comprises payroll records.
24 . The method of claim 22 , comprising:
aligning, by the data processing system using the forecasting model, the data parameters of the first data set to the second data set.
25 . The method of claim 22 , comprising:
creating, by the data processing system, the forecasting model using a non-parametric regression analysis technique.
26 . The method of claim 22 , comprising:
modeling, by the data processing system, using a non-parametric regression analysis technique, nonlinearities and interactions between the variables using a plurality of multivariate adaptive regression splines.
27 . The method of claim 22 , comprising:
identifying, by the data processing system using a plurality of multivariate adaptive regression splines, the variables.
28 . The method of claim 22 , comprising:
combining, by the data processing system, the plurality of basis functions.
29 . The method of claim 22 , comprising:
removing, by the data processing system, a selected basis function of the plurality of basis functions to maintain the residual error of the forecasting model below a second threshold.
30 . The method of claim 22 , wherein the threshold can be at least one of:
the residual error too small to reduce further; or a maximum number of basis functions.
31 . The method of claim 22 , comprising:
predicting, by the data processing system using the spline of the forecasting model, the data parameters of the second data set at a higher granularity than comprised by the second data set.
32 . A system, comprising:
a data processing system comprising a processor, coupled with memory, to: receive a first data set having a first granularity that is greater than a second granularity of a second data set, comprising data parameters of the first data set; identify variables of the data parameters to predict granularity of the second data set; generate a set of basis functions from the variables to build a forecasting model to map nonlinearities of the variables; select a plurality of basis functions from the set of basis functions to control a residual error of the forecasting model based on a threshold; generate a spline of the forecasting model based on the plurality of basis functions; and predict, using the spline of the forecasting model, data parameters of the second data set that is less granular than the first data set.
33 . The system of claim 32 , wherein the data processing system is configured to:
combine the plurality of basis functions.
34 . The system of claim 32 , wherein the threshold can be at least one of:
the residual error too small to reduce further; or a maximum number of basis functions.
35 . The system of claim 32 , wherein the data processing system is configured to:
create the forecasting model using a non-parametric regression analysis technique.
36 . The system of claim 32 , wherein the data processing system is configured to:
model, using a non-parametric regression analysis technique, nonlinearities and interactions between the variables using a plurality of multivariate adaptive regression splines.
37 . The system of claim 32 , wherein the data processing system is configured to:
identify, using a plurality of multivariate adaptive regression splines, the variables.
38 . The system of claim 32 , wherein the data processing system is further configured to:
remove a selected basis function of the plurality of basis functions to maintain the residual error of the forecasting model at a second threshold.
39 . The system of claim 32 , wherein the data processing system is further configured to:
predict, using the spline of the forecasting model, the data parameters of the second data set at a higher granularity than comprised by the second data set.
40 . A non-transitory computer-readable medium storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to:
receive a first data set having a first granularity that is greater than a second granularity of a second data set, comprising data parameters of the first data set; identify variables of the data parameters to predict granularity of the second data set; generate a set of basis functions from the variables to build a forecasting model to map nonlinearities of the variables; select a plurality of basis functions from the set of basis functions to control a residual error of the forecasting model based on a threshold; generate a spline of the forecasting model based on the plurality of basis functions; and predict, using the spline of the forecasting model, data parameters of the second data set that is less granular than the first data set.
41 . The non-transitory computer-readable medium of claim 40 , wherein the instructions further comprise instructions to:
create the forecasting model using a non-parametric regression analysis technique; and model, using the non-parametric regression analysis technique, nonlinearities and interactions between the variables using a plurality of multivariate adaptive regression splines.Join the waitlist — get patent alerts
Track US2023230040A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.