Systems and methods for staffing master controller
Abstract
Apparatuses, systems, and methods described herein include receiving by a quantum hybrid branch simulation engine, a request for a next task by a branch personnel or worker. The branch simulation engine may receive data representing a simulated branch environment and apply a deep reinforcement learning (DRL) staffing operations model to predict a next action or task. The model may weigh impacts of a reward to substantially maximize the reward in light of a policy to determine the next action and update a database to store/represent a next state of the simulated or real-life branch environment after the next action is performed.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method, comprising:
receiving by a branch simulation engine, from a database, a first dataset that represents a simulated branch environment; applying by the branch simulation engine, a deep reinforcement learning (DRL) staffing operations model to the first dataset to predict a next action or task for a branch worker, wherein the DRL staffing operations model weighs impacts of a reward to substantially maximize the reward in light of a policy to determine the next action; iteratively updating, by the branch simulation engine, the database that includes the first dataset to include a second dataset to represent a next state of the simulated branch environment after the next action is performed; providing one or more policies including the policy or instructions learned by the branch simulation engine from the DRL staffing operations model to a master DRL agent, wherein the master DRL agent is configured to recommend a real-life action in a real-life branch environment and comprises at least a first neural network (NN) and a second neural network (NN), wherein the first NN is coupled to receive an input from the second NN, and wherein the second NN is a classical NN that computes a real-life reward and a probability of a long-term real-life reward from the real-life action, wherein the master DRL agent generates an output configured to cause a computing device to dispatch a digital work assignment corresponding to a selected real-life action determined based on the computed probability of the real-life reward or long-term real-life reward.
2 . The computer-implemented method of claim 1 wherein the branch simulation engine computes and ranks a plurality of next possible actions or tasks that are performed by a simulated branch worker according to the policy.
3 . The computer-implemented method of claim 1 wherein the branch simulation engine comprises a quantum hybrid simulation engine including classical computing devices to assist in preparing data coupled to a quantum computing device that assists in determining an effect of the action or task on the simulated branch environment.
4 . The computer-implemented method of claim 1 , wherein the policy is determined by a policy machine of the branch simulation engine trained to output a next action for a real-life branch agent, where inputs to the policy machine include data representing a current real-world branch environment.
5 . The computer-implemented method of claim 1 , wherein the policy comprises a policy of a staffing organization and/or branch and includes variables representing objectives and values such as, at least, one or more of maximizing profit, worker safety, customer satisfaction, or environmental impact.
6 . The computer-implemented method of claim 1 , wherein the first and the second dataset include attributes representing a branch environment comprising one or more of open job orders, high value or lower value customers, worker availability, customer preferences, or current pricing strategies.
7 . The computer-implemented method of claim 1 wherein a reward network of the DRL staffing operations model implements a reward function and updates the staffing operations model based upon a computed reward value impact on the policy, including a policy of the staffing organization.
8 . A system, comprising:
one or more processors; and a memory that stores machine-readable instructions that when executed by the one or more processors cause the system to: receive a request to generate an instruction for a next task in a real-life branch environment; access a first dataset including attributes of a stored simulated branch environment similar to a second dataset of attributes of the real-life branch environment; based at least in part on the first dataset and the second dataset, request or cause generation of a next action or sequence of actions for a real-life worker by a master staffing engine, wherein the master staffing engine is trained on policies and instructions learned from a branch simulation engine that predicts a plurality of simulated branch environments similar to the real-life branch environment, wherein the master staffing engine comprises a master deep reinforcement learning (DRL) agent that includes at least a first neural network (NN) and a second neural network (NN) and wherein the first NN is coupled to receive an input from the second NN and wherein the second NN is a classical NN that computer a real-life reward and a probability of a long-term real-life reward from the next action in the real-life branch environment; generate, for transmission to a computing device used by the real-life worker, a data structure that causes the computer device to render a dynamically updated work queue on a display, wherein the work queue comprises an ordered list of potential real-life actions, and wherein a priority of each of the potential real-life actions within the ordered list is determined based on the computed real-life reward or probability of the long-term real-life reward.
9 . The system of claim 8 wherein the classical neural network is trained via reinforcement learning to determine a reward based on a reward function related to customer satisfaction, high-value customer satisfaction, worker retention, customer complaints, or worker utilization.
10 . The system of claim 8 wherein the one or more processors are coupled to or includes a quantum processing unit (QPU).
11 . The system of claim 8 wherein the system includes a hybrid-classical quantum neural network including a quantum portion including quantum layers.
12 . A method, comprising:
receiving a request from a real-life worker or branch to generate by a master staffing engine, an instruction for a next task in a real-life branch environment; loading into the master staffing engine, a first dataset including attributes of a stored simulated branch environment similar to a second dataset of attributes of the real-life branch environment; based at least in part on the first dataset and the second dataset, causing generation by the master staffing engine, of a next action or sequence of actions for the real-life worker, wherein the master staffing engine is trained on policies and instructions learned from a branch simulation engine that predicts a plurality of simulated branch environments similar to the real-life branch environment, wherein the master staffing engine comprises a master DRL agent that includes at least a first neural network (NN) and a second neural network (NN) and wherein the first NN is coupled to receive an input from the second NN and wherein the second NN is a classical NN that computer a real-life reward and a probability of a long-term real-life reward from the next action in the real-life branch environment; and generating an output configured to cause a computing device to dispatch a digital work assignment corresponding to a selected real-life action determined based on the computer probability of the real-life reward or long-term real-life reward.
13 . The method of claim 12 , wherein the real-life reward is computed using a reward function based upon a real-life result.
14 . The method of claim 12 , further comprising comparing the real-life branch environment to the stored simulated branch environment before loading the first dataset.
15 . The method of claim 12 , wherein the sequence of actions comprise real-life next actions of an ordered list including actions allocated to be performed by the real-life worker and bots.
16 . The method of claim 12 , wherein real-life observations of a result in the real-life branch environment are used to update a model of the master staffing engine.
17 . The method of claim 12 , wherein the branch simulation engine comprises a quantum hybrid simulation engine including classical computing devices to assist in preparing data coupled to a quantum computing device that assists in determining an effect of the action or task on a first simulated branch environment.Join the waitlist — get patent alerts
Track US12602634B1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.