Adversarial Cooperative Imitation Learning for Dynamic Treatment
Abstract
Methods and systems for responding to changing conditions include training a model, using a processor, using trajectories that resulted in a positive outcome and trajectories that resulted in a negative outcome. Training is performed using an adversarial discriminator to train the model to generate trajectories that are similar to historical trajectories that resulted in a positive outcome, and using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical trajectories that resulted in a negative outcome. A dynamic response regime is generated using the trained model and environment information. A response to changing environment conditions is performed in accordance with the dynamic response regime.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for responding to changing conditions, comprising:
training a model, using a processor, including trajectories that resulted in a positive outcome and trajectories that resulted in a negative outcome, by using an adversarial discriminator to train the model to generate trajectories that are similar to historical trajectories that resulted in a positive outcome, and by using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical trajectories that resulted in a negative outcome, and including iteratively training the adversarial discriminator, the cooperative discriminator, and the dynamic response regime using a three-party optimization until a predetermined number of iterations has been reached; generating a dynamic response regime using the trained model and environment information; and responding to changing environment conditions in accordance with the dynamic response regime.
2 . The method of claim 1 , wherein the historical trajectories include patient treatment trajectories.
3 . The method of claim 2 , wherein the positive outcomes are positive patient health outcomes, and the negative outcomes are negative patient health outcomes.
4 . The method of claim 2 , wherein the environment information and the environment conditions reflect information about a patient being treated.
5 . The method of claim 1 , wherein the adversarial discriminator, the cooperative discriminator, and the dynamic response regime are implemented as multiple-layer perceptrons.
6 . The method of claim 1 , wherein training the model comprises training an environment model that encodes environment information as a vector in a latent space.
7 . The method of claim 1 , wherein the model is implemented as a variational auto-encoder network.
8 . The method of claim 1 , wherein responding to changing environment conditions comprises automatically performing a responsive action to correct a negative condition.
9 . A system for responding to changing conditions, comprising:
a machine learning model, configured to generate a dynamic response regime for using environment information; a model trainer, configured to train the machine learning model, including trajectories that resulted in a positive outcome and trajectories that resulted in a negative outcome, by using an adversarial discriminator to train the machine learning model to generate trajectories that are similar to historical trajectories that resulted in a positive outcome, and by using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical trajectories that resulted in a negative outcome, and to iteratively train the adversarial discriminator, the cooperative discriminator, and the dynamic response regime using a three-party optimization until a predetermined number of iterations has been reached; and a response interface, configured to trigger a response to changing environment conditions in accordance with the dynamic response regime.
10 . The system of claim 9 , wherein the historical trajectories that resulted in a positive outcome and the historical trajectories that resulted in a negative outcome include patient treatment trajectories.
11 . The system of claim 10 , wherein the positive outcomes are positive patient health outcomes, and the negative outcomes are negative patient health outcomes.
12 . The system of claim 9 , wherein the environment information and the environment conditions reflect information about a patient being treated.
13 . The system of claim 9 , wherein the model trainer is further configured to iteratively train the adversarial discriminator, the cooperative discriminator, and the dynamic response regime using a three-party optimization.
14 . The system of claim 9 , wherein the adversarial discriminator, the cooperative discriminator, and the dynamic response regime are implemented as multiple-layer perceptrons in the machine learning model.
15 . The system of claim 9 , wherein the model trainer is further configured to train an environment model that encodes the environment information as a vector in a latent space.
16 . The system of claim 15 , wherein the environment model is implemented as a variational auto-encoder network in the machine learning model.
17 . The system of claim 9 , wherein the response interface is further configured to automatically perform a responsive action to correct a negative condition.Join the waitlist — get patent alerts
Track US2024005163A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.