Adversarial Cooperative Imitation Learning for Dynamic Treatment
Abstract
Methods and systems for responding to changing conditions include training a model, using a processor, using trajectories that resulted in a positive outcome and trajectories that resulted in a negative outcome. Training is performed using an adversarial discriminator to train the model to generate trajectories that are similar to historical trajectories that resulted in a positive outcome, and using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical trajectories that resulted in a negative outcome. A dynamic response regime is generated using the trained model and environment information. A response to changing environment conditions is performed in accordance with the dynamic response regime.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for responding to changing conditions, comprising:
training a model on historical treatment trajectories, using a processor, including trajectories that resulted in a positive health outcome and trajectories that resulted in a negative health outcome, by using an adversarial discriminator to train the model to generate trajectories that are similar to historical treatment trajectories that resulted in a positive health outcome, and by using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical treatment trajectories that resulted in a negative health outcome, and including iteratively training the adversarial discriminator, the cooperative discriminator, and a dynamic treatment regime using a three-party optimization; generating the dynamic treatment regime for a patient using the trained model and patient information; and treating the patient in accordance with the dynamic treatment regime, in a manner that is responsive to changing patient conditions, by triggering one or more medical devices to administer a treatment to the patient.
2 . The method of claim 1 , wherein the adversarial discriminator, the cooperative discriminator, and the dynamic treatment regime are implemented as multiple-layer perceptrons.
3 . The method of claim 1 , wherein training the model comprises training an environment model that encodes the patient information as a vector in a latent space.
4 . The method of claim 3 , wherein the environment model is implemented as a variational auto-encoder network.
5 . The method of claim 1 , wherein responding to changing patient conditions comprises automatically performing a responsive action to correct a negative condition.
6 . A system for responding to changing conditions, comprising:
a machine learning model, configured to generate a dynamic treatment regime for using patient information; a model trainer, configured to train the machine learning model, including trajectories that resulted in a positive health outcome and trajectories that resulted in a negative health outcome, by using an adversarial discriminator to train the machine learning model to generate trajectories that are similar to historical treatment trajectories that resulted in a positive health outcome, and by using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical treatment trajectories that resulted in a negative health outcome, and to iteratively train the adversarial discriminator, the cooperative discriminator, and a dynamic treatment regime using a three-party optimization; and a response interface, configured to trigger one or more medical devices to treat the patient in accordance with the dynamic treatment regime in a manner that is responsive to changing patient conditions.
7 . The system of claim 6 , wherein the adversarial discriminator, the cooperative discriminator, and the dynamic treatment regime are implemented as multiple-layer perceptrons in the machine learning model.
8 . The system of claim 6 , wherein the model trainer is further configured to train an environment model that encodes the patient information as a vector in a latent space.
9 . The system of claim 8 , wherein the environment model is implemented as a variational auto-encoder network in the machine learning model.
10 . The system of claim 6 , wherein the response interface is further configured to automatically perform a responsive action to correct a negative condition.Join the waitlist — get patent alerts
Track US2023376773A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.