US2024005163A1PendingUtilityA1

Adversarial Cooperative Imitation Learning for Dynamic Treatment

Assignee: NEC LAB AMERICA INCPriority: Aug 29, 2019Filed: Jul 31, 2023Published: Jan 4, 2024
Est. expiryAug 29, 2039(~13.1 yrs left)· nominal 20-yr term from priority
G06N 3/0499G06N 3/0455G06N 3/094G06N 3/092G06N 3/0475G06N 3/084G16H 50/30G06N 3/045G06N 5/046G06N 20/20G16H 50/20G16H 20/00G06N 3/006G06N 3/088G06N 3/047G06N 7/01
78
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and systems for responding to changing conditions include training a model, using a processor, using trajectories that resulted in a positive outcome and trajectories that resulted in a negative outcome. Training is performed using an adversarial discriminator to train the model to generate trajectories that are similar to historical trajectories that resulted in a positive outcome, and using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical trajectories that resulted in a negative outcome. A dynamic response regime is generated using the trained model and environment information. A response to changing environment conditions is performed in accordance with the dynamic response regime.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for responding to changing conditions, comprising:
 training a model, using a processor, including trajectories that resulted in a positive outcome and trajectories that resulted in a negative outcome, by using an adversarial discriminator to train the model to generate trajectories that are similar to historical trajectories that resulted in a positive outcome, and by using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical trajectories that resulted in a negative outcome, and including iteratively training the adversarial discriminator, the cooperative discriminator, and the dynamic response regime using a three-party optimization until a predetermined number of iterations has been reached;   generating a dynamic response regime using the trained model and environment information; and   responding to changing environment conditions in accordance with the dynamic response regime.   
     
     
         2 . The method of  claim 1 , wherein the historical trajectories include patient treatment trajectories. 
     
     
         3 . The method of  claim 2 , wherein the positive outcomes are positive patient health outcomes, and the negative outcomes are negative patient health outcomes. 
     
     
         4 . The method of  claim 2 , wherein the environment information and the environment conditions reflect information about a patient being treated. 
     
     
         5 . The method of  claim 1 , wherein the adversarial discriminator, the cooperative discriminator, and the dynamic response regime are implemented as multiple-layer perceptrons. 
     
     
         6 . The method of  claim 1 , wherein training the model comprises training an environment model that encodes environment information as a vector in a latent space. 
     
     
         7 . The method of  claim 1 , wherein the model is implemented as a variational auto-encoder network. 
     
     
         8 . The method of  claim 1 , wherein responding to changing environment conditions comprises automatically performing a responsive action to correct a negative condition. 
     
     
         9 . A system for responding to changing conditions, comprising:
 a machine learning model, configured to generate a dynamic response regime for using environment information;   a model trainer, configured to train the machine learning model, including trajectories that resulted in a positive outcome and trajectories that resulted in a negative outcome, by using an adversarial discriminator to train the machine learning model to generate trajectories that are similar to historical trajectories that resulted in a positive outcome, and by using a cooperative discriminator to train the model to generate trajectories that are dissimilar to historical trajectories that resulted in a negative outcome, and to iteratively train the adversarial discriminator, the cooperative discriminator, and the dynamic response regime using a three-party optimization until a predetermined number of iterations has been reached; and   a response interface, configured to trigger a response to changing environment conditions in accordance with the dynamic response regime.   
     
     
         10 . The system of  claim 9 , wherein the historical trajectories that resulted in a positive outcome and the historical trajectories that resulted in a negative outcome include patient treatment trajectories. 
     
     
         11 . The system of  claim 10 , wherein the positive outcomes are positive patient health outcomes, and the negative outcomes are negative patient health outcomes. 
     
     
         12 . The system of  claim 9 , wherein the environment information and the environment conditions reflect information about a patient being treated. 
     
     
         13 . The system of  claim 9 , wherein the model trainer is further configured to iteratively train the adversarial discriminator, the cooperative discriminator, and the dynamic response regime using a three-party optimization. 
     
     
         14 . The system of  claim 9 , wherein the adversarial discriminator, the cooperative discriminator, and the dynamic response regime are implemented as multiple-layer perceptrons in the machine learning model. 
     
     
         15 . The system of  claim 9 , wherein the model trainer is further configured to train an environment model that encodes the environment information as a vector in a latent space. 
     
     
         16 . The system of  claim 15 , wherein the environment model is implemented as a variational auto-encoder network in the machine learning model. 
     
     
         17 . The system of  claim 9 , wherein the response interface is further configured to automatically perform a responsive action to correct a negative condition.

Join the waitlist — get patent alerts

Track US2024005163A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.