Apparatus and methods for online training of robots
Abstract
Robotic devices may be trained by a user guiding the robot along a target trajectory using a correction signal. A robotic device may comprise an adaptive controller configured to generate control commands based on one or more of the trainer input, sensory input, and/or performance measure. Training may comprise a plurality of trials. During an initial portion of a trial, the trainer may observe robot's operation and refrain from providing the training input to the robot. Upon observing a discrepancy between the target behavior and the actual behavior during the initial trial portion, the trainer may provide a teaching input (e.g., a correction signal) configured to affect robot's trajectory during subsequent trials. Upon completing a sufficient number of trials, the robot may be capable of navigating the trajectory in absence of the training input.
Claims
exact text as granted — not AI-modified1 .- 22 . (canceled)
23 . A robot, comprising:
a controllable actuator; a sensor module configured to provide information related to an environment surrounding the robot; and an adaptive controller configured to produce a control instruction for the controllable actuator in accordance with the information provided by the sensor module, the control instruction being configured to cause the robot to execute a target task autonomously; wherein:
the target task is learned by the robot through a demonstration by a trainer in one or more trials;
the adaptive controller is further configured to receive a corrective signal from the trainer in an operating mode in which substantially no learning by the robot occurs; and
the corrective signal is configured to alter an observed behavior of the robot while the robot executes the target task autonomously.
24 . The robot of claim 23 , wherein the corrective signal originates from a remote device disposed remotely from the robot.
25 . The robot of claim 23 , wherein the demonstration comprises a supervised learning process based on a sensory context and a teaching signal provided by the trainer.
26 . The robot of claim 25 , wherein the target task is demonstrated over a plurality of trials.
27 . The robot of claim 23 , wherein the target task comprises travelling along a target trajectory.
28 . The robot of claim 23 , wherein the sensor module comprises at least one of a camera, lidar, and sonar.
29 . The robot of claim 23 , wherein the adaptive controller is configured to cause the robot to avoid an obstacle detected by the sensor module.
30 . The robot of claim 29 , wherein the corrective signal includes a command to the adaptive controller to avoid an obstacle.
31 . A method of operating a robot, comprising:
receiving a sensory context from a sensor of the robot; at a first time instance, executing a first action with the robot in accordance with a first sensory context; at a second time instance subsequent to the first time instance, determining with an adaptive controller whether to execute the first action based on: (i) a second sensory context received from the sensor, and (ii) a user input received from a user interface; and executing the first action with the robot in accordance with the determination of the adaptive controller; wherein:
increasing or decreasing a probability of execution of the first action is based on the user input at the second time instance, the user input having an effectiveness value predetermined by the adaptive controller.
32 . The method of claim 31 , wherein the first sensory context and the second sensory context are at substantially similar location.
33 . The method of claim 31 , further comprising at the second time instance, detecting an obstacle in the second sensory context.
34 . The method of claim 31 , wherein the first action is following a first trajectory.
35 . The method of claim 31 , wherein the user input includes instructions to follow a second trajectory.
36 . The method of claim 31 , further comprising recognizing objects in the first sensory context and second sensory context.
37 . The method of claim 31 , further comprising generating commands from the adaptive controller for operation of the robot, the commands based at least in part on training during one or more trials.
38 . The method of claim 31 , further comprising performing an online learning process, wherein the execution of the first action is at least a portion of the online learning process.
39 . An adaptive controller apparatus, comprising:
one or more processors configured to execute computer program instructions that, when executed, cause a robot to: at a first time instance, following a first trajectory in accordance with a first sensory context; at a second time instance subsequent to the first time instance, determine with an adaptive controller whether to execute the first action based on: (i) a second sensory context, and (ii) a user input received from a user interface; and execute a second trajectory with the robot in accordance with the determination of the adaptive controller; wherein:
the first trajectory is learned by the robot through a demonstration by a trainer in one or more trials;
the user input is a corrective signal configured to alter the observed behavior of the adaptive controller in an operating mode in which the robot operates autonomously and no learning occurs; and
the second trajectory comprises an altered behavior in accordance in accordance to the corrective signal and a resuming of the first trajectory.
40 . The adaptive controller apparatus of claim 39 , wherein the first sensory context and the second sensory context are at substantially similar location.
41 . The adaptive controller apparatus of claim 39 , wherein the one or more processors are further configured to perform an online learning process, the first time instance occurring during at least a portion of the online learning process.
42 . The adaptive controller apparatus of claim 39 , wherein the second sensory context includes a detected obstacle.Join the waitlist — get patent alerts
Track US2017095923A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.