Verifying an action proposed by a reinforcement learning model
Abstract
The present disclosure provides a computer-implemented method for determining whether to perform an action proposed by a model. The model is developed using a reinforcement learning process. The method comprises classifying at least one of a plurality of inputs to the model as being supportive or resistant to an action proposed by the model. The method further comprises comparing the classification of the at least one of the plurality of inputs to domain knowledge to determine whether or not the proposed action conflicts with the domain knowledge, and, in response to determining that the proposed action does not conflict with the domain knowledge, initiating the proposed action. In this context, the domain knowledge is indicative of a relationship between the proposed action and the at least one of the plurality of inputs.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method for determining whether to perform an action proposed by a model developed using a reinforcement learning process, the method comprising:
classifying at least one of a plurality of inputs to the model as being supportive or resistant to a proposed action by the model; comparing the classification of the at least one of the plurality of inputs to domain knowledge to determine whether or not the proposed action conflicts with the domain knowledge, wherein the domain knowledge is indicative of a relationship between the proposed action and the at least one of the plurality of inputs; and in response to determining that the proposed action does not conflict with the domain knowledge, initiating the proposed action.
2 .- 18 . (canceled)
19 . An apparatus for determining whether to perform an action proposed by a model developed using a reinforcement learning process, the apparatus comprising a processor and a machine-readable medium, wherein the machine-readable medium contains instructions executable by the processor such that the apparatus is operable to:
classify at least one of a plurality of inputs to the model as being supportive or resistant to a proposed action by the model; compare the classification of the at least one of the plurality of inputs to domain knowledge to determine whether or not the proposed action conflicts with the domain knowledge, wherein the domain knowledge is indicative of a relationship between the proposed action and the at least one of the plurality of inputs; and in response to determining that the proposed action does not conflict with the domain knowledge, initiate the proposed action.
20 . The apparatus of claim 19 , wherein the apparatus is operable to classify at least one of the plurality of inputs as being supportive or resistant using an explainable artificial intelligence, XAI, process.
21 . The apparatus of claim 19 , wherein the apparatus is further operable to:
determine a relative importance of one of the plurality of inputs to the proposed action compared to at least one other input in the plurality of inputs; and select the input for comparison with the domain knowledge based on its relative importance.
22 . The apparatus of claim 21 , wherein the apparatus is operable to determine the relative importance of the input using an XAI process.
23 . The apparatus of claim 19 , wherein the apparatus is operable to compare the classification of the at least one of the plurality of inputs to domain knowledge to determine whether or not the proposed action conflicts with the domain knowledge by:
determining that a conflict occurs in response to determining that one or more inputs have a classification that contradicts the domain knowledge.
24 . The apparatus of claim 19 , wherein the apparatus is further operable to:
map one or more other inputs to the model to one or more events for comparison with the domain knowledge; and compare the set of events with the domain knowledge to determine whether or not the proposed action conflicts with the domain knowledge.
25 . The apparatus of claim 24 , wherein the apparatus is operable to map one or more other inputs to one or more events using a mapping model developed using a machine learning process.
26 . The apparatus of claim 19 , wherein the apparatus is a node in a communication network.
27 . The apparatus of claim 26 , wherein the communication network comprises a radio access network and the plurality of inputs comprises one or more metrics of the radio access network and the proposed action comprises configuring an operational parameter of the radio access network.
28 . The apparatus of claim 26 , wherein the node comprises a base station or a core network node.
29 . The apparatus of claim 27 , wherein the apparatus is further operable to:
map other metrics of the radio access network that are input to the model to one or more events for comparison with the domain knowledge; and compare the set of events with the domain knowledge to determine whether or not the proposed action conflicts with the domain knowledge.
30 . The apparatus of claim 29 , wherein the other metrics comprise data representing received signal power at a base station in the radio access network over a period of time and data representing a plurality of performance metrics for a cell served by the base station over the time period and wherein the apparatus is operable to map one or more other metrics to one or more events using a mapping model developed using a machine learning process.
31 . The apparatus of claim 30 , wherein the machine learning process is a multi-task learning process.
32 . The apparatus of claim 19 , wherein the reinforcement learning process is a policy optimisation process or a q-learning process.Join the waitlist — get patent alerts
Track US2024249199A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.