Intelligent virtual assistant training through phased observational learning tasks
Abstract
Disclosed embodiments pertain to training an intelligent virtual assistant through phased observational learning tasks. A pre-trained language model can be updated offline to produce a second language model with self-supervised learning based on transcripts of historical interactions between one or more customers, one or more customer service agents, and one or more data stores. The second language model can be evaluated and determined to satisfy a predetermined performance threshold. Subsequently, the second language model can be updated online to produce a third language model with reinforcement learning based on received customer input and similarity between a response provided by a customer service agent and a predicted response generated by the second language model. The third language model can then be deployed with an intelligent virtual assistant to respond to received user input.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An intelligent virtual assistant system, comprising:
a processor coupled to a memory that stores instructions that, when executed by the processor, cause the processor to:
update a pre-trained language model offline to produce a second language model with self-supervised learning based on transcripts of historical interactions between one or more customers, one or more customer service agents, and one or more data stores;
determine that the second language model satisfies a predetermined performance threshold;
update the second language model online to produce a third language model with reinforcement learning based on received customer input and a similarity between a response provided by a customer service agent and a predicted response generated by the second language model; and
deploy the third language model to respond to received user input.
2 . The intelligent virtual assistant system of claim 1 , wherein the instructions further cause the processor to:
generate a generic entity placeholder for at least one aspect in the transcripts of historical interactions associated with a subset of customer service agents; and update the pre-trained language model using the transcripts with the self-supervised learning to generate a response template with the generic entity placeholder.
3 . The intelligent virtual assistant system of claim 2 , wherein the instructions further cause the processor to:
identify a data retrieval action of a customer service agent within a conversation in one or more transcripts; add reference information associated with the data retrieval action in the one or more transcripts; and update the pre-trained language model using the one or more transcripts with the self-supervised learning to generate the response template with the reference information.
4 . The intelligent virtual assistant system of claim 3 , wherein the instructions further cause the processor to update the pre-trained language model to generate a complete response using the reference information to populate one or more entity placeholders in the response template.
5 . The intelligent virtual assistant system of claim 4 , wherein the instructions further cause the processor to:
detect one or more errors that fail to satisfy a second predetermined performance threshold based on historical customer input and comparison of the complete response with a response to a customer service agent; and initiate additional offline updating of the pre-trained language model.
6 . The intelligent virtual assistant system of claim 1 , wherein the instructions further cause the processor to compute a performance score of the third language model based on one or more feedback signals associated with user interaction.
7 . The intelligent virtual assistant system of claim 6 , wherein the instructions further cause the processor to:
deploy the third language model to respond to a subset of the additional received user input; and adjust subset size based on a performance score.
8 . The intelligent virtual assistant system of claim 6 , wherein the instructions further cause the processor to:
determine that the performance score fails to satisfy a predetermined minimum threshold; and initiate further updating of the third language model.
9 . The intelligent virtual assistant system of claim 1 , wherein the instructions further cause the processor to update the second language model online based on customer service agent feedback.
10 . A method of generating an intelligent virtual assistant, comprising:
updating a pre-trained language model offline to produce a second language model with self-supervised learning based on transcripts of historical interactions between one or more customers, one or more customer service agents, and one or more data stores; determining that the second language model satisfies a predetermined performance threshold; updating the second language model online to produce a third language model with reinforcement learning based on received customer input and a similarity between a response provided by a customer service agent and a predicted response generated by the second language model; and deploying the third language model to respond to received user input.
11 . The method of claim 10 , further comprising:
generating a generic entity placeholder for at least one aspect in the transcripts of historical interactions associated with a subset of customer service agents; and updating the pre-trained language model using the transcripts with the self-supervised learning to generate a response template with the generic entity placeholder.
12 . The method of claim 11 , further comprising:
identifying a data retrieval action of a customer service agent within a conversation in one or more transcripts; adding reference information associated with the data retrieval action in the one or more transcripts; and updating the pre-trained language model using the one or more transcripts with the self-supervised learning to generate the response template with the reference information.
13 . The method of claim 12 , further comprising updating the pre-trained language model to generate a complete response using the reference information to populate one or more entity placeholders in the response template.
14 . The method of claim 13 , further comprising:
detecting one or more errors that fail to satisfy a second predetermined performance threshold based on historical customer input and comparison of the complete response with a response to a customer service agent; and initiating additional offline updating of the pre-trained language model.
15 . The method of claim 10 , further comprising computing a performance score of the third language model based on one or more feedback signals associated with user interaction.
16 . The method of claim 15 , further comprising:
deploying the third language model to respond to a subset of the additional received user input; and adjusting subset size based on the performance score.
17 . The method of claim 15 , further comprising:
determining that the performance score fails to satisfy a predetermined minimum threshold; and initiating further updating of the third language model.
18 . An intelligent virtual assistant method, comprising:
receiving a user input; invoking a language model to infer a response to the user input, the language model having been trained by:
updating a pre-trained language model offline to produce a second language model with self-supervised learning based on transcripts of historical interactions between one or more customers, one or more customer service agents, and one or more data stores;
determining that the second language model satisfies a predetermined performance threshold;
updating the second language model online to produce a third language model with reinforcement learning based on received customer input and similarity between a response provided by a customer service agent and a predicted response generated by the second language model; and
outputting the response to the user input.
19 . The method of claim 18 , further comprising computing a performance score of the language model based on one or more feedback signals associated with user interaction.
20 . The method of claim 19 , further comprising:
determining that the performance score fails to satisfy a predetermined minimum threshold; and initiating further training of the language model.Join the waitlist — get patent alerts
Track US2024355318A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.