Multi-modal model for dynamically responsive virtual characters
Abstract
The disclosed embodiments relate to a method for controlling a virtual character (or “avatar”) using a multi-modal model. The multi-modal model may process various input information relating to a user and process the input information using multiple internal models. The multi-modal model may combine the internal models to make believable and emotionally engaging responses by the virtual character. The link to a virtual character may be embedded on a web browser and the avatar may be dynamically generated based on a selection to interact with the virtual character by a user. A report may be generated for a client that provides insights as to characteristics of users interacting with a virtual character associated with the client.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . (canceled)
3 . (canceled)
4 . (canceled)
5 . (canceled)
6 . (canceled)
7 . (canceled)
8 . (canceled)
9 . (canceled)
10 . (canceled)
11 . (canceled)
12 . (canceled)
13 . (canceled)
14 . (canceled)
15 . (canceled)
16 . (canceled)
17 . (canceled)
18 . (canceled)
19 . (canceled)
20 . (canceled)
21 . A method of generating a virtual character comprising:
receiving multi-modal input information including environmental information representing a real-world environment, speech information, and facial expression information; implementing a first internal model of a plurality of models and a second internal model of the plurality of models to identify a first characteristic of the multi-modal input information by the first internal model and to identify a second characteristic of the multi-modal input information by the second internal model; determining whether the first characteristic is within a threshold similarity to the second characteristic; selecting a selected characteristic based on determining that the first characteristic is within the threshold similarity of the second characteristic; accessing a library of actions associated with the virtual character to select an action that matches the selected characteristic, the action including an animation and an associated audio to be performed by the virtual character; and causing display of the virtual character on a user device, the virtual character displayed on the user device as augmented in a display of the real-world environment, the virtual character performing the action and outputting the associated audio.
22 . The method of claim 21 , wherein selecting the selected characteristic is further based on a knowledge model of the virtual character that includes a persona of the virtual character.
23 . The method of claim 22 , further comprising:
receiving information relating to the virtual character from an external source; and updating the knowledge model based on the information relating to the virtual character.
24 . The method of claim 21 , wherein the knowledge model further includes information indicative of prior interactions between a user of the user device and the virtual character.
25 . The method of claim 21 , further comprising:
embedding a link in a web browser of the user device; receiving an indication from the user device that the link has been selected; and responsive to the link being selected, transmitting a stream of data to the user device, the stream of data including media files for performing the action and outputting the associated audio, wherein the virtual character is displayed in the web browser of the user device.
26 . The method of claim 25 , further comprising:
prior to embedding the link, transmitting a first batch of the stream of data at a first time, the first batch including information to initially generate the virtual character on the user device, wherein the virtual character is displayed in the web browser of the user device within one second of receiving the indication from the user device that the link has been selected.
27 . The method of claim 21 , wherein the first internal model is a speech recognition model configured to parse a speech sentiment from the speech information, wherein the second internal model is a facial feature recognition model configured to detect a facial feature sentiment based on the facial expression information, wherein the selected characteristic is a common sentiment among the speech sentiment and the facial feature sentiment, and wherein the action is determined based on the common sentiment.
28 . The method of claim 21 , wherein the multi-modal input information includes a plurality of information streams, the method further comprising:
augmenting each of the plurality of information streams with another of the plurality of information streams to produce a plurality of augmented information streams, wherein the first characteristic and the second characteristic are identified from the plurality of augmented information streams.
29 . The method of claim 21 , further comprising:
inspecting the environmental information to identify a portion of the real-world environment corresponding to a floor; and causing display of the virtual character on the user device as positioned on the floor in the display of the real-world environment.
30 . The method of claim 21 , wherein the plurality of models include a natural language understanding model configured to derive context and meaning from audio information, an awareness model configured to identify environmental information, and a social simulation model configured to identify data relating to a user and other virtual characters.
31 . A device comprising:
a plurality of sensors; a processor; and a memory storing instructions, execution of which by the processor causes the device to perform operations comprising:
receiving multi-modal input information via the plurality of sensors, the multi-modal input information including environmental information representing a real-world environment, speech information, and facial expression information;
implementing a first internal model of a plurality of models and a second internal model of the plurality of models to identify a first characteristic of the multi-modal input information by the first internal model and to identify a second characteristic of the multi-modal input information by the second internal model;
determining whether the first characteristic is within a threshold similarity to the second characteristic;
selecting a selected characteristic based on determining that the first identified characteristic is within the threshold similarity of the second identified characteristic;
accessing a library of actions associated with a virtual character to select an action that matches the selected characteristic, the action including an animation and an associated audio to be performed by the virtual character; and
displaying the virtual character on the device, the virtual character displayed on the device as augmented in a display of the real-world environment, the virtual character performing the action and outputting the associated audio.
32 . The device of claim 31 , wherein selecting the selected characteristic is further based on a knowledge model of the virtual character that includes a persona of the virtual character.
33 . The device of claim 31 , the operations further comprising: receiving a link embedded in a web browser executing on the device;
receiving a user-input indicating that the link has been selected; transmitting an indication from the device that the link has been selected; and receiving a stream of data at the device, the stream of data including media files for performing the action and outputting the associated audio, wherein the virtual character is displayed in the web browser executing on the device.
34 . The device of claim 31 , wherein the first internal model is a speech recognition model configured to parse a speech sentiment from the speech information, wherein the second internal model is a facial feature recognition model configured to detect a facial feature sentiment based on the facial expression information, wherein the selected characteristic is a common sentiment among the speech sentiment and the facial feature sentiment, and wherein the action is determined based on the common sentiment.
35 . The device of claim 31 , the operations further comprising:
inspecting the environmental information to identify a portion of the real-world environment corresponding to a floor; and displaying the virtual character on the device as positioned on the floor in the display of the real-world environment.
36 . The device of claim 31 , wherein the plurality of models include a natural language understanding model configured to derive context and meaning from audio information, an awareness model configured to identify environmental information, and a social simulation model configured to identify data relating to a user and other virtual characters.
37 . A non-transitory computer-readable storage medium storing instructions, execution of which by a processor of a computing system causes the computing system to:
receive, at an input layer from a plurality of sensors, a plurality of signals that each quantify a property of a physical environment being measured by the plurality of sensors; process the plurality of signals to produce a plurality of information streams corresponding to the plurality of signals; correct errors detected in a first information stream of the plurality of information streams by synthesizing the first information stream with a second information stream of the plurality of information streams; produce, at an output layer, a first augmented information stream based on the synthesis of the first information stream with the second information stream; and cause a virtual character to perform an action based on the first augmented information stream.
38 . The non-transitory computer-readable storage medium of claim 37 , wherein the plurality of sensors includes an image sensor, an audio sensor, and an olfactory sensor.
39 . The non-transitory computer-readable storage medium of claim 37 , wherein the first information stream includes audio-based speech recognition information, and wherein the second information stream includes computer vision-based lip reading information or olfactory information.
40 . The non-transitory computer-readable storage medium of claim 37 , wherein the errors detected in the first information stream are corrected by further synthesizing the first information stream with information about a fictional world of the virtual character.Join the waitlist — get patent alerts
Track US2023145369A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.