US2020202859A1PendingUtilityA1
Generating interactive audio-visual representations of individuals
Est. expiryJul 24, 2038(~12 yrs left)· nominal 20-yr term from priority
Inventors:Pegah Aarabi
G06N 3/047G06N 3/045G06F 16/48G06N 3/0475G06N 3/0455G06N 3/094G06N 5/04G06N 3/088G06N 3/006G06N 20/00G10L 15/1815G10L 15/26G10L 15/22G10L 2015/088G10L 2015/223
43
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A system for generating an audio-visual representation of an individual is provided. The system includes an audio-visual representation generator to obtain audio-visual data of an individual communicating responses to prompts. The generator includes a recording analyzer and recording processor to segment the audio-visual data into responsive audio-video segments, or includes a machine learning model to generate artificial audio-visual responses, which simulate the individual communicating a response to the input prompt.
Claims
exact text as granted — not AI-modified1 . A system for generating an interactive audio-visual representation of an individual, the system comprising:
a memory storage unit to store a plurality of audio-video recordings of an individual communicating responses to prompts; a recording analyzer to segment the plurality of audio-video recordings into a plurality of audio-video segments according to topics referenced in the responses or the prompts; a communication interface to receive a linguistic input; a recording processor to analyze the linguistic input and generate one or more responsive audio-video segments from the plurality of audio-video segments to be provided in response to the linguistic input; and an audio-visual media generator to generate a playback of the one or more responsive audio-video segments as an audio-visual representation of the individual responding to the linguistic input.
2 . The system of claim 1 , wherein the one or more responsive audio-video segments comprises a plurality of responsive audio-video segments, the recording processor comprises a video segment resequencer, and wherein the recording processor generates the plurality of responsive audio-video segments at least in part by the video segment resequencer selecting and resequencing a plurality of selected responsive audio-video segments from the plurality of audio-video segments.
3 . The system of claim 1 , wherein:
the recording analyzer comprises a video segment labeler to generate keyword labels for the audio-video segments indicating topics covered in the audio-video segments; the communication interface comprises an input labeler to generate keyword labels for the linguistic input; and the recording processor generates the one or more responsive audio-video segments by matching keyword labels of the audio-video segments with keyword labels of the linguistic input.
4 . The system of claim 3 , wherein the linguistic input comprises an auditory input, and wherein the communication interface comprises a text transcriber to transcribe the auditory input into a text input, and wherein the input labeler generates keyword labels for the linguistic input by generating keyword labels for the input text.
5 . The system of claim 1 , wherein the recording processor comprises a natural language processor to determine a meaning of the linguistic input.
6 . The system of claim 1 , wherein the playback comprises an audio-video compilation of the one or more responsive audio-video segments.
7 . The system of claim 1 , wherein the plurality of audio-video recordings comprises a plurality of video recording threads, each respective video recording thread captured by a different respective video recording device, and wherein the playback comprises an augmented reality representation generated with the one or more responsive audio-video segments.
8 . The system of claim 1 , wherein the plurality of audio-video recordings comprises a plurality of video recording threads, each respective video recording thread captured by a different respective recording device, and wherein the playback comprises a virtual reality representation generated with the one or more responsive audio-video segments.
9 . The system of claim 1 , wherein a prompt of the prompts includes a question to elucidate an aspect of personality of the individual.
10 . (canceled)
11 . (canceled)
12 . A system for generating an interactive audio-visual representation of an individual, the system comprising:
a memory storage unit to store genuine audio-visual responses to prompts, each genuine audio-visual response comprising a segment of an audio-video recording of the individual communicating a response to a prompt; a communication interface to receive a linguistic input; a machine learning model to generate an artificial audio-visual response to the linguistic input to simulate how the individual may respond to the linguistic input, the machine learning model trained with the genuine audio-visual responses to generate artificial audio-visual responses to simulate how the individual may respond to linguistic inputs; and an audio-visual media generator to generate media as an audio-visual representation of the individual based on the artificial audio-visual response.
13 . (canceled)
14 . (canceled)
15 . (canceled)
16 . A system for generating an interactive audio-visual representation of an individual, the system comprising:
one or more recording devices to capture audio-visual data of an individual communicating responses to prompts; an audio-visual representation generator to obtain the audio-visual data, analyze the audio-visual data, receive an input prompt, and generate an audio-visual response to the input prompt based on analysis of the audio-visual data to simulate the individual communicating a response to the input prompt; and a media device to output the audio-visual response.
17 . A system of claim 16 , wherein the audio-visual representation generator comprises:
a recording analyzer to segment the audio-visual data into a plurality of audio-video segments according to topics referenced in the responses or the prompts; and a recording processor to analyze the input prompt and generate one or more responsive audio-video segments from the plurality of audio-video segments as the audio-visual response.
18 . The system of claim 16 , wherein the audio-visual representation generator comprises:
a machine learning model to generate an artificial audio-visual response to the input prompt as the audio-visual response, the machine learning model including a generative adversarial network, the generative adversarial network adversarially trained with the audio-visual data to generate artificial audio-visual responses to simulate how the individual may respond to input prompts.
19 . The system of claim 16 , wherein the audio-visual data comprises a plurality of visual recording threads, each respective visual recording thread captured by a different respective recording device, and wherein the media device outputs the audio-visual response in an augmented reality representation.
20 . The system of claim 16 , wherein the audio-visual data comprises a plurality of visual recording threads, each respective visual recording thread captured by a different respective recording device, and wherein the media device outputs the audio-visual response in a virtual reality representation.Join the waitlist — get patent alerts
Track US2020202859A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.