US2020202859A1PendingUtilityA1

Generating interactive audio-visual representations of individuals

Assignee: AARABI PEGAHPriority: Jul 24, 2018Filed: Mar 4, 2020Published: Jun 25, 2020
Est. expiryJul 24, 2038(~12 yrs left)· nominal 20-yr term from priority
Inventors:Pegah Aarabi
G06N 3/047G06N 3/045G06F 16/48G06N 3/0475G06N 3/0455G06N 3/094G06N 5/04G06N 3/088G06N 3/006G06N 20/00G10L 15/1815G10L 15/26G10L 15/22G10L 2015/088G10L 2015/223
43
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system for generating an audio-visual representation of an individual is provided. The system includes an audio-visual representation generator to obtain audio-visual data of an individual communicating responses to prompts. The generator includes a recording analyzer and recording processor to segment the audio-visual data into responsive audio-video segments, or includes a machine learning model to generate artificial audio-visual responses, which simulate the individual communicating a response to the input prompt.

Claims

exact text as granted — not AI-modified
1 . A system for generating an interactive audio-visual representation of an individual, the system comprising:
 a memory storage unit to store a plurality of audio-video recordings of an individual communicating responses to prompts;   a recording analyzer to segment the plurality of audio-video recordings into a plurality of audio-video segments according to topics referenced in the responses or the prompts;   a communication interface to receive a linguistic input;   a recording processor to analyze the linguistic input and generate one or more responsive audio-video segments from the plurality of audio-video segments to be provided in response to the linguistic input; and   an audio-visual media generator to generate a playback of the one or more responsive audio-video segments as an audio-visual representation of the individual responding to the linguistic input.   
     
     
         2 . The system of  claim 1 , wherein the one or more responsive audio-video segments comprises a plurality of responsive audio-video segments, the recording processor comprises a video segment resequencer, and wherein the recording processor generates the plurality of responsive audio-video segments at least in part by the video segment resequencer selecting and resequencing a plurality of selected responsive audio-video segments from the plurality of audio-video segments. 
     
     
         3 . The system of  claim 1 , wherein:
 the recording analyzer comprises a video segment labeler to generate keyword labels for the audio-video segments indicating topics covered in the audio-video segments;   the communication interface comprises an input labeler to generate keyword labels for the linguistic input; and   the recording processor generates the one or more responsive audio-video segments by matching keyword labels of the audio-video segments with keyword labels of the linguistic input.   
     
     
         4 . The system of  claim 3 , wherein the linguistic input comprises an auditory input, and wherein the communication interface comprises a text transcriber to transcribe the auditory input into a text input, and wherein the input labeler generates keyword labels for the linguistic input by generating keyword labels for the input text. 
     
     
         5 . The system of  claim 1 , wherein the recording processor comprises a natural language processor to determine a meaning of the linguistic input. 
     
     
         6 . The system of  claim 1 , wherein the playback comprises an audio-video compilation of the one or more responsive audio-video segments. 
     
     
         7 . The system of  claim 1 , wherein the plurality of audio-video recordings comprises a plurality of video recording threads, each respective video recording thread captured by a different respective video recording device, and wherein the playback comprises an augmented reality representation generated with the one or more responsive audio-video segments. 
     
     
         8 . The system of  claim 1 , wherein the plurality of audio-video recordings comprises a plurality of video recording threads, each respective video recording thread captured by a different respective recording device, and wherein the playback comprises a virtual reality representation generated with the one or more responsive audio-video segments. 
     
     
         9 . The system of  claim 1 , wherein a prompt of the prompts includes a question to elucidate an aspect of personality of the individual. 
     
     
         10 . (canceled) 
     
     
         11 . (canceled) 
     
     
         12 . A system for generating an interactive audio-visual representation of an individual, the system comprising:
 a memory storage unit to store genuine audio-visual responses to prompts, each genuine audio-visual response comprising a segment of an audio-video recording of the individual communicating a response to a prompt;   a communication interface to receive a linguistic input;   a machine learning model to generate an artificial audio-visual response to the linguistic input to simulate how the individual may respond to the linguistic input, the machine learning model trained with the genuine audio-visual responses to generate artificial audio-visual responses to simulate how the individual may respond to linguistic inputs; and   an audio-visual media generator to generate media as an audio-visual representation of the individual based on the artificial audio-visual response.   
     
     
         13 . (canceled) 
     
     
         14 . (canceled) 
     
     
         15 . (canceled) 
     
     
         16 . A system for generating an interactive audio-visual representation of an individual, the system comprising:
 one or more recording devices to capture audio-visual data of an individual communicating responses to prompts;   an audio-visual representation generator to obtain the audio-visual data, analyze the audio-visual data, receive an input prompt, and generate an audio-visual response to the input prompt based on analysis of the audio-visual data to simulate the individual communicating a response to the input prompt; and   a media device to output the audio-visual response.   
     
     
         17 . A system of  claim 16 , wherein the audio-visual representation generator comprises:
 a recording analyzer to segment the audio-visual data into a plurality of audio-video segments according to topics referenced in the responses or the prompts; and   a recording processor to analyze the input prompt and generate one or more responsive audio-video segments from the plurality of audio-video segments as the audio-visual response.   
     
     
         18 . The system of  claim 16 , wherein the audio-visual representation generator comprises:
 a machine learning model to generate an artificial audio-visual response to the input prompt as the audio-visual response, the machine learning model including a generative adversarial network, the generative adversarial network adversarially trained with the audio-visual data to generate artificial audio-visual responses to simulate how the individual may respond to input prompts.   
     
     
         19 . The system of  claim 16 , wherein the audio-visual data comprises a plurality of visual recording threads, each respective visual recording thread captured by a different respective recording device, and wherein the media device outputs the audio-visual response in an augmented reality representation. 
     
     
         20 . The system of  claim 16 , wherein the audio-visual data comprises a plurality of visual recording threads, each respective visual recording thread captured by a different respective recording device, and wherein the media device outputs the audio-visual response in a virtual reality representation.

Join the waitlist — get patent alerts

Track US2020202859A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.