US2021125236A1PendingUtilityA1

System and method for synthesizing spoken dialogue using deep learning

Assignee: 10TALES INCPriority: Apr 7, 2003Filed: Nov 11, 2020Published: Apr 29, 2021
Est. expiryApr 7, 2023(expired)· nominal 20-yr term from priority
Inventors:David J. Russek
G06Q 10/40H04N 21/4126G06V 20/41H04N 21/6125H04N 21/4312G06F 16/285H04N 21/812G06Q 30/02G06F 16/58G06Q 30/00G06N 20/00H04N 21/4532H04N 21/4788H04N 21/25883G06Q 30/0271G06Q 30/0277G06F 16/9535H04N 21/44218G06Q 50/01G06K 9/00718G06Q 10/48G06Q 10/42G06F 16/9536
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Disclosed are a system, method and software to associate attributes with digital media assets using deep learning. Digital media contains specific assets, such as audio, that can be replaced with other assets. The system, method and software allow for personalizing digital media content based in part on neural network analysis in order to generate synthesized audio, such as speech, spoken prompts, and/or dialogue.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for personalizing media content using synthesized audio, comprising:
 generating, by a server, a script configured to output a digital media stream;   capturing, by a computing device communicatively coupled to the server, a user interaction with the digital media stream;   analyzing, by the server, the user interaction to determine a user affinity for the digital media stream;   selecting, by the server, an audio asset based on the user affinity; and   updating, by the server, the script to insert the audio asset into the digital media stream.   
     
     
         2 . The method of  claim 1 , wherein the audio asset is a speech-based audio asset. 
     
     
         3 . The method of  claim 1 , wherein the audio asset is a song. 
     
     
         4 . The method of  claim 1 , wherein the user affinity is selected from a group comprising at least one of a preference, an emotional value, a cognitive value, and a social value. 
     
     
         5 . The method of  claim 1 , wherein the user interaction is selected from a group comprising at least one of a viewing habit, a purchase, a selection, and an answer to a question. 
     
     
         6 . The method of  claim 1 , wherein the selecting utilizes an artificial intelligence method to identify the digital audio asset that matches the user affinity. 
     
     
         7 . The method of  claim 1 , wherein the selecting utilizes a neural network to identify the digital audio asset that matches the user affinity. 
     
     
         8 . The method of  claim 1 , wherein the user interaction is an input indicating a like or dislike by a user. 
     
     
         9 . A system for personalizing media content using synthesized audio, comprising:
 a content distributor configured to generate a script, the script configured to output a digital media stream;   a server coupled to the content distributor, the server configured to capture a user interaction with the digital media stream; and   a correlation algorithm communicatively coupled to the server, the correlation algorithm configured to correlate a user affinity with the user interaction,   wherein the content distributor is further configured to update the script to insert an audio asset into the digital media stream, wherein the audio asset is selected based on the user affinity.   
     
     
         10 . The system of  claim 9 , wherein the audio asset is a speech-based audio asset. 
     
     
         11 . The system of  claim 9 , wherein the audio asset is a spoken prompt. 
     
     
         12 . The system of  claim 9 , wherein the user affinity is selected from a group comprising at least one of a preference, an emotional value, a cognitive value, and a social value. 
     
     
         13 . The system of  claim 9 , wherein the user interaction is selected from a group comprising at least one of a viewing habit, a purchase, a selection, and an answer to a question. 
     
     
         14 . The system of  claim 9 , wherein the content distributor utilizes an artificial intelligence method to select the digital audio asset that matches the user affinity. 
     
     
         15 . The system of  claim 9 , wherein the content distributor utilizes a neural network to select the digital audio asset that matches the user affinity. 
     
     
         16 . A method for providing interactive synthesized audio content, comprising:
 generating, by a server, a script to output a digital audio stream in the form of a spoken question;   capturing, by a computing device communicatively coupled to the server, a user response to the spoken question;   selecting, by the server, an audio asset based on the user response;   updating, by the server, the script to insert the audio asset, thereby generating a second digital audio stream; and   outputting to the computing device, by the server, the second digital audio stream.   
     
     
         17 . The method of  claim 16 , wherein the user response is in the form of audio captured from a user. 
     
     
         18 . The method of  claim 16 , wherein the user response is the form of an emotion captured from a user. 
     
     
         19 . The method of  claim 16 , wherein the user response is an input indicating a like or dislike by a user. 
     
     
         20 . The method of  claim 16 , wherein the selecting utilizes an artificial intelligence method to identify the audio asset based on the user response.

Join the waitlist — get patent alerts

Track US2021125236A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.