US2019073993A1PendingUtilityA1

Artificially generated speech for a communication session

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Feb 2, 2017Filed: Oct 31, 2018Published: Mar 7, 2019
Est. expiryFeb 2, 2037(~10.5 yrs left)· nominal 20-yr term from priority
H04S 2420/01H04S 7/30H04L 65/1069G10L 13/033H04M 3/2236G10L 13/047H04L 43/08H04M 2201/40G10L 13/08H04M 7/0084H04L 65/752H04L 65/762H04R 2227/003G10L 19/0018H04R 27/00G10L 13/04H04L 65/80H04M 7/006
46
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A device is disclosed, which includes a processor and a memory in communication with the processor. The memory includes executable instructions that, when executed by the processor, cause the processor to control the device to perform functions of capturing a speech by a user; generating audio data representing the captured speech by a user; generating, based on the audio data, text data representing at least a portion of the captured speech; and transmitting, via a communication channel, the audio data and text data to the remote device. The device thus can provide the text data representing the captured speech when a quality of the audio signal received by the remote device is below a predetermined level.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A device comprising:
 a processor; and   a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor, cause the processor to control the device to perform functions of:
 capturing a speech by a user; 
 generating audio data representing the captured speech by a user; 
 generating, based on the audio data, text data representing at least a portion of the captured speech; and 
 transmitting, via a communication channel, the audio data and text data to a remote device. 
   
     
     
         2 . The device of  claim 1 , further comprising a microphone configured to capture the speech by the user. 
     
     
         3 . The device of  claim 1 , further comprising a camera configured to capture an image of the user, wherein the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform a function of generating video data presenting the captured image of the user. 
     
     
         4 . The device of  claim 3 , wherein the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform a function of encoding the audio data, video data and text data for transmission to the remote device via the communication network. 
     
     
         5 . The device of  claim 1 , wherein the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform a function of receiving, via the communication channel, a feedback signal from the remote device, wherein the transmission of the text data to the remote device is initiated when the feedback signal indicates an occurrence of a predetermined condition at the remote device. 
     
     
         6 . The device of  claim 5 , wherein the predetermined condition includes a condition that a quality of the audio data received by the remote device is below a predetermined level. 
     
     
         7 . The device of  claim 1 , wherein to transmit the audio data and text data, the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform functions of:
 transmitting the audio data to the remote device via a first communication modality; and   transmitting the text data to the remote device via a second communication modality having a higher robustness than the first communication modality.   
     
     
         8 . The device of  claim 1 , wherein to transmit the audio data and text data, the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform functions of:
 transmitting the audio data to the remote device in a first quality of service level; and   transmitting the text data to the remote device in a second quality of service level that is higher than the first quality of service level.   
     
     
         9 . A method comprising:
 capturing a speech by a user;   generating audio data representing the captured speech by a user;   generating, based on the audio data, text data representing at least a portion of the captured speech; and   transmitting, via a communication channel, the audio data and text data to a remote device.   
     
     
         10 . The method of  claim 9 , further comprising:
 capturing an image of the user; and   generating video data presenting the captured image of the user.   
     
     
         11 . The method of  claim 10 , further comprising encoding the audio data, video data and text data for transmission to the remote device via the communication network. 
     
     
         12 . The method of  claim 9 , further comprising receiving, via the communication channel, a feedback signal from the remote device, wherein the transmission of the text data to the remote device is initiated when the feedback signal indicates an occurrence of a predetermined condition at the remote device. 
     
     
         13 . The method of  claim 12 , wherein the predetermined condition includes a condition that a quality of the audio data received by the remote device is below a predetermined level. 
     
     
         14 . The method of  claim 9 , wherein the transmitting the audio data and text data comprises:
 transmitting the audio data to the remote device via a first communication modality; and   transmitting the text data to the remote device via a second communication modality having a higher robustness than the first communication modality.   
     
     
         15 . The method of  claim 9 , wherein the transmitting the audio data and text data comprises:
 transmitting the audio data to the remote device in a first quality of service level; and   transmitting the text data to the remote device in a second quality of service level that is higher than the first quality of service level.   
     
     
         16 . A device comprising:
 means for capturing a speech by a user;   means for generating audio data representing the captured speech by a user;   means for generating, based on the audio data, text data representing at least a portion of the captured speech; and   means for transmitting, via a communication channel, the audio data and text data to a remote device.   
     
     
         17 . The device of  claim 16 , further comprising means for receiving, via the communication channel, a feedback signal from the remote device, wherein the transmitting means initiates the transmission of the text data to the remote device when the feedback signal indicates an occurrence of a predetermined condition at the remote device. 
     
     
         18 . The device of  claim 17 , wherein the predetermined condition includes a condition that a quality of the audio data received by the remote device is below a predetermined level. 
     
     
         19 . The device of  claim 16 , wherein the transmitting means comprises:
 means for transmitting the audio data to the remote device via a first communication modality; and   means for transmitting the text data to the remote device via a second communication modality having a higher robustness than the first communication modality.   
     
     
         20 . The device of  claim 16 , wherein the transmitting means comprises:
 means for transmitting the audio data to the remote device in a first quality of service level; and   means for transmitting the text data to the remote device in a second quality of service level that is higher than the first quality of service level.

Join the waitlist — get patent alerts

Track US2019073993A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.