Artificially generated speech for a communication session
Abstract
A device is disclosed, which includes a processor and a memory in communication with the processor. The memory includes executable instructions that, when executed by the processor, cause the processor to control the device to perform functions of capturing a speech by a user; generating audio data representing the captured speech by a user; generating, based on the audio data, text data representing at least a portion of the captured speech; and transmitting, via a communication channel, the audio data and text data to the remote device. The device thus can provide the text data representing the captured speech when a quality of the audio signal received by the remote device is below a predetermined level.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A device comprising:
a processor; and a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor, cause the processor to control the device to perform functions of:
capturing a speech by a user;
generating audio data representing the captured speech by a user;
generating, based on the audio data, text data representing at least a portion of the captured speech; and
transmitting, via a communication channel, the audio data and text data to a remote device.
2 . The device of claim 1 , further comprising a microphone configured to capture the speech by the user.
3 . The device of claim 1 , further comprising a camera configured to capture an image of the user, wherein the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform a function of generating video data presenting the captured image of the user.
4 . The device of claim 3 , wherein the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform a function of encoding the audio data, video data and text data for transmission to the remote device via the communication network.
5 . The device of claim 1 , wherein the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform a function of receiving, via the communication channel, a feedback signal from the remote device, wherein the transmission of the text data to the remote device is initiated when the feedback signal indicates an occurrence of a predetermined condition at the remote device.
6 . The device of claim 5 , wherein the predetermined condition includes a condition that a quality of the audio data received by the remote device is below a predetermined level.
7 . The device of claim 1 , wherein to transmit the audio data and text data, the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform functions of:
transmitting the audio data to the remote device via a first communication modality; and transmitting the text data to the remote device via a second communication modality having a higher robustness than the first communication modality.
8 . The device of claim 1 , wherein to transmit the audio data and text data, the instructions further include instructions that, when executed by the processor, cause the processor to control the device to perform functions of:
transmitting the audio data to the remote device in a first quality of service level; and transmitting the text data to the remote device in a second quality of service level that is higher than the first quality of service level.
9 . A method comprising:
capturing a speech by a user; generating audio data representing the captured speech by a user; generating, based on the audio data, text data representing at least a portion of the captured speech; and transmitting, via a communication channel, the audio data and text data to a remote device.
10 . The method of claim 9 , further comprising:
capturing an image of the user; and generating video data presenting the captured image of the user.
11 . The method of claim 10 , further comprising encoding the audio data, video data and text data for transmission to the remote device via the communication network.
12 . The method of claim 9 , further comprising receiving, via the communication channel, a feedback signal from the remote device, wherein the transmission of the text data to the remote device is initiated when the feedback signal indicates an occurrence of a predetermined condition at the remote device.
13 . The method of claim 12 , wherein the predetermined condition includes a condition that a quality of the audio data received by the remote device is below a predetermined level.
14 . The method of claim 9 , wherein the transmitting the audio data and text data comprises:
transmitting the audio data to the remote device via a first communication modality; and transmitting the text data to the remote device via a second communication modality having a higher robustness than the first communication modality.
15 . The method of claim 9 , wherein the transmitting the audio data and text data comprises:
transmitting the audio data to the remote device in a first quality of service level; and transmitting the text data to the remote device in a second quality of service level that is higher than the first quality of service level.
16 . A device comprising:
means for capturing a speech by a user; means for generating audio data representing the captured speech by a user; means for generating, based on the audio data, text data representing at least a portion of the captured speech; and means for transmitting, via a communication channel, the audio data and text data to a remote device.
17 . The device of claim 16 , further comprising means for receiving, via the communication channel, a feedback signal from the remote device, wherein the transmitting means initiates the transmission of the text data to the remote device when the feedback signal indicates an occurrence of a predetermined condition at the remote device.
18 . The device of claim 17 , wherein the predetermined condition includes a condition that a quality of the audio data received by the remote device is below a predetermined level.
19 . The device of claim 16 , wherein the transmitting means comprises:
means for transmitting the audio data to the remote device via a first communication modality; and means for transmitting the text data to the remote device via a second communication modality having a higher robustness than the first communication modality.
20 . The device of claim 16 , wherein the transmitting means comprises:
means for transmitting the audio data to the remote device in a first quality of service level; and means for transmitting the text data to the remote device in a second quality of service level that is higher than the first quality of service level.Join the waitlist — get patent alerts
Track US2019073993A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.