Semiautomated relay method and apparatus
Abstract
A system includes a first user device configured to perform captioning session operations, a call-assistant (CA) device remote from the first user device, and a remote relay server separate from the CA device. The relay server initiates a captioning process, receives, from the first user device, a request to initiate a captioning session, establishes the session, assigns the session to the CA, receives first audio data from the first user device derived from a second user device, directs the first audio data to the CA device, receives, from the CA device, second audio data related to the first audio data and derived from CA speech, accesses an ASR engine trained to the CA voice, generates captioned text including a transcription of the second audio data, generates screen information including the transcription, directs the screen information to the CA device, and directs the captioned text to the first user device.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system comprising:
a first user device of a first user, the first user device being configured to perform operations related to a captioning session; a call-assistant (CA) device of a CA of a captioning system, the CA device being remotely located from the first user device; and a relay server communicatively coupled to and remotely located from the first user device, the relay server separate from the CA device, the relay server configured to: initiate a captioning process wherein the captioning process is configured to run a captioning software application, the captioning process specific to a specific CA; receive, from the first user device, a request to initiate a captioning session; establish the captioning session with the first user device; assign the captioning session to the CA; receive, from the first user device and in response to establishing the captioning session, first audio data that is derived from a second user device that is participating in a communication session with the first user device; direct the first audio data to the CA device in response to the captioning session being assigned to the CA; receive from the CA device, second audio data that is related to the first audio data and that is derived from speech of the CA; access, with the captioning software application, an automated speech recognition (ASR) engine that is trained to the voice of the CA based on the captioning session being assigned to the CA and based on the captioning process being specific to the CA; generate, with the ASR engine that is trained to the voice of the CA, captioned text that includes a transcription of the second audio data; generate, based on the transcription, screen information related to the captioning software application, the screen information including the transcription; direct the screen information to the CA device; and direct the captioned text to the first user device.
2 . The system of claim 1 , wherein:
the CA device is configured to: receive, from the CA, an error correction to correct at least one error in the transcription; and direct the error correction to the relay server; and the relay server is configured to: modify, with the captioning software, the captioned text based on the error correction; modify the screen information based on the modified captioned text; direct the modified screen data to the CA device; and direct the modified caption data to the first user device.
3 . The system of claim 1 , wherein the captioning process includes one or more of the following: an instance of an operating system, a CA interface, and an instance of the captioning software application.
4 . The system of claim 1 , wherein the captioning process is configured to:
establish the captioning session with the first user device; receive the first audio data from the first user device; direct the first audio data to the CA device; receive the second audio data from the CA device; generate the screen information; direct the screen information to the CA device; and direct the captioned text to the first user device.
5 . The system of claim 1 , wherein the CA device is configured as a user interface to access the captioning software application run by the server.
6 . The system of claim 1 , wherein CA device presents a customized computing environment for the call assistant based on the ASR trained to the call assistant's voice.
7 . A system comprising:
one or more processors; and one or more computer-readable storage media communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform operations related to a captioning session, the operations comprising: receive, from a first user device, first audio data that is derived from a second user device that is performing a communication session with the first user device, the first user device being configured to perform operations related to a captioning session; direct the first audio data to a remotely located call-assistant (CA) device; receive from the CA device, second audio data that is related to the first audio data and that is derived from speech of a CA using the CA device; access, with a captioning software application running in a captioning process, an automated speech recognition (ASR) engine that is trained to the voice of the CA; initiate the captioning process as a customized captioning process for the CA based on the ASR engine that is trained to the voice of the CA; generate, with the captioning software application, caption text that includes a transcription of the second audio data, the captioning software application being configured to use the ASR engine to generate the caption text; generate, based on the transcription, screen data related to the captioning software application, the screen data including the transcription; direct the screen data to the CA device; and direct the caption text to the first user device.
8 . The system of claim 7 , wherein the operations further comprise:
receive, by the captioning process from the first user device, a request to initiate the captioning session; and establish, by the captioning process, the captioning session with the first user device.
9 . The system of claim 7 , wherein the operations further comprise assign the captioning session to the CA.
10 . The system of claim 7 , wherein the operations further comprise:
receive, from the CA device, a modification command related to modification of the transcription; modify, with the captioning software application, the caption text based on the modification command; modify the screen data based on the modified caption data; direct the modified screen data to the CA device; and direct the modified caption data to the first user device.
11 . The system of claim 7 , wherein the captioning process includes one or more of the following: an instance of an operating system, a CA interface, and an instance of the captioning software application.
12 . The system of claim 7 , wherein the captioning software application is configured to:
receive the first audio data from the first user device; direct the first audio data to the CA device; receive the second audio data from the CA device; generate the screen data; direct the screen data to the CA device; and direct the caption data to the first user device.
13 . A method of performing captioning operations, the method being performed by a computing system and comprising:
initiating a first captioning process for use by a first call assistant (CA), the first captioning process being associated with the first CA and being configured to run a first instance of a captioning software application that includes a first ASR engine that is trained to the voice of the first CA; initiating a second captioning process for use by a second CA, the second captioning process being associated with the second CA and being configured to run a second instance of a captioning software application that includes a second ASR engine that is trained to the voice of the second CA; assigning a first captioning session to the first CA and the first captioning interface; assigning a second captioning session to the second CA and the second captioning interface; receiving, by the first instance of the captioning software application, first audio data from a remotely located first CA device of the first CA, the first audio data being derived from speech of the first CA; receiving, by the second instance of the captioning software application, second audio data from a remotely located second CA device of the second CA, the second audio data being derived from speech of the second CA; generating, with the first instance of the captioning software application, first caption data that includes a first transcription of the first audio data, the first instance of the captioning software application being configured to use the first ASR engine that is trained to the voice of the first CA to generate the first caption data; generating, with the second instance of the captioning software application, second caption data that includes a second transcription of the second audio data, the second instance of the captioning software application being configured to use the second ASR engine trained to the voice of the second CA to generate the second caption data; generating, based on the first transcription, first screen data related to the first instance of the captioning software application, the first screen data including the first transcription; generating, based on the second transcription, second screen data related to the second instance of the captioning software application, the second screen data including the second transcription; directing the first screen data to the first call-assistant device; directing the second screen data to the second call-assistant device; directing the first caption data to a first user device participating in a first communication session with a first other user device; and directing the second caption data to a second user device participating in a second communication session with a second other user device.
14 . The method of claim 13 , further comprising:
receiving, from the first CA device, a first modification command related to modification of the first transcription; modifying, with the first instance of the captioning software application, the first caption data based on the first modification command; modifying the first screen data based on the modified first caption data; communicating the modified first screen data to the first CA device; communicating the modified first caption data to the first user device; receiving, from the second CA device, a second modification command related to modification of the second transcription; modifying, with the second instance of the captioning software application, the second caption data based on the second modification command; modifying the second screen data based on the modified second caption data; communicating the modified second screen data to the second CA device; and communicating the modified second caption data to the second user device.
15 . The method of claim 13 , wherein each of the first and second captioning processes includes at least one of: an instance of an operating system, a user interface, and an instance of the captioning software application.
16 . The method of claim 13 , wherein the first and second instances of the captioning software application are run on a remote server.
17 . A method of performing captioning operations, the method being performed by a computing system and comprising:
initiating a captioning process wherein the captioning process is configured to run a captioning software application, the captioning process specific to a specific CA; receive, from the first user device, a request to initiate a captioning session; establish the captioning session with the first user device; assign the captioning session to the CA; receive, from the first user device and in response to establishing the captioning session, first audio data that is derived from a second user device that is participating in a communication session with the first user device; direct the first audio data to the CA device in response to the captioning session being assigned to the CA; receive from the CA device, second audio data that is related to the first audio data and that is derived from speech of the CA; access, with the captioning software application, an automated speech recognition (ASR) engine that is trained to the voice of the CA based on the captioning session being assigned to the CA and based on the captioning process being specific to the CA; generate, with the ASR engine that is trained to the voice of the CA, captioned text that includes a transcription of the second audio data; generate, based on the transcription, screen information related to the captioning software application, the screen information including the transcription; direct the screen information to the CA device; and direct the captioned text to the first user device.
18 . The method of claim 17 , further comprising:
receiving, by the captioning process from the first user device, a request to initiate the captioning session; and establishing, by the captioning process, the captioning session with the first user device.
19 . The method of claim 17 , wherein the captioning process includes one or more of the following: an instance of an operating system, a CA interface, and an instance of the captioning software application.Join the waitlist — get patent alerts
Track US2021234959A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.