Semiautomated relay method and apparatus
Abstract
A method to transcribe communications during a voice communication session between first and second devices includes obtaining an audio message originating at the first device. Upon initiating a captioning session, the message is provided to an automated speech recognition system to generate a first caption set for at least some of the message without any captioning assistance from a call assistant. The first caption set is made available to the second device for presentation on a display screen. In response to obtaining a first indication via the second device to change the captioning process, the audio message is provided to a second speech recognition system to generate a second caption set based on at least some of the message, where a call assistant performs error corrections to generate the second caption set. The second caption set then is made available to the second device for presentation on the display screen.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method to transcribe communications, the method comprising the steps of:
obtaining an audio message originating at a first device during a voice communication session between the first device and a second device, wherein the second device includes a display screen; upon initiation of a captioning session, providing the audio message to an automated speech recognition (ASR) system to generate a first caption set for at least a portion of the audio message without any captioning assistance from a call assistant (CA); making the first caption set available to the second device for presentation on the display screen; in response to obtaining a first indication via the second device to change a captioning process, providing the audio message to a second speech recognition system to generate a second caption set based on at least a portion of the audio message wherein a CA performs at least some error corrections to generate the second caption set; and making the second caption set available to the second device for presentation on the display screen.
2 . The method of claim 1 wherein the second device includes a processor that runs program code to provide the automated speech recognition system.
3 . The method of claim 2 wherein the second speech recognition system is remote from the second device, the method further including, in response to obtaining the first indication, transmitting the first caption set generated subsequent to the first indication to the second speech recognition system for consideration by the CA.
4 . The method of claim 3 wherein the second device presents the first caption set via the display screen throughout the captioning process and uses the second caption set to make error corrections in the first caption set presented on the display screen.
5 . The method of claim 4 wherein the CA visually examines the first caption set on a second display screen and makes changes to the first caption set to perform error correction resulting in the second caption set.
6 . The method of claim 1 wherein the automated speech recognition system is remote from the second device.
7 . The method of claim 1 wherein the second speech recognition system is remote from the second device.
8 . The method of claim 1 further including presenting a CA help option to a user of the second device and wherein the first indication occurs when the CA help option is selected.
9 . The method of claim 8 further comprising the step of, during the captioning session and prior to obtaining the first indication, presenting an indicator via the second device visually indicating that captions are being generated via an ASR engine.
10 . The method of claim 9 further comprising the step of, subsequent to obtaining the first indication, presenting an indicator via the second device visually indicating that the captions are being generated via CA assistance.
11 . The method of claim 1 wherein, subsequent to obtaining the first indication, the method further includes providing the first caption set to the CA, audibly broadcasting the audio message to the CA, and receiving error corrections to the first caption set to generate the second caption set.
12 . The method of claim 11 further including, while the second caption set is made available to the second device, obtaining a second indication to change the captioning process, and, in response to obtaining the second indication, ceasing providing the first caption set to the second device and to the CA, switching to a captioning process wherein, as the CA listens to the broadcast audio message, the method further include receiving CA input to generate a third caption set and receiving error corrections to the third caption set to generate the second caption set.
13 . The method of claim 12 further including providing the third caption set to the second device for presentation via the display screen prior to generation of the second caption set.
14 . The method of claim 12 wherein the step of receiving CA input to generate the third caption set includes receiving textual input from the CA.
15 . The method of claim 12 wherein the step of receiving CA input to generate the third caption set includes receiving a revoiced audio signal from the CA that corresponds to the audio message.
16 . The method of claim 12 further including presenting the third caption set on a display screen for the CA to view and receiving error corrections to the third caption set presented on the display screen.
17 . The method of claim 11 wherein the step of making the second caption set available to the second device includes transmitting at least the corrected captions to the second device.
18 . The method of claim 1 wherein error corrections are used by the ASR system to further train the ASR system.
19 . The method of claim 1 wherein the ASR system automatically trains without any CA input during an ongoing call to increase captioning accuracy.
20 . The method of claim 1 wherein, after the first indication is obtained, the second speech recognition system remains operational during a remainder of an ongoing call to generate the second caption set.
21 . The method of claim 1 further including, prior to the obtaining the first indication, assessing an accuracy level of the first caption set and presenting an accuracy indication via the second device.
22 . The method of claim 21 wherein the accuracy indication is expressed as a percentage.
23 . A method to transcribe communications, the method comprising the steps of:
obtaining an audio message originating at a first device during a voice communication session between the first device and a second device, wherein the second device includes a display screen; upon initiation of a captioning session, providing the audio message to an automated speech recognition (ASR) system to generate a first caption set for at least a portion of the audio message without any captioning assistance from a call assistant (CA); making the first caption set available to the second device for presentation on the display screen; presenting the first caption set to a CA via a display screen; broadcasting the audio message via a speaker to the CA; receiving a first indication from the CA to change a captioning process; in response to the first indication, enabling the CA to input information resulting in a second caption set that corresponds to the audio message; and making at least portions of the second caption set available to the second device for presentation on the display screen.
24 . The method of claim 23 wherein the information input by the CA resulting in the second caption set includes error corrections to captions corresponding to the audio message.
25 . The method of claim 23 wherein the second device includes a processor that runs program code to provide the automated speech recognition system.
26 . The method of claim 25 wherein the display screen used by the CA is remote from the second device.
27 . The method of claim 23 wherein the automated speech recognition system is remote from the second device.
28 . A method to transcribe communications, the method comprising the steps of:
obtaining an audio message originating at a first device during a voice communication session between the first device and a second device; upon initiation of a captioning session, providing the audio message to an automated speech recognition (ASR) system to generate a first caption set for at least a portion of the audio message without any captioning assistance from a call assistant (CA); making the first caption set available to a display screen for presentation; generating a confidence factor indicating how likely it is that at least a portion of the first caption set accurately reflects an associated portion of the audio message; and presenting the confidence factor via the display screen along with the at least a portion of the first caption set.
29 . The method of claim 28 further including presenting an interface selection tool for indicating that a call assistant (CA) captioning process should commence.
30 . The method of claim 29 further including presenting an interface selection tool for indicating that a different captioning process should commence.
31 . The method of claim 30 wherein a call assistant uses a computing device that includes the display screen.Join the waitlist — get patent alerts
Track US2021058510A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.