US2020244800A1PendingUtilityA1

Semiautomated relay method and apparatus

Assignee: ULTRATEC INCPriority: Feb 28, 2014Filed: Apr 7, 2020Published: Jul 30, 2020
Est. expiryFeb 28, 2034(~7.6 yrs left)· nominal 20-yr term from priority
G10L 15/26G10L 15/01H04M 2201/60G10L 15/1815H04M 1/2475H04M 2203/2061H04M 3/42391H04M 2201/40G10L 25/48G10L 25/60G10L 15/265
73
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes obtaining a hearing user's (HU's) (first) voice signal including speech from a first device during a communication session between first and second devices, the session configured for verbal communication. The method also includes obtaining a first text string that is a transcription of the first voice signal via a first automatic speech recognition engine using the first voice signal, obtaining a second text string that is a transcription of the first voice signal via input obtained from a device associated with a call assistant, generating an output text string from the first and second text strings that includes one or more words from each of the first and second text strings, and providing the output text string as a transcription of the HU's voice signal for presentation during the session substantially concurrently with the presentation of the HU's voice signal by the second device.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 obtaining a hearing user's (HU's) voice signal originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the HU's voice signal includes speech, the HU's voice signal being a first voice signal;   obtaining a first text string that is a transcription of the first voice signal, the first text string generated by a first automatic speech recognition (ASR) engine using the first voice signal;   obtaining a second text string that is a transcription of the first voice signal, the second text string generated based on input obtained from a device associated with a call assistant (CA);   generating an output text string from the first text string and the second text string, the output text string includes one or more first words from the first text string and one or more second words from the second text string; and   providing the output text string as a transcription of the hearing user's voice signal for presentation during the communication session substantially concurrently with the presentation of the HU's voice signal by the second device.   
     
     
         2 . The method of  claim 1 , wherein the first ASR engine is not trained for transcribing a particular voice signal. 
     
     
         3 . The method of  claim 1 , wherein generating the output text string further includes:
 aligning the first text string and the second text string; and   comparing the aligned first and second text strings.   
     
     
         4 . The method of  claim 1 , wherein the step of providing includes providing the output text string, without providing the first text string and the second text to the second device. 
     
     
         5 . The method of  claim 1 , further comprising correcting at least one word in one or more of: the output text string, the first text string, and the second text string based on input obtained from a device associated with the CA. 
     
     
         6 . The method of  claim 1 , wherein the input obtained from the device is based on a third text string generated by the first ASR engine using the HU's voice signal. 
     
     
         7 . The method of  claim 6  wherein the first text string and the third text string are both estimates generated by the first ASR engine for substantially the same portion of the HU's voice signal. 
     
     
         8 . The method of  claim 1 , further comprising obtaining a third text string that is a transcription of the HU's voice signal, the third text string generated by a third ASR engine, wherein the output text string is generated from the first text string, the second text string, and the third text string. 
     
     
         9 . The method of  claim 6  wherein the first ASR engine persistently and automatically generates corrections to the first text string based on context within the first text string. 
     
     
         10 . The method of  claim 1  further including the steps of tracking a delay between generation of the second text string and the first text string, the step of generating an output text string including selecting one or more words from the first text string when the delay is less than a threshold value and selecting one or more works from the second text string when the delay is greater than a threshold value. 
     
     
         11 . At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of  claim 1 . 
     
     
         12 . A method comprising:
 obtaining a hearing user's (HU's) voice signal originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the HU's voice signal includes speech;   obtaining a first text string that is a transcription of the HU's voice signal, the first text string generated using an automatic speech recognition (ASR) system using the HU's voice signal;   obtaining a second text string that is a transcription of a second voice signal, the second voice signal including a revoicing of the HU's voice signal by a call assistant and the second text string generated by the ASR system using the second voice signal;   generating an output text string from the first text string and the second text string; and   using the output text string as a transcription of the speech.   
     
     
         13 . The method of  claim 12 , wherein the ASR system includes first and second ASR engines, the first ASR engine used to generate the first text string and the second ASR engine used to generate the second text string, the second ASR engine trained to the voice of the call assistant. 
     
     
         14 . The method of  claim 12 , wherein the output text string includes one or more first words from the first text string and one or more second words from the second text string. 
     
     
         15 . The method of  claim 12 , further comprising correcting at least one word in one or more of: the output text string, the first text string, and the second text string based on input obtained from a device associated with the call assistant. 
     
     
         16 . The method of  claim 15 , wherein the input obtained from the device is based on a third text string generated by the automatic speech recognition technology using the HU's voice signal. 
     
     
         17 . The method of  claim 16 , wherein the first text string and the third text string are both hypothesis generated by the automatic speech recognition technology for the substantially same portion of the HU's voice signal. 
     
     
         18 . The method of  claim 12 , further comprising obtaining a third text string that is a transcription of the HU's voice signal or the second voice signal, the third text string generated by the automatic speech recognition system, wherein the output text string is generated from the first text string, the second text string, and the third text string. 
     
     
         19 . The method of  claim 1  wherein the step of obtaining a second text string includes receiving a second voice signal that is a revoicing by the CA of the HU's voice signal, using a second ASR engine to transcribe the second voice signal to the second text string, the second ASR engine trained to the voice of the CA. 
     
     
         20 . At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of  claim 12 . 
     
     
         21 . The method of  claim 8  wherein the first text string and the third text string are both estimates for different segments of the HU's voice signal. 
     
     
         22 . The method of  claim 21  wherein the different segments of the HU's voice signal partially overlap. 
     
     
         23 . A method comprising:
 obtaining a hearing user's (HU's) voice signal originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the HU's voice signal includes speech;   obtaining a first text string that is a transcription of the HU's voice signal, the first text string generated by a first automatic speech recognition (ASR) engine using the HU's voice signal;   obtaining a second text string that is a transcription of a second voice signal, the second voice signal including a revoicing of the HU's voice signal by a call assistant and the second text string generated by a second ASR engine using the second voice signal, the second ASR engine trained to the voice of the call assistant;   generating an output text string from the first text string and the second text string, the output text string includes one or more first words from the first text string and one or more second words from the second text string; and   providing the output text string, without providing the first text string and the second text string, as a transcription of the speech to the second device for presentation during the communication session.   
     
     
         24 . The method of  claim 21  wherein the step of providing the output text string includes providing the output text string substantially concurrently with the presentation of the HU's voice signal by the second device. 
     
     
         25 . The method of  claim 21  wherein the step of generating an output text string includes generating the output text string based on an alignment between the first text string and the second text string. 
     
     
         26 . A method comprising:
 obtaining a hearing user's (HU's) voice signal originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the HU's voice signal includes speech;   obtaining a first text string that is a transcription of the HU's voice signal, the first text string generated by a first automatic speech recognition (ASR) engine using the HU's voice signal wherein the first ASR engine is not trained to a specific user's voice signal;   obtaining a second text string that is a transcription of a second voice signal, the second voice signal including a revoicing of the HU's voice signal by a call assistant and the second text string generated by a second ASR engine using the second voice signal wherein the second ASR engine is trained for the captioning the voice of the call assistant;   obtaining a third text string that is a transcription of the HU's voice signal, the third text string generated by a third ASR engine;   generating an output text string from the first text string, the second text string, and the third text string; and   providing the output text string as a transcription of the speech to the second device for presentation during the communication session substantially concurrently with the presentation of the HU's voice signal by the second device.   
     
     
         27 . A system comprising:
 one or more processors; and   at least one non-transitory computer-readable media coupled to the one or more processors, the at least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by the one or more processors cause the system to perform operations, the operations comprising:   obtain a hearing user's (HU's) voice signal originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the HU's voice signal includes speech;   obtain a first text string that is a transcription of the HU's voice signal, the first text string generated using automatic speech recognition (ASR) technology using the HU's voice signal;   obtain a second text string that is a transcription of a second voice signal, the second voice signal including a revoicing of the HU's voice signal and the second text string generated by the ASR technology using the second voice signal;   obtain a third text string that is a transcription of the first audio data, the third text string generated by the automatic speech recognition technology;   generate an output text string from the first text string, the second text string, and the third text string; and   provide the output text string as a transcription of the speech.

Join the waitlist — get patent alerts

Track US2020244800A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.