Transcription of communications
Abstract
A system may include a camera configured to obtain an image of a user, at least one processor, and at least one non-transitory computer-readable media communicatively coupled to the at least one processor. The non-transitory computer-readable media configured to store one or more instructions that when executed cause or direct the system to perform operations. The operations may include establish a communication session between the system and a device. The communication session may be configured such that the device provides audio for the system. The operations may further include compare the image to a particular user image associated with the system and select a first method of transcription generation from among two or more methods of transcription generation based on the comparison of the image to the particular user image. The operations may also include present, a transcription of the audio generated using the selected first method of transcription generation.
Claims
exact text as granted — not AI-modified1 . A method comprising:
obtaining audio of a communication session between a first device and a second device; obtaining an image of a user of the first device; comparing the image to a particular image; selecting a first speech recognition system from among two or more speech recognition systems based on the comparison of the image to the particular image, each of the two or more speech recognition systems including a different speech engine trained to automatically recognize speech in audio; and obtaining a transcription of the audio, the transcription of the audio generated using the selected first speech recognition system.
2 . The method of claim 1 , further comprising presenting the transcription of the audio in real-time during the communication session.
3 . The method of claim 2 , wherein the presentation of the transcription is performed by the second device.
4 . The method of claim 1 , further comprising establishing the communication session between the first device and the second device, wherein the image is obtained after establishing the communication session.
5 . The method of claim 1 , wherein the first speech recognition system is selected from among the two or more speech recognition systems based on the image not matching the particular image and the first speech recognition system automatically recognizes speech and is not specifically trained for the user.
6 . The method of claim 1 , wherein the first speech recognition system is selected from among the two or more speech recognition systems based on the image matching the particular image and the first speech recognition system automatically recognizes speech and is trained for the user.
7 . At least one non-transitory computer-readable media configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform the method of claim 1 .
8 . A system comprising:
at least one processor; and at least one non-transitory computer-readable media communicatively coupled to the at least one processor and configured to store one or more instructions that when executed by the at least one processor cause or direct the system to perform operations comprising:
obtain audio of a communication session between the system and a device;
obtain an image of a user of the device;
compare the image to a particular image;
select a first speech recognition system from among two or more speech recognition systems based on the comparison of the image to the particular image, each of the two or more speech recognition systems including a different speech engine trained to automatically recognize speech in audio; and
obtain a transcription of the audio, the transcription of the audio generated using the selected first speech recognition system.
9 . The system of claim 8 , wherein the operations further comprise present the transcription of the audio in real-time during the communication session.
10 . The system of claim 9 , wherein the transcription of the audio is presented by the system.
11 . The system of claim 8 , wherein the image of the user of the device is directed to the system for presentation of the image by the system.
12 . The system of claim 8 , wherein the operations further comprise establish the communication session between the system and the device, wherein the image is obtained after establishing the communication session.
13 . The system of claim 8 , wherein the first speech recognition system is selected from among the two or more speech recognition systems based on the image not matching the particular image and the first speech recognition system automatically recognizes speech and is not specifically trained for the user.
14 . The system of claim 8 , wherein the first speech recognition system is selected from among the two or more speech recognition systems based on the image matching the particular image and the first speech recognition system automatically recognizes speech and is trained for the user.
15 . A method comprising:
obtaining audio to be directed to a first device associated with a first user, the audio generated by a second device based on speech of a second user; obtaining an image of the second user; comparing the image to a particular image; selecting a first speech recognition system from among two or more speech recognition systems based on the comparison of the image to the particular image, each of the two or more speech recognition systems including a different speech engine trained to automatically recognize speech in audio; and obtaining a transcription of the audio, the transcription of the audio generated using the selected first speech recognition system.
16 . The method of claim 15 , wherein the transcription of the audio is generated for presentation to the first user of the first device.
17 . The method of claim 15 , wherein the first speech recognition system is selected from among the two or more speech recognition systems based on the image matching the particular image and the first speech recognition system automatically recognizes speech and is trained for the second user.
18 . The method of claim 15 , wherein the audio is directed to the first device during a communication session between the first device and the second device.
19 . The method of claim 15 , further comprising directing the transcription of the audio to the first device for presentation of the transcription of the audio by the first device.
20 . The method of claim 15 , wherein the image of the second user is directed to the first device for presentation of the image by the first device.Join the waitlist — get patent alerts
Track US2020184973A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.