Contextual speech recognition of virtual meetings
Abstract
A method includes receiving audio data of a virtual meeting and identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text. The method also includes causing the speech recognition system to be modified based on the previously unrecognized content. The method further includes causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
receiving audio data of a virtual meeting; identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text; causing the speech recognition system to be modified based on the previously unrecognized content; and causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content.
2 . The method of claim 1 , wherein the plurality of content items comprises at least one of:
names of participants of the virtual meeting; documents related to the virtual meeting; text shared between participants of the virtual meeting; or documents associated with an organization of a participant of the virtual meeting.
3 . The method of claim 1 , wherein an image is shared during the virtual meeting, and wherein the plurality of content items comprises text derived from processing the image using optical character recognition.
4 . The method of claim 1 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises:
generating one or more possible pronunciations for the previously unrecognized content; and adding the one or more possible pronunciations to a lexicon of the speech recognition system.
5 . The method of claim 1 , wherein the speech recognition system comprises one or more machine learning models trained to convert speech data to corresponding text data.
6 . The method of claim 5 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises:
generating one or more possible pronunciations for the previously unrecognized content; generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output; and retraining a first machine learning model of the one or more machine learning models using the generated training data.
7 . The method of claim 5 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises:
generating one or more possible pronunciations for the previously unrecognized content; generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output; training a new machine learning model using the generated training data to recognize the previously unrecognized content; and adding the new machine learning model to the one or more machine learning models of the speech recognition system.
8 . The method of claim 5 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises providing a representation of the previously unrecognized content to at least a first machine learning model of the one or more machine learning models.
9 . The method of claim 1 , wherein the text from causing the audio data of the virtual meeting to be converted using the modified speech recognition system is at least one of:
live captions visible during the virtual meeting; a transcription of the virtual meeting; or a summary of the virtual meeting generated using one or more machine learning models.
10 . A system comprising:
a memory device; and a processing device coupled to the memory device, the processing device to perform operations comprising:
receiving audio data of a virtual meeting;
identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text;
causing the speech recognition system to be modified based on the previously unrecognized content; and
causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content.
11 . The system of claim 10 , wherein the plurality of content items comprises at least one of:
names of participants of the virtual meeting; documents related to the virtual meeting; text shared between participants of the virtual meeting; or documents associated with an organization of a participant of the virtual meeting.
12 . The system of claim 10 , wherein an image is shared during the virtual meeting, and wherein the plurality of content items comprises text derived from processing the image using optical character recognition.
13 . The system of claim 10 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises:
generating one or more possible pronunciations for the previously unrecognized content; and adding the one or more possible pronunciations to a lexicon of the speech recognition system.
14 . The system of claim 10 , wherein the speech recognition system comprises one or more machine learning models trained to convert speech data to corresponding text data.
15 . The system of claim 14 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises:
generating one or more possible pronunciations for the previously unrecognized content; generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output; and retraining a first machine learning model of the one or more machine learning models using the generated training data.
16 . The system of claim 14 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises:
generating one or more possible pronunciations for the previously unrecognized content; generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output; . a new machine learning model using the generated training data to recognize the previously unrecognized content; and . the new machine learning model to the one or more machine learning models of the speech recognition system.
17 . The system of claim 14 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises providing a representation of the previously unrecognized content to at least a first machine learning model of the one or more machine learning models.
18 . The system of claim 10 , wherein the text generated based on the audio data of the virtual meeting using the modified speech recognition system is at least one of:
live captions visible during the virtual meeting; a transcription of the virtual meeting; or a summary of the virtual meeting generated using one or more machine learning models.
19 . A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising:
receiving audio data of a virtual meeting; identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text; causing the speech recognition system to be modified based on the previously unrecognized content; and causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content.
20 . The non-transitory computer-readable storage medium of claim 19 , wherein the plurality of content items comprises at least one of:
names of participants of the virtual meeting; documents related to the virtual meeting; text shared between participants of the virtual meeting; or documents associated with an organization of a participant of the virtual meeting.Join the waitlist — get patent alerts
Track US2026046375A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.