US2026046375A1PendingUtilityA1

Contextual speech recognition of virtual meetings

Assignee: GOOGLE LLCPriority: Aug 8, 2024Filed: Aug 8, 2024Published: Feb 12, 2026
Est. expiryAug 8, 2044(~18 yrs left)· nominal 20-yr term from priority
Inventors:COHEN TAL
H04N 7/152G10L 15/26H04M 2203/552H04M 2201/40H04M 3/568H04L 51/02H04L 12/1831H04L 12/1827H04L 12/1822H04L 51/066H04N 7/15G10L 15/065G10L 15/063G10L 2015/0638G10L 15/16G10L 2015/228H04N 7/157G10L 15/22
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes receiving audio data of a virtual meeting and identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text. The method also includes causing the speech recognition system to be modified based on the previously unrecognized content. The method further includes causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising: 
 receiving audio data of a virtual meeting;   identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text;   causing the speech recognition system to be modified based on the previously unrecognized content; and   causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content.   
     
     
         2 . The method of  claim 1 , wherein the plurality of content items comprises at least one of: 
 names of participants of the virtual meeting;   documents related to the virtual meeting;   text shared between participants of the virtual meeting; or   documents associated with an organization of a participant of the virtual meeting.   
     
     
         3 . The method of  claim 1 , wherein an image is shared during the virtual meeting, and wherein the plurality of content items comprises text derived from processing the image using optical character recognition. 
     
     
         4 . The method of  claim 1 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises: 
 generating one or more possible pronunciations for the previously unrecognized content; and   adding the one or more possible pronunciations to a lexicon of the speech recognition system.   
     
     
         5 . The method of  claim 1 , wherein the speech recognition system comprises one or more machine learning models trained to convert speech data to corresponding text data. 
     
     
         6 . The method of  claim 5 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises: 
 generating one or more possible pronunciations for the previously unrecognized content;   generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output; and   retraining a first machine learning model of the one or more machine learning models using the generated training data.   
     
     
         7 . The method of  claim 5 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises: 
 generating one or more possible pronunciations for the previously unrecognized content;   generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output;   training a new machine learning model using the generated training data to recognize the previously unrecognized content; and   adding the new machine learning model to the one or more machine learning models of the speech recognition system.   
     
     
         8 . The method of  claim 5 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises providing a representation of the previously unrecognized content to at least a first machine learning model of the one or more machine learning models. 
     
     
         9 . The method of  claim 1 , wherein the text from causing the audio data of the virtual meeting to be converted using the modified speech recognition system is at least one of: 
 live captions visible during the virtual meeting;   a transcription of the virtual meeting; or   a summary of the virtual meeting generated using one or more machine learning models.   
     
     
         10 . A system comprising: 
 a memory device; and   a processing device coupled to the memory device, the processing device to perform operations comprising: 
 receiving audio data of a virtual meeting; 
 identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text; 
 causing the speech recognition system to be modified based on the previously unrecognized content; and 
 causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content. 
   
     
     
         11 . The system of  claim 10 , wherein the plurality of content items comprises at least one of: 
 names of participants of the virtual meeting;   documents related to the virtual meeting;   text shared between participants of the virtual meeting; or   documents associated with an organization of a participant of the virtual meeting.   
     
     
         12 . The system of  claim 10 , wherein an image is shared during the virtual meeting, and wherein the plurality of content items comprises text derived from processing the image using optical character recognition. 
     
     
         13 . The system of  claim 10 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises: 
 generating one or more possible pronunciations for the previously unrecognized content; and   adding the one or more possible pronunciations to a lexicon of the speech recognition system.   
     
     
         14 . The system of  claim 10 , wherein the speech recognition system comprises one or more machine learning models trained to convert speech data to corresponding text data. 
     
     
         15 . The system of  claim 14 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises: 
 generating one or more possible pronunciations for the previously unrecognized content;   generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output; and   retraining a first machine learning model of the one or more machine learning models using the generated training data.   
     
     
         16 . The system of  claim 14 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises: 
 generating one or more possible pronunciations for the previously unrecognized content;   generating training data comprising the one or more possible pronunciations as inputs and the previously unrecognized content as target output;   . a new machine learning model using the generated training data to recognize the previously unrecognized content; and   . the new machine learning model to the one or more machine learning models of the speech recognition system.   
     
     
         17 . The system of  claim 14 , wherein causing the speech recognition system to be modified based on the previously unrecognized content comprises providing a representation of the previously unrecognized content to at least a first machine learning model of the one or more machine learning models. 
     
     
         18 . The system of  claim 10 , wherein the text generated based on the audio data of the virtual meeting using the modified speech recognition system is at least one of: 
 live captions visible during the virtual meeting;   a transcription of the virtual meeting; or   a summary of the virtual meeting generated using one or more machine learning models.   
     
     
         19 . A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising: 
 receiving audio data of a virtual meeting;   identifying, within a plurality of content items related to the virtual meeting, content not previously recognized by a speech recognition system designated to convert the audio data of the virtual meeting into text;   causing the speech recognition system to be modified based on the previously unrecognized content; and   causing the audio data of the virtual meeting to be converted into the text using the modified speech recognition system, wherein the text comprises at least part of the previously unrecognized content.   
     
     
         20 . The non-transitory computer-readable storage medium of  claim 19 , wherein the plurality of content items comprises at least one of: 
 names of participants of the virtual meeting;   documents related to the virtual meeting;   text shared between participants of the virtual meeting; or   documents associated with an organization of a participant of the virtual meeting.

Join the waitlist — get patent alerts

Track US2026046375A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.