Meeting transcription using custom lexicons based on document history
Abstract
A collaborative content management system allows multiple users to access and modify collaborative documents. When audio data is recorded by or uploaded to the system, the audio data may be transcribed or summarized to improve accessibility and user efficiency. Text transcriptions are associated with portions of the audio data representative of the text, and users can search the text transcription and access the portions of the audio data corresponding to search queries for playback. An outline can be automatically generated based on a text transcription of audio data and embedded as a modifiable object within a collaborative document. The system associates hot words with actions to modify the collaborative document upon identifying the hot words in the audio data. Collaborative content management systems can also generate custom lexicons for users based on documents associated with the user for use in transcribing audio data, ensuring that text transcription is more accurate.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method comprising:
identifying, by a content creation system, a respective account for each speaker of a plurality of speakers, each respective account associated with one or more documents stored by the content creation system, each document of the one or more documents associated with information identifying a set of accounts as collaborators having accessed the document; determining, by the content creation system, a plurality of collaborative documents on which each speaker is a collaborator based on the information identifying the set of accounts as collaborators, each collaborative document of the plurality of collaborative documents having been determined for inclusion in the plurality of collaborative documents based on it being stored in association with accounts of all of the plurality of speakers; and generating, by the content creation system for the plurality of speakers, the custom lexicon based on the plurality of collaborative documents on which each speaker is a collaborator, wherein audio data received during a meeting of the plurality of speakers is transcribed using the custom lexicon.
2 . The computer-implemented method of claim 1 , wherein generating the custom lexicon further comprises:
identifying, by the content creation system, a set of words or n-grams included within the plurality of collaborative documents; and modifying, by the content creation system, a default lexicon to include the identified set of words or n-grams to generate the custom lexicon.
3 . The computer-implemented method of claim 1 , wherein generating the custom lexicon further comprises selecting among a plurality of lexicons associated with each speaker of the plurality of speakers based on a subject matter of the audio data.
4 . The computer-implemented method of claim 1 , wherein generating the custom lexicon further comprises selecting among a plurality of lexicons associated with each speaker of the plurality of speakers based on one or more characteristics of the one or more speakers.
5 . The computer-implemented method of claim 1 , wherein each of the plurality of collaborative documents is associated with a subject matter of the audio data.
6 . The computer-implemented method of claim 1 , wherein each of the plurality of collaborative documents is selected based on a characteristic of the meeting.
7 . The computer-implemented method of claim 1 , wherein a second custom lexicon is accessed in response to the custom lexicon not including text associated with a spoken word.
8 . A system comprising:
one or more processors; and a non-transitory computer-readable storage medium storing executable instructions that, when executed by the one or more processors, cause the system to perform steps comprising:
identifying, by a content creation system, a respective account for each speaker of a plurality of speakers, each respective account associated with one or more documents stored by the content creation system, each document of the one or more documents associated with information identifying a set of accounts as collaborators having accessed the document;
determining, by the content creation system, a plurality of collaborative documents on which each speaker is a collaborator based on the information identifying the set of accounts as collaborators, each collaborative document of the plurality of collaborative documents having been determined for inclusion in the plurality of collaborative documents based on it being stored in association with accounts of all of the plurality of speakers; and
generating, by the content creation system for the plurality of speakers, the custom lexicon based on the plurality of collaborative documents on which each speaker is a collaborator, wherein audio data received during a meeting of the plurality of speakers is transcribed using the custom lexicon.
9 . The system of claim 8 , wherein generating the custom lexicon further comprises:
identifying, by the content creation system, a set of words or n-grams included within the plurality of collaborative documents; and modifying, by the content creation system, a default lexicon to include the identified set of words or n-grams to generate the custom lexicon.
10 . The system of claim 8 , wherein generating the custom lexicon further comprises selecting among a plurality of lexicons associated with each speaker of the plurality of speakers based on a subject matter of the audio data.
11 . The system of claim 8 , wherein generating the custom lexicon further comprises selecting among a plurality of lexicons associated with each speaker of the plurality of speakers based on one or more characteristics of the one or more speakers.
12 . The system of claim 8 , wherein each of the plurality of collaborative documents is associated with a subject matter of the audio data.
13 . The system of claim 8 , wherein each of the plurality of collaborative documents is selected based on a characteristic of the meeting.
14 . The system of claim 8 , wherein a second custom lexicon is accessed in response to the custom lexicon not including text associated with a spoken word.
15 . A non-transitory computer-readable medium comprising memory with instructions encoded thereon that, when executed, cause one or more processors to perform operations, the instructions comprising instructions to:
identify, by a content creation system, a respective account for each speaker of a plurality of speakers, each respective account associated with one or more documents stored by the content creation system, each document of the one or more documents associated with information identifying a set of accounts as collaborators having accessed the document; determine, by the content creation system, a plurality of collaborative documents on which each speaker is a collaborator based on the information identifying the set of accounts as collaborators, each collaborative document of the plurality of collaborative documents having been determined for inclusion in the plurality of collaborative documents based on it being associated with accounts of all of the plurality of speakers; and generate, by the content creation system for the plurality of speakers, the custom lexicon based on the plurality of collaborative documents on which each speaker is a collaborator, wherein audio data received during a meeting of the plurality of speakers is transcribed using the custom lexicon.
16 . The non-transitory computer-readable medium of claim 15 , wherein the instructions to generate the custom lexicon further comprise instructions to:
identify, by the content creation system, a set of words or n-grams included within the plurality of collaborative documents; and modify, by the content creation system, a default lexicon to include the identified set of words or n-grams to generate the custom lexicon.
17 . The non-transitory computer-readable medium of claim 15 , wherein the instructions to generate the custom lexicon further comprise instructions to select among a plurality of lexicons associated with each speaker of the plurality of speakers based on a subject matter of the audio data.
18 . The non-transitory computer-readable medium of claim 15 , wherein the instructions to generate the custom lexicon further comprise instructions to select among a plurality of lexicons associated with each speaker of the plurality of speakers based on one or more characteristics of the one or more speakers.
19 . The non-transitory computer-readable medium of claim 15 , wherein each of the plurality of collaborative documents is associated with a subject matter of the audio data.
20 . The non-transitory computer-readable medium of claim 15 , wherein each of the plurality of collaborative documents is selected based on a characteristic of the meeting.Join the waitlist — get patent alerts
Track US2023042473A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.