Generation of timed text using speech-to-text technology, and applications thereof
Abstract
Embodiments relate to generation of timed text in web video. In an embodiment, a computer-implemented method generates timed text for online video. In the method, a request to play a timed text track of a video incorporated into a web video service is received from a client computing device. Prior to receipt of the request, audio of the video is processed to determine intermediate timed text data. The intermediate timed text data lacks a complete text transcription of the audio, but includes data to enable the complete text transcription to be generated when playing the video. In response to receipt of the request, a text transcription of the audio is determined using the intermediate data with an automated speech-to-text algorithm. Finally, the text transcription of the audio is sent to the client computing device for display along with the video.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
generating intermediate timed text data based on audio of a video, the intermediate timed text data comprising timing data to enable a text transcription to be generated when playing the video; determining the text transcription of the audio based on the intermediate timed text data in response to a request to play the video; and transmitting the text transcription of the audio to a client computing device for display along with the video.
2 . The method of claim 1 , wherein the intermediate timed text data lacks any text transcription of the audio.
3 . The method of claim 1 , wherein the generating the intermediate timed text data comprises:
determining a plurality of segments, each segment corresponding to a time period in the audio, wherein the intermediate timed text data comprises the plurality of segments.
4 . The method of claim 3 , wherein respective segments in the plurality of segments specify a type of sound played during a corresponding time period in the audio.
5 . The method of claim 1 , wherein the intermediate timed text data comprises data to enable a complete text transcription to be generated when playing the video,
wherein determining the text transcription comprises determining the text transcription of the audio in real time with the audio, and wherein transmitting the text transcription comprises transmitting the text transcription in real time to display along with the video.
6 . The method of claim 1 , further comprising:
providing an interface to enable a user to select one of a plurality of transcriptions to play, wherein the plurality of transcriptions comprises a user-generated transcription and an automatically generated transcription.
7 . The method of claim 1 , further comprising:
receiving the video to incorporate into an online video service, wherein the generating occurs in response to receipt of the video.
8 . The method of claim 1 , wherein the timing data comprises:
time codes indicative of when to display respective portions of the transcript to align the transcript with the audio of the video.
9 . A system, comprising:
a memory to store a video; a processor to:
generate intermediate timed text data based on audio of the video, the intermediate timed text data comprising timing data to enable a text transcription to be generated when playing the video;
determine the text transcription of the audio based on the intermediate timed text data in response to a request to play the video; and
transmit the text transcription of the audio to a client computing device for display along with the video.
10 . The system of claim 9 , wherein the intermediate timed text data lacks any text transcription of the audio.
11 . The system of claim 9 , wherein the processor is to generate the intermediate timed text data by:
determining a plurality of segments, each segment corresponding to a time period in the audio, wherein the intermediate timed text data comprises the plurality of segments.
12 . The system of claim 9 , wherein the intermediate timed text data comprises data to enable a complete text transcription to be generated when playing the video,
wherein determining the text transcription comprises determining the text transcription of the audio in real time with the audio, and wherein transmitting the text transcription comprises transmitting the text transcription in real time to display along with the video.
13 . The system of claim 9 , wherein the process is further to:
providing an interface to enable a user to select one of a plurality of transcriptions to play, wherein the plurality of transcriptions comprises a user-generated transcription and an automatically generated transcription.
14 . The system of claim 9 , wherein the timing data comprises:
time codes indicative of when to display respective portions of the transcript to align the transcript with the audio of the video.
15 . A non-transitory computer readable storage medium having instructions that, when executed by a processing device, cause the processing device to perform operations comprising
generating intermediate timed text data based on audio of a video, the intermediate timed text data comprising timing data to enable a text transcription to be generated when playing the video; determining the text transcription of the audio based on the intermediate timed text data in response to a request to play the video; and transmitting the text transcription of the audio to a client computing device for display along with the video.
16 . The non-transitory computer readable storage medium of claim 15 , wherein the intermediate timed text data lacks any text transcription of the audio.
17 . The non-transitory computer readable storage medium of claim 15 , wherein the generating comprises determining a plurality of segments, each segment corresponding to a time period in the audio, wherein the intermediate timed text data comprises the plurality of segments.
18 . The non-transitory computer readable storage medium of claim 15 , wherein the intermediate timed text data comprises data to enable a complete text transcription to be generated when playing the video,
wherein determining the text transcription comprises determining the text transcription of the audio in real time with the audio, and wherein transmitting the text transcription comprises transmitting the text transcription in real time to display along with the video.
19 . The non-transitory computer readable storage medium of claim 15 , wherein the operations further comprise:
providing an interface to enable a user to select one of a plurality of transcriptions to play, wherein the plurality of transcriptions comprises a user-generated transcription and an automatically generated transcription.
20 . The non-transitory computer readable storage medium of claim 15 , wherein the timing data comprises:
time codes indicative of when to display respective portions of the transcript to align the transcript with the audio of the video.Join the waitlist — get patent alerts
Track US2014142941A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.