Method and apparatus for annotating video content with metadata generated using speech recognition technology
Abstract
A method and apparatus is provided for annotating video content with metadata generated using speech recognition technology. The method begins by rendering video content on a display device. A segment of speech is received from a user such that the speech segment annotates a portion of the video content currently being rendered. The speech segment is converted to a text-segment and the text-segment is associated with the rendered portion of the video content. The text segment is stored in a selectively retrievable manner so that it is associated with the rendered portion of the video content.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A computer-implemented method comprising:
providing a particular portion of a video for output; receiving an utterance while the particular portion of the video is being provided for output; obtaining, from an automated speech recognizer, a transcription of the utterance; generating video metadata based on the transcription; and updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata.
3 . The method of claim 1 , wherein updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating the header of the frame to include a sequence header, a GOP header, user data, and an I-frame header.
4 . The method of claim 1 , wherein:
providing a particular portion of a video for output comprises:
providing, by a set-top box, the particular portion of the video for output,
receiving an utterance while the particular portion of the video is being provided for output comprises:
receiving, by the set-top box, the utterance while the particular portion of the video is being provided for output,
obtaining, from an automated speech recognizer, a transcription of the utterance comprises:
obtaining, from the automated speech recognizer on the set-top box, the transcription of the utterance,
generating video metadata based on the transcription comprises:
generating, by the set-top box, the video metadata based on the transcription, and
updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating, by the set-top box, the header of the frame of data that corresponds to the particular portion of the video, to include the video metadata.
5 . The method of claim 1 , wherein generating video metadata based on the transcription comprises:
generating the video metadata based on a particular video standard.
6 . The method of claim 1 , wherein updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating the header of the frame of data to include the video metadata and an associated time-stamp as user data bits.
7 . The method of claim 1 , comprising:
generating a caption or a subtitle based on the video metadata; and storing the caption or the subtitle for display with the particular portion of the video.
8 . The method of claim 1 , comprising:
providing a predetermined user prompt, wherein the utterance is received in response to the predetermined user prompt.
9 . A system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
providing a particular portion of a video for output;
receiving an utterance while the particular portion of the video is being provided for output;
obtaining, from an automated speech recognizer, a transcription of the utterance;
generating video metadata based on the transcription; and
updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata.
10 . The system of claim 9 , wherein updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating the header of the frame to include a sequence header, a GOP header, user data, and an I-frame header.
11 . The system of claim 9 , wherein:
providing a particular portion of a video for output comprises:
providing, by a set-top box, the particular portion of the video for output,
receiving an utterance while the particular portion of the video is being provided for output comprises:
receiving, by the set-top box, the utterance while the particular portion of the video is being provided for output,
obtaining, from an automated speech recognizer, a transcription of the utterance comprises:
obtaining, from the automated speech recognizer on the set-top box, the transcription of the utterance,
generating video metadata based on the transcription comprises:
generating, by the set-top box, the video metadata based on the transcription, and
updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating, by the set-top box, the header of the frame of data that corresponds to the particular portion of the video, to include the video metadata.
12 . The system of claim 9 , wherein generating video metadata based on the transcription comprises:
generating the video metadata based on a particular video standard.
13 . The system of claim 9 , wherein updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating the header of the frame of data to include the video metadata and an associated time-stamp as user data bits.
14 . The system of claim 9 , wherein the operations further comprise:
generating a caption or a subtitle based on the video metadata; and storing the caption or the subtitle for display with the particular portion of the video.
15 . The system of claim 9 , wherein the operations further comprise:
providing a predetermined user prompt, wherein the utterance is received in response to the predetermined user prompt.
16 . A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
providing a particular portion of a video for output; receiving an utterance while the particular portion of the video is being provided for output; obtaining, from an automated speech recognizer, a transcription of the utterance; generating video metadata based on the transcription; and updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata.
17 . The medium of claim 16 , wherein updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating the header of the frame to include a sequence header, a GOP header, user data, and an I-frame header.
18 . The medium of claim 16 , wherein:
providing a particular portion of a video for output comprises:
providing, by a set-top box, the particular portion of the video for output,
receiving an utterance while the particular portion of the video is being provided for output comprises:
receiving, by the set-top box, the utterance while the particular portion of the video is being provided for output,
obtaining, from an automated speech recognizer, a transcription of the utterance comprises:
obtaining, from the automated speech recognizer on the set-top box, the transcription of the utterance,
generating video metadata based on the transcription comprises:
generating, by the set-top box, the video metadata based on the transcription, and
updating a header of a frame of data that corresponds to the particular portion of the video, to include the video metadata comprises:
updating, by the set-top box, the header of the frame of data that corresponds to the particular portion of the video, to include the video metadata.
19 . The medium of claim 16 , wherein generating video metadata based on the transcription comprises:
generating the video metadata based on a particular video standard.
20 . The medium of claim 16 , wherein the operations further comprise:
generating a caption or a subtitle based on the video metadata; and storing the caption or the subtitle for display with the particular portion of the video.
21 . The medium of claim 16 , wherein the operations further comprise:
providing a predetermined user prompt, wherein the utterance is received in response to the predetermined user prompt.Join the waitlist — get patent alerts
Track US2014331137A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.