System and method for extracting text captions from video and generating video summaries
Abstract
Caption boxes which are embedded in video content can be located and the text within the caption boxes decoded. Real time processing is enhanced by locating caption box regions in the compressed video domain and performing pixel based processing operations within the region of the video frame in which a caption box is located. The captions boxes are further refined by identifying word regions within the caption boxes and then applying character and word recognition processing to the identified word regions. Domain based models are used to improve text recognition results. The extracted caption box text can be used to detect events of interest in the video content and a semantic model applied to extract a segment of video of the event of interest.
Claims
exact text as granted — not AI-modified1 . A system for generating an event based summary of one or more fields or frames of video content including caption boxes embedded therein, comprising:
extracting means for extracting caption boxes from at least a portion of the fields or frames of the video content; change identifying means, coupled to the extracting means and receiving the caption boxes therefrom, for identifying one or more changes in the content of the extracted caption boxes which are indicative of an event of interest; and model applying means, coupled to the extracting means and change identifying means and receiving the caption boxes and the one or more changes therefrom, for applying a semantic model to select a portion of the video content preceding each of the one or more changes.
2 . The system of claim 1 , wherein the video content comprises video of a baseball game, and wherein the semantic model comprises a model which identifies the portion of the video content of the event of interest as residing between a pitching event and non-active views.
3 . The system of claim 2 , wherein the events of interest comprise score events and last pitch events.
4 . The system of claim 1 , wherein the extracting means further comprises:
location means for determining at least one expected location of a caption box in one or more fields or frames of the video content; determining means, coupled to the location means and receiving the at least one expected location therefrom, for determining at least one caption box mask within the expected location; frame identifying means, coupled to the determining means and location means and receiving the at least one caption box mask and the at least one expected location therefrom, for identifying one or more fields or frames in the video content as caption fields or frames if the current field or frame exhibits substantial correlation to the at least one caption box mask within the expected caption box;
word region identifying means, coupled to the frame identifying means and receiving the identified caption fields or frames therefrom, for at least a portion of the caption fields or frames, identifying word regions within the confines of the expected location; and
text character means, coupled to the word region identifying means and receiving the word regions therefrom, for each word region, identifying text characters within the region; and processing the identified text characters.
5 . The system of claim 4 , further comprising comparing means, coupled to the text character means and receiving the text characters therefrom, for comparing the text characters in the word region against a domain specific model to enhance word recognition.
6 . The system of claim 4 , wherein the location means includes means for evaluating motion features of the video field or frame in the compressed domain, evaluating texture features of the video field or frame in the compressed domain, and identifying regions having low motion features and high texture features as candidate caption box regions.
7 . The system of claim 5 , wherein the location means includes means for evaluating motion features of the video field or frame in the compressed domain, evaluating texture features of the video field or frame in the compressed domain, and identifying regions having low motion features and high texture features as candidate caption box regions.
8 . The system of claim 4 , further comprising removal means, coupled to the frame identifying means, the determining means and the location means and receiving the identified caption fields or frames, the at least one caption box mask and the at least one expected location therefrom, for evaluating the identified caption fields or frames, within the caption box location, for changes in content and removing caption fields or frames from word region processing which do not exhibit a change in content.
9 . The system of claim 4 , further comprising interval means, coupled to the frame identifying means and receiving the identified caption fields or frames therefrom, for selecting a subset of caption fields or frames based on a predetermined time interval and sending that subset to the word region identifying means.
10 . The system of claim 4 , further comprising number means, coupled to the frame identifying means and receiving the identified caption fields or frames therefrom, for determining a subset of caption fields or frames by selecting caption fields or frames based on a predetermined number of intervening caption fields or frames and sending that subset to the word region identifying means.
11 . A non-transitory computer-readable storage medium storing a program for causing a computer to implement a method for generating an event based summary of one or more fields or frames of video content including caption boxes embedded therein, comprising:
extracting caption boxes from at least a portion of the fields or frames of the video content; identifying one or more changes in the content of the extracted caption boxes which are indicative of an event of interest; and applying a semantic model to select a portion of the video content preceding each of the one or more changes.
12 . The non-transitory computer-readable storage medium of claim 11 , wherein the extracting caption boxes further comprises:
determining at least one expected location of a caption box in one or more fields or frames of the video content; determining at least one caption box mask within the expected location; identifying one or more fields or frames in the video content as caption fields or frames if the current field or frame exhibits substantial correlation to the at least one caption box mask within the expected caption box location; for at least a portion of the caption fields or frames, identifying word regions within the confines of the expected location; for each word region, identifying text characters within the region, and processing the identified text characters.
13 . The non-transitory computer-readable storage medium of claim 12 , wherein the extracting caption boxes includes the step of comparing the text characters in the word region against a domain specific model to enhance word recognition.
14 . A system for generating an event based summary of one or more fields or frames of video content including caption boxes embedded therein, comprising:
a caption box extractor for extracting caption boxes from at least a portion of the fields or frames of the video content; a caption change identification processor, coupled to the caption box extractor and receiving the caption boxes therefrom, for identifying one or more changes in the content of the extracted caption boxes which are indicative of an event of interest; and a model processor, coupled to the caption box extractor and caption change identification processor and receiving the caption boxes and the one or more changes therefrom, for applying a semantic model to select a portion of the video content preceding each of the one or more changes.
15 . The system of claim 14 , wherein the video content comprises video of a baseball game, and wherein the semantic model comprises a model which identifies the portion of the video content of the event of interest as residing between a pitching event and non-active views.
16 . The system of claim 14 , wherein the event of interest comprises one or more score events and last pitch events.
17 . The system of claim 14 , wherein the caption box extractor further comprises:
a location processor for determining at least one expected location of a caption box in one or more fields or frames of the video content; a caption mask determining processor, coupled to the location processor and receiving the at least one expected location therefrom, for determining at least one caption box mask within the expected location; a frame identification processor, coupled to the caption mask determining processor and location processor and receiving the at least one caption box mask and the at least one expected location therefrom, for identifying one or more fields or frames in the video content as caption fields or frames if the current field or frame exhibits substantial correlation to the at least one caption box mask within the expected caption box; a word region identification processor, coupled to the frame identification processor and receiving the identified caption fields or frames therefrom, for at least a portion of the caption fields or frames, identifying word regions within the confines of the expected location; and a text character identification processor, coupled to the word region identification processor and receiving the word regions therefrom, for each word region, identifying text characters within the region; and processing the identified text characters.
18 . The system of claim 17 , further comprising a text character comparison processor, coupled to the text character identification processor and receiving the text characters therefrom, for comparing the text characters in the word region against a domain specific model to enhance word recognition.
19 . The system of claim 17 , wherein the location processor includes a processor for evaluating motion features of the video field or frame in the compressed domain, evaluating texture features of the video field or frame in the compressed domain, and identifying regions having low motion features and high texture features as candidate caption box regions.
20 . The system of claim 18 , wherein the location processor includes a processor for evaluating motion features of the video field or frame in the compressed domain, evaluating texture features of the video field or frame in the compressed domain, and identifying regions having low motion features and high texture features as candidate caption box regions.
21 . The system of claim 17 , further comprising a caption removal processor, coupled to the frame identification processor, the caption mask determining processor and the location processor and receiving the identified caption fields or frames, the at least one caption box mask and the at least one expected location therefrom, for evaluating the identified caption fields or frames, within the caption box location, for changes in content and removing caption fields or frames from word region processing which do not exhibit a change in content.
22 . The system of claim 17 , further comprising an interval-based caption selection processor, coupled to the frame identification processor and receiving the identified caption fields or frames therefrom, for selecting a subset of caption fields or frames based on a predetermined time interval and sending that subset to the word region identification processor.
23 . The system of claim 17 , further comprising a number-based caption selection processor, coupled to the frame identification processor and receiving the identified caption fields or frames therefrom, for determining a subset of caption fields or frames by selecting caption fields or frames based on a predetermined number of intervening caption fields or frames and sending that subset to the word region identification processor.Join the waitlist — get patent alerts
Track US2013293776A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.