System and method for identifying and segmenting repeating media objects embedded in a stream
Abstract
An “object extractor” automatically identifies and segments repeating media objects in a media stream. “Objects” are any section of non-negligible duration, i.e., a song, video, advertisement, jingle, etc., which would be considered to be a logical unit by a human listener or viewer. Identification and segmentation of repeating objects is achieved by directly comparing sections of the media stream to identify matching portions of the stream, then aligning the matching portions to identify object endpoints. Alternately, a suite of object dependent algorithms is employed to target particular aspects of the stream for identifying possible objects within the stream. Confirmation of possible objects as repeating objects is achieved by automatically searching for potentially matching objects in a dynamic object database, followed by a detailed comparison to one or more of the potentially matching objects. Object endpoints are then determined by automatic alignment and comparison to other copies of that object.
Claims
exact text as granted — not AI-modified1 . A method for identifying pairs of temporal endpoints for delimiting media objects which repeat in media stream, comprising;
identifying at least one instance when one or more unique media objects repeats in a media stream; aligning a portion of the media stream centered around one or more repeating instances of each unique media object with portions of the media stream centered around one or more other repeating instances of that unique media object; and comparing the aligned portions of the media stream to determine pairs of temporal endpoints for delimiting each repeating instance of each unique media object in the media stream.
2 . The method of claim 1 wherein comparing the aligned portions of the media stream comprises tracing backward and forward in each of the aligned portions of the media stream to determine locations within the aligned portions of the media stream where the aligned portions are still approximately equivalent to each other.
3 . The method of claim 2 wherein furthest backward and forward points of the aligned portions of the media stream where the aligned portions are still approximately equivalent to each other correspond to the temporal endpoints delimiting each repeating instance of each unique media object in the media stream.
4 . The method of claim 1 wherein comparing the aligned portions of the media stream comprises comparing a low-dimensional version of the aligned portions of the media stream.
5 . The method of claim 1 further comprising storing at least one representative copy of each repeating unique media object on a computer readable medium.
6 . The method of claim 1 further comprising storing the pairs of temporal endpoints for each repeating media object on a computer readable medium.
7 . The method of claim 6 wherein the stored pairs of temporal endpoints are used to prevent repeated searching of portions of the media stream already determined to correspond to repeating instances of media objects in the media stream.
8 . The method of claim 1 wherein the media stream is an audio media stream.
9 . The method of claim 1 wherein the media stream is a video stream.
10 . The method of claim 1 wherein the media objects are any of songs, music, advertisements, video clips, station identifiers, speech, images, and image sequences.
11 . A system for determining temporal endpoints of media objects embedded in a media stream, comprising steps for:
(a) extracting a segment of the media stream; (b) comparing the extracted segment of the media stream to the remainder of the media stream for identifying repeating content in the media stream where at least a part of the extracted segment matches one or more other parts of the media stream; (c) for one or more instances of repeating content within the media stream, aligning a portion of the media stream centered on the one or more instances of repeating content with portions of the media stream centered on one or more other instances of matching repeating content; (d) comparing each of the aligned portions of the media stream to identify a pair of temporal endpoints for defining temporal boundaries of each repeating media object.
12 . The system of claim 11 further comprising flagging each portion of the media stream between each pair of temporal endpoints as being identified.
13 . The system of claim 12 further comprising steps for:
extracting at least one new segment of the media stream from one or more portions of the media stream which have not been flagged as being identified; and repeating steps (b) through (d) for each new segment of the media stream.
14 . The system of claim 11 wherein comparing each of the aligned portions of the media stream comprises tracing backwards and forwards in each of the aligned portions of the media stream to determine forward and backward positions where the various aligned portions of the media stream begin to diverge from each other.
15 . The method of claim 14 wherein the determined positions within the aligned portions of the media stream correspond to the pair of temporal endpoints for defining the temporal boundaries of each repeating media object in the media stream.
16 . A computer-readable medium having computer executable instructions for locating repeating media objects within a media stream, comprising:
selecting a portion of the media stream; sequentially comparing the selected portion of the media stream to subsequent portions of the media stream to identify one or more instances within the media stream wherein a part of the selected portion at least partially matches a part of any of the subsequent portions; extracting a segment of the media stream centered on each matching part of the compared portions; simultaneously aligning two or more of the extracted segments with each other; and determining locations within the media stream of repeating media objects by determining where the simultaneously aligned segments of the media stream diverge in a forward and backwards direction from the center of each matching part of the compared portions.
17 . The computer-readable medium of claim 16 wherein sequentially comparing the selected portion of the media stream to subsequent portions of the media stream comprises comparing low dimensional versions of the portions of the media stream.
18 . The computer-readable medium of claim 16 wherein extracting the segments of the media stream comprises extracting low dimensional versions of the segments of the media stream.
19 . The computer-readable medium of claim 16 wherein the media stream is an audio media stream.
20 . The computer-readable medium of claim 16 wherein the media stream is a combined audio/video stream, and wherein only the audio component of the media stream is compared to identify audio/video media objects embedded in the media stream.
21 . The computer-readable medium of claim 16 wherein the media objects are any of songs, music, advertisements, video clips, station identifiers, speech, images, and image sequences.Join the waitlist — get patent alerts
Track US2005063667A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.