System and Method for AV Sync Correction by Remote Sensing
Abstract
Method and system for measuring audio video synchronization. This is done by having a remote sensing station to measure and process the audio and images from a multimedia terminal that displays a television program. The remote sensing station uses an audio sensor to sense the audio signal from the multimedia station and an image sensor to sense the images displayed on the multimedia station. An AV processing circuit in the remote sensing terminal processes the signals from the audio sensor and the sensed images. The delay information is then communicated to the multimedia station to adjust the AV synchronization.
Claims
exact text as granted — not AI-modified1 . A system for determining, correcting or determining and correcting time synchronization of an entertainment or information program having motion images carried by a video signal and associated sound carried by an audio signal, where provided are a multimedia terminal including at least one audio transducer responding to the program audio signal and providing audible program sound, a video display responding to the program video signal and displaying visible program images, and a first communication port; a remote sensing station including an audio sensor, an image sensor, and a second communication port the system comprising:
a remote sensing station including an associated image sensor for sensing visible program moving images displayed on a video display of a multimedia terminal and providing a sensed video signal in response; said remote sensing station including an associated audio sensor for sensing associated audible program audio provided by an audio transducer associated with said multimedia terminal and providing a sensed audio signal in response; said remote sensing station including an electronic processor device executing a program and operating to: a) in response to said sensed video signal identifying the presence of video MuEvs therein; b) in response to said sensed audio signal identifying the presence of audio MuEvs therein; c) in response to said video MuEvs and said audio MuEvs estimating the temporal mismatch of said audible program audio and said visible program moving images and in response thereto creating AV timing data; d) displaying said AV timing data for use by a user.
2 . A system as in claim 1 wherein said multimedia terminal includes a delay circuit responsive to said AV data to correct the audio video synchronization of said visible program moving images and said audible program audio provided by the multimedia terminal based on said estimated temporal mismatch by delaying the earlier of the video signal carrying said visible program moving images displayed on a video display or the audio signal carrying said associated audible program audio provided by an audio transducer
3 . A system as in claim 1 wherein said visible program moving images are made up of a temporal sequence of frames of individual still images which from time to time include a temporal sequence of still frames including an image of a talking person and each said video MuEv consists of a type of shape of the lips of said talking person in a given video frame which said type of shape corresponds to a spoken sound chosen from a group comprised only of known vowel sounds or a group comprised only of known consonants sounds.
4 . A system as in claim 1 wherein said visible program moving images are made up of a temporal sequence of frames of individual still images which from time to time include a temporal sequence of still frames including an image of a talking person and each said video MuEv consists of a type of shape of the lips of said talking person in a single video frame which said type of shape corresponds to a spoken sound chosen from a group comprised only of known vowel sounds and known consonants sounds.
5 . A system as in claim 1 wherein said associated audible program audio provided by an audio transducer associated with said multimedia terminal is from time to time made up of audio sounds which are made by a talking person and each said audio MuEv consists of a type of spoken sound chosen from a group comprised only of known vowel sounds or a group comprised only of known consonants sounds.
6 . A system as in claim 1 wherein said associated audible program audio provided by an audio transducer associated with said multimedia terminal is from time to time made up of audio sounds which are made by a talking person and each said audio MuEv consists of a type of spoken sound chosen only from a group comprised of known vowel sounds and known consonants sounds.
7 . A system as in claim 1 wherein said visible program moving images are made up of a temporal sequence of frames of individual still images which from time to time include a temporal sequence of still frames which include an image of a talking person and each said video MuEv consists of a type of shape of the lips of said talking person in a single video frame which shape corresponds only to one of sounds AA, EE, OO, “s”, “v”, “z”, “f”, “p”, “b” or “m”.
8 . A system as in claim 1 wherein said associated audible program audio provided by an audio transducer associated with said multimedia terminal is made up of audio sounds which from time to time is made by a talking person and each said audio MuEv consists only of one of the spoken sounds AA, EE, OO, “s”, “v”, “z”, “f”, “p”, “b” or “m”.
9 . A system for determining, correcting or determining and correcting time synchronization of a television program having motion images carried by a video signal and associated sound carried by an audio signal, where provided are a television including at least one audio transducer responding to the program audio signal to provide audible program sound, a video display responding to the program video signal and visibly displaying program images, and a first communication port; a sensing station the system comprising:
a television responsive to a video signal for visually displaying television program moving images and responsive to an associated audio signal for providing audible television sound; a sensing station having electronic circuitry including: e) an image processing circuit responsive to said video signal and identifying the presence of video MuEvs therein; f) an audio processing circuit responsive to said audio signal and identifying the presence of audio MuEvs therein; g) a comparison circuit for comparing the temporal timing of said video MuEvs relative to the temporal timing of said audio MuEvs to estimate the temporal mismatch of said visually displayed television program moving images and said audible television audio and in response thereto creating A/V timing data; h) a variable delay circuit responsive to said A/V timing data and operating to delay the earlier of said video signal and said audio signal before they are used by the television to generate said displayed program moving images and said audible television sound respectively in order that the respective displayed images and audible sound will be temporally aligned.
10 . A system as in claim 9 wherein said video signal made up of a sequence of frames of individual still images which from time to time include a sequence of still frames including an image of a talking person and each said video MuEv consists of a type of shape of the lips of said talking person in a single video frame which type of shape corresponds to a spoken sound chosen from a group comprised only of known vowel sounds or a group comprised only of known consonants sounds.
11 . A system as in claim 9 wherein said video signal is made up of a sequence of frames of individual still images which from time to time include a sequence of still frames including an image of a talking person and each said video MuEv consists of a type of shape of the lips of said talking person in a single video frame which shape corresponds to a spoken sound chosen from a group comprised only of known vowel sounds and known consonants sounds.
12 . A system as in claim 9 wherein said audio signal from time to time includes audio sounds made by a talking person and each said audio MuEv consists of a spoken sound chosen from a group comprised only of known vowel sounds or a group comprised of known consonants sounds.
13 . A system as in claim 9 wherein said audio signal from time to time is made up of audio sounds made by a talking person and each said audio MuEv consists of a spoken sound chosen from a group comprised only of known vowel sounds and known consonants sounds.
14 . A system as in claim 9 wherein said video signal is made up of a temporal sequence of frames of individual still images which from time to time include a temporal sequence of still frames including an image of a talking person and each said video MuEv consists of a shape of the lips of said talking person in a single video frame which shape corresponds only to one of sounds AA, EE, OO, “s”, “v”, “z”, “f”, “p”, “b” or “m”.
15 . A system as in claim 9 wherein said audio signal from time to time is made up of audio sounds which are made by a talking person and each said audio MuEv consists only of one of the spoken sounds AA, EE, OO, “s”, “v”, “z”, “f”, “p”, “b” or “m”.
16 . A system for determining, correcting or determining and correcting time synchronization of a television program having motion images carried by a video signal and associated sound carried by an audio signal, where provided are a television including at least one audio transducer responding to the program audio signal to provide audible program sound, a video display responding to the program video signal and visibly displaying program images, and a first communication port; a sensing station the system comprising:
a video display responsive to a video signal and providing a visible display of motion images; an audio transducer responsive to a related audio signal and providing an audible sound program; a video processing circuit responsive to said video signal and providing a temporal series of video MuEvs; an audio processing circuit responsive to said audio signal and providing a temporal series of audio MuEvs; an AV timing circuit responsive to said video MuEvs and said audio MuEvs to determine the temporal timing relationship between the movement of lips of a speaker whose image is carried by said video signal and the sound corresponding to said movement of lips which sound is carried by said audio signal and in response to said temporal timing relationship creating AV timing data; displaying said AV timing data on said video display or on a remote control device which operates to control said video display, said audio display or both.
17 . A system as claimed in claim 16 wherein said video processing circuit, said audio processing circuit and said AV timing circuit are remotely located relative to said video display and said audio transducer with said AV timing data being transmitted from said AV timing circuit to said video display via a wireless RF communications link.
18 . A system as claimed in claim 16 wherein said video processing circuit, said audio processing circuit and said AV timing circuit are part of a remote sensing station remotely located relative to said video display and said audio transducer with said AV timing data being transmitted from said remote sensing station to said video display via a wireless RF communications link for display on said video display.
19 . A system as claimed in claim 16 wherein said video processing circuit, said audio processing circuit and said AV timing circuit are incorporated in a remote sensing station which is part of a remote control for said video display, said remote control having a display, and on command from a user said AV timing data is displayed on said remote control display or transmitted from said remote control to said video display via a wireless RF communications link for display on said video display.
20 . A system as claimed in claim 16 wherein said video processing circuit, said audio processing circuit and said AV timing circuit are incorporated in a remote sensing station which is part of a remote control for one or both of said video display and said audio transducer, said remote control having a display:
said AV timing data is transmitted from said remote control to one or both of said video display and said audio transducer via one or more of a wireless RF or wireless optical communications link;
one or more of said video display and said audio transducer incorporating a variable delay circuit operating in response to said video signal or said audio signals respectively as well as in response to said AV timing data to delay the earlier of said video signal or said audio signal to reduce temporal timing errors between said visible display of motion images and said audible sound program.Join the waitlist — get patent alerts
Track US2016316108A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.