US2025227320A1PendingUtilityA1
Methods and systems for synchronization of closed captions with content output
Est. expiryMar 18, 2042(~15.6 yrs left)· nominal 20-yr term from priority
Inventors:Christopher Stone
H04N 21/44004H04N 21/8547H04N 21/4884H04N 21/43074H04N 21/466H04N 21/44008H04N 21/440236
74
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Alignment between closed caption and audio/video content may be improved by determining text associated with a portion of the audio or a portion of the video and comparing the determined text to a portion of closed caption text. Based on the comparison, a delay may be determined and the audio/video content may be buffered based on the determined delay.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A non-transitory computer-readable medium storing instructions that, when executed, cause:
receiving content, wherein the content comprises audio and closed caption text; determining text associated with at least a portion of the audio; determining, based on a comparison of the determined text to at least a portion of the closed caption text, a synchronization point corresponding to one or more matching portions; determining, based on the synchronization point, first timing data associated with the closed caption text and second timing data associated with the determined text; determining, based on a comparison of the first timing data and the second timing data, a misalignment between the audio and the closed caption text; and realigning, based on the determined misalignment, the audio and the closed caption text.
2 . The non-transitory computer-readable medium recited in claim 1 , wherein the instructions that, when executed, cause realigning the audio and the closed caption text cause slowing down the audio.
3 . The non-transitory computer-readable medium recited in claim 1 , wherein the instructions that, when executed, cause realigning the audio and the closed caption text cause speeding up the audio.
4 . The non-transitory computer-readable medium recited in claim 1 , wherein the instructions that, when executed, cause realigning the audio and the closed caption text cause buffering the audio.
5 . The non-transitory computer-readable medium recited in claim 1 , wherein the instructions that, when executed, cause determining the text cause decoding, by a player based on an audio-to-text translation, the at least the portion of the audio by the player.
6 . The non-transitory computer-readable medium recited in claim 1 , wherein the instructions that, when executed, cause determining the text cause converting descriptive audio of the content to text.
7 . The non-transitory computer-readable medium recited in claim 1 , wherein the content comprises an audiovisual stream.
8 . The non-transitory computer-readable medium recited in claim 1 , wherein the closed caption text comprises decoded closed captions.
9 . The non-transitory computer-readable medium recited in claim 1 , wherein the closed caption text comprises one or more subtitles.
10 . A system comprising:
a first computing device configured to send content; and a second computing device configured to:
receive the content, wherein the content comprises audio and closed caption text;
determine text associated with at least a portion of the audio;
determine, based on a comparison of the determined text to at least a portion of the closed caption text, a synchronization point corresponding to one or more matching portions;
determine, based on the synchronization point, first timing data associated with the closed caption text and second timing data associated with the determined text;
determine, based on a comparison of the first timing data and the second timing data, a misalignment between the audio and the closed caption text; and
realign, based on the determined misalignment, the audio and the closed caption text.
11 . The system recited in claim 10 , wherein the second computing device is configured to realign the audio and the closed caption text by slowing down the audio.
12 . The system recited in claim 10 , wherein the second computing device is configured to realign the audio and the closed caption text by speeding up the audio.
13 . The system recited in claim 10 , wherein the second computing device is configured to realign the audio and the closed caption text by buffering the audio.
14 . The system recited in claim 10 , wherein the second computing device is configured to determine the text by decoding, based on an audio-to-text translation, the at least the portion of the audio.
15 . The system recited in claim 10 , wherein the second computing device is configured to determine the text by converting descriptive audio of the content to text.
16 . The system recited in claim 10 , wherein the content comprises an audiovisual stream.
17 . The system recited in claim 10 , wherein the closed caption text comprises decoded closed captions.
18 . The system recited in claim 10 , wherein the closed caption text comprises one or more subtitles.
19 . A non-transitory computer-readable medium storing instructions that, when executed, cause:
receiving content, wherein the content comprises video, audio, and closed caption text; determining, based on an audio-to-text conversion of at least a portion of the audio, text; determining, based on a comparison of the determined text to at least a portion of the closed caption text, a delay; and buffering, based on the determined delay, at least one of the audio, the video, or the closed caption text.
20 . The non-transitory computer-readable medium recited in claim 19 , wherein the instructions, when executed, further cause:
outputting the buffered audio or the buffered video; and outputting the closed caption text.
21 . The non-transitory computer-readable medium recited in claim 19 , wherein the instructions, when executed, further cause synchronizing output of the buffered audio or the buffered video with output of the closed caption text.
22 . The non-transitory computer-readable medium recited in claim 19 , wherein the instructions that, when executed, cause determining the text cause decoding, by a player based on the audio-to-text translation of the at least the portion of the audio, the at least the portion of the audio.
23 . The non-transitory computer-readable medium recited in claim 19 , wherein the closed caption text comprises decoded closed captions.
24 . The non-transitory computer-readable medium recited in claim 19 , wherein the audio-to-text conversion comprises performing visual speech recognition on at least a portion of the video.
25 . The non-transitory computer-readable medium recited in claim 19 , wherein the content comprises an audiovisual stream.
26 . The non-transitory computer-readable medium recited in claim 19 , wherein the buffering facilitates alignment of the audio or the video with the closed caption text.
27 . The non-transitory computer-readable medium recited in claim 19 , wherein the closed caption text comprises one or more subtitles.
28 . A system comprising:
a first computing device configured to send content; and a second computing device configured to:
receive the content, wherein the content comprises video, audio, and closed caption text;
determine, based on an audio-to-text conversion of at least a portion of the audio, text;
determine, based on a comparison of the determined text to at least a portion of the closed caption text, a delay; and
buffer, based on the determined delay, at least one of the audio, the video, or the closed caption text.
29 . The system recited in claim 28 , wherein the second computing device is further configured to:
output the buffered audio or the buffered video; and output the closed caption text.
30 . The system recited in claim 28 , wherein the second computing device is further configured to synchronize output of the buffered audio or the buffered video with output of the closed caption text.
31 . The system recited in claim 28 , wherein the second computing device is configured to determine the text by decoding, based on the audio-to-text translation of the at least the portion of the audio, the at least the portion of the audio.
32 . The system recited in claim 28 , wherein the closed caption text comprises decoded closed captions.
33 . The system recited in claim 28 , wherein the audio-to-text conversion comprises performing visual speech recognition on at least a portion of the video.
34 . The system recited in claim 28 , wherein the content comprises an audiovisual stream.
35 . The system recited in claim 28 , wherein the buffering facilitates alignment of the audio or the video with the closed caption text.
36 . The system recited in claim 28 , wherein the closed caption text comprises one or more subtitles.Join the waitlist — get patent alerts
Track US2025227320A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.