US2018068690A1PendingUtilityA1

Data processing apparatus, data processing method

Assignee: SONY CORPPriority: Jan 9, 2009Filed: Nov 13, 2017Published: Mar 8, 2018
Est. expiryJan 9, 2029(~2.5 yrs left)· nominal 20-yr term from priority
G06F 16/5866H04N 9/8227H04N 21/4884H04N 9/8715H04N 21/8456H04N 9/8233H04N 21/440236H04N 9/8211H04N 21/439G06F 16/00G06F 16/7867H04N 5/783G11B 27/28G06F 17/30268G06F 17/30H04N 21/4508G06F 17/3082G06F 16/587
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A data processing apparatus includes a text acquisition mechanism acquiring texts to be used as keywords which will be subject to audio retrieval, the texts being related to contents corresponding to contents data including image data and audio data; a keyword acquisition mechanism acquiring the keywords from the texts; an audio retrieval mechanism retrieving utterance of the keywords from the audio data of the contents data and acquiring timing information representing the timing of the utterance of the keywords of which the utterance is retrieved; and a playback control mechanism generating, from image data around the time represented by the timing information, representation image data of a representation image which will be displayed together with the keywords and performing playback control of displaying the representation image corresponding to the representation image data together with the keywords which are uttered at the time represented by the timing information.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . An apparatus, comprising:
 circuitry configured to:
 obtain a keyword from text data of content data; 
 detect timing information, which represents at least one of timing of appearance or timing of utterance of the keyword, based on the content data; and 
 generate digest data based on the timing information and the content data. 
   
     
     
         2 . The apparatus according to  claim 1 , wherein:
 the content data further includes caption data, and   the circuitry is further configured to acquire the caption data of the content data as the text data.   
     
     
         3 . The apparatus according to  claim 2 , wherein the circuitry is further configured to retrieve the utterance of the keyword with respect to audio data around a time at which a caption corresponding to the caption data is displayed. 
     
     
         4 . The apparatus according to  claim 1 , wherein the circuitry is further configured to acquire metadata of contents corresponding to the content data as the text data. 
     
     
         5 . The apparatus according to  claim 1 , wherein the circuitry is further configured to acquire input from a user as the text data. 
     
     
         6 . The apparatus according to  claim 5 , wherein the circuitry is further configured to acquire input from a keyboard operated by the user or a result of speech recognition of user's speech as the text data. 
     
     
         7 . A method, comprising
 obtaining a keyword from text data of content data;   detecting timing information, which represents at least one of timing of appearance or timing of utterance of the keyword, based on the content data; and   generating digest data based on the timing information and the content data.   
     
     
         8 . The method according to  claim 7 , wherein:
 the content data further includes caption data, and   the caption data of the content data is acquired as the text data.   
     
     
         9 . The method according to  claim 8 , further comprising retrieving the utterance of the keyword with respect to audio data around a time at which a caption corresponding to the caption data is displayed. 
     
     
         10 . The method according to  claim 7 , further comprising acquiring metadata of contents corresponding to the content data as the text data. 
     
     
         11 . The method according to  claim 7 , wherein further comprising acquiring input from a user as the text data. 
     
     
         12 . The method according to  claim 11 , wherein further comprising acquiring input from a keyboard operated by the user or a result of speech recognition of user's speech as the text data. 
     
     
         13 . A non-transitory computer-readable medium having stored thereon computer-readable instructions, which when executed by a computer, cause the computer to execute operations, the operations comprising:
 obtaining a keyword from text data of content data;   detecting timing information, which represents at least one of timing of appearance or timing of utterance of the keyword, based on the content data; and   generating digest data based on the timing information and the content data.

Join the waitlist — get patent alerts

Track US2018068690A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.