System and method for rich media annotation
Abstract
Disclosed herein are systems, methods, and computer readable-media for rich media annotation, the method comprising receiving a first recorded media content, receiving at least one audio annotation about the first recorded media, extracting metadata from the at least one of audio annotation, and associating all or part of the metadata with the first recorded media content. Additional data elements may also be associated with the first recorded media content. Where the audio annotation is a telephone conversation, the recorded media content may be captured via the telephone. The recorded media content, audio annotations, and/or metadata may be stored in a central repository which may be modifiable. Speech characteristics such as prosody may be analyzed to extract additional metadata. In one aspect, a specially trained grammar identifies and recognizes metadata.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method comprising:
obtaining, by a processing system including a processor, first metadata from an audio annotation of first media content, wherein the first metadata is obtained by speech recognition using a grammar, the first metadata indicative of one or more of a location, an individual or an activity;
storing, by the processing system, the first media content, the audio annotation and the first metadata;
based on weights assigned to the first metadata, identifying, by the processing system, second media content having an associated weighted second metadata corresponding to the first metadata; and
generating, by the processing system, at least one descriptive annotation for the first media content based on the first metadata and the second metadata.
2 . The method of claim 1 , wherein the first media content, the audio annotation and the first metadata are stored at a storage device.
3 . The method of claim 2 , wherein a user interface coupled to the storage device enables a user to modify one or more of the first media content, the audio annotation and the first metadata.
4 . The method of claim 3 , further comprising receiving, by the processing system, additional metadata comprising speech data provided in response to a prompt via the user interface.
5 . The method of claim 4 , wherein the prompt is based on a confidence score of the speech recognition.
6 . The method of claim 1 , further comprising:
receiving, by the processing system, additional metadata via a dialog with a user; and annotating, by the processing system, the first media content using the at least one descriptive annotation and the additional metadata to yield annotated first media content.
7 . The method of claim 6 , further comprising:
providing, by the processing system, the annotated first media content to a user device.
8 . The method of claim 1 , wherein the weights are assigned respectively to first metadata from each of a plurality of audio annotations of the first media content, based on a speech recognition certainty, and further comprising:
determining, by the processing system in accordance with the weights, a most likely correct audio annotation of the plurality of audio annotations.
9 . The method of claim 1 , wherein the audio annotation is recorded via a cellular phone.
10 . The method of claim 1 , wherein the first media content comprises user-generated video content.
11 . A device comprising:
a processing system including a processor; and a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations, the operations comprising: obtaining first metadata from an audio annotation of first media content, wherein the first metadata is obtained by speech recognition using a grammar;
storing the first media content, the audio annotation and the first metadata;
based on weights assigned to the first metadata, identifying second media content having an associated weighted second metadata corresponding to the first metadata; and
generating at least one descriptive annotation for the first media content based on the first metadata and the second metadata.
12 . The device of claim 11 , wherein the first media content, the audio annotation and the first metadata are stored at a storage device.
13 . The device of claim 12 , wherein a user interface coupled to the storage device enables a user to modify one or more of the first media content, the audio annotation and the first metadata.
14 . The device of claim 13 , wherein the operations further comprise receiving additional metadata comprising speech data provided in response to a prompt via the user interface, wherein the prompt is based on a confidence score of the speech recognition.
15 . The device of claim 11 , wherein the first metadata is indicative of one or more of a location, an individual or an activity.
16 . The device of claim 11 , wherein the operations further comprise:
receiving additional metadata via a dialog with a user; and annotating the first media content using the at least one descriptive annotation and the additional metadata to yield annotated first media content.
17 . The device of claim 11 , wherein the weights are assigned respectively to first metadata from each of a plurality of audio annotations of the first media content, based on a speech recognition certainty, and wherein the operations further comprise:
determining, in accordance with the weights, a most likely correct audio annotation of the plurality of audio annotations.
18 . A non-transitory machine-readable medium comprising executable instructions that, when executed by a processing system including a processor, facilitate performance of operations, the operations comprising:
obtaining first metadata from an audio annotation of first media content, wherein the first metadata is obtained by speech recognition, the first metadata indicative of one or more of a location, an individual or an activity;
storing the first media content, the audio annotation and the first metadata;
based on weights assigned to the first metadata, identifying second media content having an associated weighted second metadata corresponding to the first metadata; and
generating at least one descriptive annotation for the first media content based on the first metadata and the second metadata.
19 . The non-transitory machine-readable medium of claim 18 , wherein the first media content, the audio annotation and the first metadata are stored at a storage device, wherein a user interface coupled to the storage device enables a user to modify one or more of the first media content, the audio annotation and the first metadata.
20 . The non-transitory machine-readable medium of claim 18 , wherein the operations further comprise:
receiving additional metadata via a dialog with a user; and annotating the first media content using the at least one descriptive annotation and the additional metadata to yield annotated first media content.Join the waitlist — get patent alerts
Track US2021294833A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.