Variable audio for audio-visual content
Abstract
Various implementations disclosed herein include devices, systems, and methods that that modify audio of played back AV content based on context in accordance with some implementations. In some implementations audio-visual content of a physical environment is obtained, and the audio-visual content includes visual content and audio content that includes a plurality of audio portions corresponding to the visual content. In some implementations, a context for presenting the audio-visual content is determined, and a temporal relationship between one or more audio portions of the plurality of audio portions and the visual content is determined based on the context. Then, synthesized audio-visual content is presented based on the temporal relationship.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
at an electronic device having a processor:
obtaining audio-visual content of a physical environment, wherein the audio-visual content comprises visual content and audio content comprising a plurality of audio portions corresponding to the visual content;
determining a context for presenting the audio-visual content;
in accordance with the determined context, selecting an audio characteristic for presenting the plurality of audio portions with the visual content; and
presenting the visual content with the selected audio characteristics.
2 . The method of claim 1 , wherein selecting the audio characteristic comprises selecting spatialized audio.
3 . The method of claim 1 , wherein selecting the audio characteristic comprises selecting a portion of the plurality of audio portions, wherein, for the visual content, different portions of the plurality of audio portions of the audio content are selected for different contexts.
4 . The method of claim 1 , wherein selecting the audio content comprises selecting a portion of audio of the plurality of audio portions based on metadata for different portions of the plurality of audio portions identifying sources of the different portions.
5 . The method of claim 1 , wherein selecting the audio content comprises selecting a portion of the plurality of audio portions based on metadata for different portions of the plurality of audio portions identifying types of the different portions of audio content.
6 . The method of claim 1 , wherein the plurality of audio portions comprises audio of a user of an audio-visual (AV) capture device, low frequency audio, ambient audio, or a plurality of spatialized audio streams, wherein the visual content comprises at least a 2D image, a 3D image, a 2D sequence of images or a 3D sequence of images, a 3D photo, or a 3D video including corresponding audio.
7 . The method of claim 6 , further comprising semantically labelling sections of the plurality of audio portions based on metadata included with the audio-visual content of the physical environment or scene analysis of the corresponding visual content.
8 . The method of claim 7 , wherein the metadata comprises:
information related to the AV capture device including pose, movement, sensors, and sensor data of the AV capture device; information related to a user of the AV capture device including gaze, body movement, and operational inputs; information related to an environment of the AV capture device during capture; or information related to a scene or the visual content.
9 . The method of claim 1 , further comprising semantically labelling at least one section of the plurality of audio portions based on analysis of the audio content, wherein semantically labelling at least one section of the plurality of audio portions is performed by the AV capture device, a processing electronic device, or the electronic device.
10 . The method of claim 1 , wherein the audio content is decoupled from the visual content.
11 . The method of claim 1 , wherein determining the context for presenting the audio-visual content is based on actions of a user in an extended reality (XR) environment including a representation of the audio-visual content.
12 . The method of claim 1 , wherein determining the context for presenting the audio-visual content comprises determining at least whether the audio-visual content is selected based on user actions and determining a spatial distance between the user and a representation of the audio-visual content.
13 . The method of claim 1 , wherein determining the context comprises determining that a user has selected the visual content.
14 . The method of claim 1 , wherein determining the context comprises determining that a user is looking at the visual content.
15 . The method of claim 1 , wherein determining the context comprises determining that a user is looking away from the visual content.
16 . The method of claim 1 , wherein determining the context comprises determining that a user is within a threshold distance of the visual content in an extended reality (XR) environment.
17 . The method of claim 1 , wherein determining the context comprises determining that a user is more than a threshold distance from the visual content in an extended reality (XR) environment.
18 . The method of claim 1 , wherein determining the context comprises determining that a user is moving towards the visual content in an extended reality (XR) environment.
19 . A system comprising:
a non-transitory computer-readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising:
obtaining audio-visual content of a physical environment, wherein the audio-visual content comprises visual content and audio content comprising a plurality of audio portions corresponding to the visual content;
determining a context for presenting the audio-visual content;
in accordance with the determined context, selecting an audio characteristic for presenting the plurality of audio portions with the visual content; and
presenting the visual content with the selected audio characteristics.
20 . A non-transitory computer-readable storage medium, storing program instructions computer-executable on a computer to perform operations comprising:
obtaining audio-visual content of a physical environment, wherein the audio-visual content comprises visual content and audio content comprising a plurality of audio portions corresponding to the visual content; determining a context for presenting the audio-visual content; in accordance with the determined context, selecting an audio characteristic for presenting the plurality of audio portions with the visual content; and presenting the visual content with the selected audio characteristics.Join the waitlist — get patent alerts
Track US2023344973A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.