US2023344973A1PendingUtilityA1

Variable audio for audio-visual content

Assignee: APPLE INCPriority: Jul 16, 2020Filed: Jun 28, 2023Published: Oct 26, 2023
Est. expiryJul 16, 2040(~14 yrs left)· nominal 20-yr term from priority
H04N 9/802G10L 25/51G11B 27/10H04S 7/302H04N 21/47217H04N 21/4722H04N 21/439H04N 5/772H04N 9/8205H04R 2460/07H04R 1/1041H04R 5/033
61
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Various implementations disclosed herein include devices, systems, and methods that that modify audio of played back AV content based on context in accordance with some implementations. In some implementations audio-visual content of a physical environment is obtained, and the audio-visual content includes visual content and audio content that includes a plurality of audio portions corresponding to the visual content. In some implementations, a context for presenting the audio-visual content is determined, and a temporal relationship between one or more audio portions of the plurality of audio portions and the visual content is determined based on the context. Then, synthesized audio-visual content is presented based on the temporal relationship.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 at an electronic device having a processor:
 obtaining audio-visual content of a physical environment, wherein the audio-visual content comprises visual content and audio content comprising a plurality of audio portions corresponding to the visual content; 
 determining a context for presenting the audio-visual content; 
 in accordance with the determined context, selecting an audio characteristic for presenting the plurality of audio portions with the visual content; and 
 presenting the visual content with the selected audio characteristics. 
   
     
     
         2 . The method of  claim 1 , wherein selecting the audio characteristic comprises selecting spatialized audio. 
     
     
         3 . The method of  claim 1 , wherein selecting the audio characteristic comprises selecting a portion of the plurality of audio portions, wherein, for the visual content, different portions of the plurality of audio portions of the audio content are selected for different contexts. 
     
     
         4 . The method of  claim 1 , wherein selecting the audio content comprises selecting a portion of audio of the plurality of audio portions based on metadata for different portions of the plurality of audio portions identifying sources of the different portions. 
     
     
         5 . The method of  claim 1 , wherein selecting the audio content comprises selecting a portion of the plurality of audio portions based on metadata for different portions of the plurality of audio portions identifying types of the different portions of audio content. 
     
     
         6 . The method of  claim 1 , wherein the plurality of audio portions comprises audio of a user of an audio-visual (AV) capture device, low frequency audio, ambient audio, or a plurality of spatialized audio streams, wherein the visual content comprises at least a 2D image, a 3D image, a 2D sequence of images or a 3D sequence of images, a 3D photo, or a 3D video including corresponding audio. 
     
     
         7 . The method of  claim 6 , further comprising semantically labelling sections of the plurality of audio portions based on metadata included with the audio-visual content of the physical environment or scene analysis of the corresponding visual content. 
     
     
         8 . The method of  claim 7 , wherein the metadata comprises:
 information related to the AV capture device including pose, movement, sensors, and sensor data of the AV capture device;   information related to a user of the AV capture device including gaze, body movement, and operational inputs;   information related to an environment of the AV capture device during capture; or   information related to a scene or the visual content.   
     
     
         9 . The method of  claim 1 , further comprising semantically labelling at least one section of the plurality of audio portions based on analysis of the audio content, wherein semantically labelling at least one section of the plurality of audio portions is performed by the AV capture device, a processing electronic device, or the electronic device. 
     
     
         10 . The method of  claim 1 , wherein the audio content is decoupled from the visual content. 
     
     
         11 . The method of  claim 1 , wherein determining the context for presenting the audio-visual content is based on actions of a user in an extended reality (XR) environment including a representation of the audio-visual content. 
     
     
         12 . The method of  claim 1 , wherein determining the context for presenting the audio-visual content comprises determining at least whether the audio-visual content is selected based on user actions and determining a spatial distance between the user and a representation of the audio-visual content. 
     
     
         13 . The method of  claim 1 , wherein determining the context comprises determining that a user has selected the visual content. 
     
     
         14 . The method of  claim 1 , wherein determining the context comprises determining that a user is looking at the visual content. 
     
     
         15 . The method of  claim 1 , wherein determining the context comprises determining that a user is looking away from the visual content. 
     
     
         16 . The method of  claim 1 , wherein determining the context comprises determining that a user is within a threshold distance of the visual content in an extended reality (XR) environment. 
     
     
         17 . The method of  claim 1 , wherein determining the context comprises determining that a user is more than a threshold distance from the visual content in an extended reality (XR) environment. 
     
     
         18 . The method of  claim 1 , wherein determining the context comprises determining that a user is moving towards the visual content in an extended reality (XR) environment. 
     
     
         19 . A system comprising:
 a non-transitory computer-readable storage medium; and   one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising:
 obtaining audio-visual content of a physical environment, wherein the audio-visual content comprises visual content and audio content comprising a plurality of audio portions corresponding to the visual content; 
 determining a context for presenting the audio-visual content; 
 in accordance with the determined context, selecting an audio characteristic for presenting the plurality of audio portions with the visual content; and 
 presenting the visual content with the selected audio characteristics. 
   
     
     
         20 . A non-transitory computer-readable storage medium, storing program instructions computer-executable on a computer to perform operations comprising:
 obtaining audio-visual content of a physical environment, wherein the audio-visual content comprises visual content and audio content comprising a plurality of audio portions corresponding to the visual content;   determining a context for presenting the audio-visual content;   in accordance with the determined context, selecting an audio characteristic for presenting the plurality of audio portions with the visual content; and   presenting the visual content with the selected audio characteristics.

Join the waitlist — get patent alerts

Track US2023344973A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.