US2026059073A1PendingUtilityA1

Spatialized voice feedback

Assignee: ZOOM COMMUNICATIONS INCPriority: Mar 3, 2023Filed: Jun 4, 2025Published: Feb 26, 2026
Est. expiryMar 3, 2043(~16.6 yrs left)· nominal 20-yr term from priority
G10L 21/013G10L 2021/0135G10L 21/0208G10L 2021/02082H04N 7/147H04M 2201/22H04M 2201/16H04M 3/002H04M 3/568G06F 3/165H04N 7/152
73
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for providing spatialized voice feedback for video conferences are provided. A computer-implemented method includes receiving a configuration for voice feedback for a client device including a spatial audio configuration including an apparent distance and direction. The method further includes receiving a first audio stream from a remote server. The method further includes receiving a second audio stream from an audio input device of the client device including a voice of a user of the client device. The method further includes playing the first audio stream on a first channel of an audio output device connected to the client device and a modified second audio stream on a second channel of the audio output device, in which the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and direction.

Claims

exact text as granted — not AI-modified
That which is claimed is: 
     
         1 . A computer-implemented method, comprising:
 receiving, by a client device, a first configuration for voice feedback for the client device, comprising a spatial audio configuration including an apparent distance and an apparent direction;   receiving a first audio stream from a remote server;   receiving a second audio stream from an audio input device of the client device, the second audio stream comprising a voice of a user of the client device;   playing the first audio stream on a first channel of a first audio output device connected to the client device; and   playing a modified second audio stream on a second channel of the first audio output device connected to the client device, wherein the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and from the apparent direction.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein the modified second audio stream is generated using a binaural room impulse responses (“BRIR”) technique. 
     
     
         3 . The computer-implemented method of  claim 2 , wherein the BRIR technique involves convolving the second audio stream with a BRIR corresponding to the apparent distance and the apparent direction. 
     
     
         4 . The computer-implemented method of  claim 3 , wherein the BRIR corresponding to the apparent distance and the apparent direction is determined from a set of precomputed BRIRs. 
     
     
         5 . The computer-implemented method of  claim 1 , where in the spatial audio configuration is generated in response to one or more inputs to a user interface comprising a spatial audio direction dial for selecting the apparent direction and a spatial audio distance slider for selecting the apparent distance. 
     
     
         6 . The computer-implemented method of  claim 1 , wherein:
 the spatial audio configuration is a three-dimensional (“3D”) spatial audio configuration, further including an apparent altitude; and   the modified second audio stream is further configured to cause the user of the client device to hear the voice of the user coming from the apparent altitude.   
     
     
         7 . The computer-implemented method of  claim 1 , wherein the first audio output device is an audio headphone including at least two earpieces, a first earpiece corresponding to the first channel and a second earpiece corresponding to the second channel. 
     
     
         8 . The computer-implemented method of  claim 1 , wherein the remote server is a video conference provider and the client device is participating in a video conference hosted by the video conference provider, the video conference including a plurality of participating client devices. 
     
     
         9 . The computer-implemented method of  claim 1 , further comprising:
 prior to receiving the first configuration, receiving a first indication to enable voice feedback;   after playing the modified second audio stream, receiving a second indication to disable voice feedback; and   stopping the playing of the modified second audio stream.   
     
     
         10 . A non-transitory computer-readable storage medium storing processor-executable instructions configured to cause one or more processors to:
 receive, by a client device, a first configuration for voice feedback for the client device, comprising a spatial audio configuration including an apparent distance and an apparent direction;   receive a first audio stream from a remote server;   receive a second audio stream from an audio input device of the client device, the second audio stream comprising a voice of a user of the client device;   play the first audio stream on a first channel of a first audio output device connected to the client device; and   play a modified second audio stream on a second channel of the first audio output device connected to the client device, wherein the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and from the apparent direction.   
     
     
         11 . The non-transitory computer-readable storage medium of  claim 10 , wherein the modified second audio stream is generated using a binaural room impulse responses (“BRIR”) technique. 
     
     
         12 . The non-transitory computer-readable storage medium of  claim 11 , wherein the BRIR technique involves convolving the second audio stream with a BRIR corresponding to the apparent distance and the apparent direction. 
     
     
         13 . The non-transitory computer-readable storage medium of  claim 10 , wherein:
 the spatial audio configuration is a three-dimensional (“3D”) spatial audio configuration, further including an apparent altitude; and   the modified second audio stream is further configured to cause the user of the client device to hear the voice of the user coming from the apparent altitude.   
     
     
         14 . The non-transitory computer-readable storage medium of  claim 10 , wherein the first audio output device is an audio headphone including at least two earpieces, a first earpiece corresponding to the first channel and a second earpiece corresponding to the second channel. 
     
     
         15 . The non-transitory computer-readable storage medium of  claim 10 , comprising additional executable instructions configured to cause the one or more processors to:
 prior to receiving the first configuration, receive a first indication to enable voice feedback;   after playing the modified second audio stream, receive a second indication to disable voice feedback; and   stop the playing of the modified second audio stream.   
     
     
         16 . A system comprising:
 one or more non-transitory computer-readable media; and   one or more processors communicatively coupled to the one or more non-transitory computer-readable media, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable media to:
 receive, by a client device, a first configuration for voice feedback for the client device, comprising a spatial audio configuration including an apparent distance and an apparent direction; 
 receive a first audio stream from a remote server; 
 receive a second audio stream from an audio input device of the client device, the second audio stream comprising a voice of a user of the client device; 
 play the first audio stream on a first channel of a first audio output device connected to the client device; and 
 play a modified second audio stream on a second channel of the first audio output device connected to the client device, wherein the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and from the apparent direction. 
   
     
     
         17 . The system of  claim 16 , wherein the modified second audio stream is generated using a binaural room impulse responses (“BRIR”) technique. 
     
     
         18 . The system of  claim 17 , wherein the BRIR technique involves convolving the second audio stream with a BRIR corresponding to the apparent distance and the apparent direction. 
     
     
         19 . The system of  claim 16 , wherein:
 the spatial audio configuration is a three-dimensional (“3D”) spatial audio configuration, further including an apparent altitude; and   the modified second audio stream is further configured to cause the user of the client device to hear the voice of the user coming from the apparent altitude.   
     
     
         20 . The system of  claim 16 , wherein the one or more processors are configured to execute additional processor-executable instructions stored in the non-transitory computer-readable media to:
 prior to receiving the first configuration, receive a first indication to enable voice feedback;   after playing the modified second audio stream, receive a second indication to disable voice feedback; and   stop the playing of the modified second audio stream.

Join the waitlist — get patent alerts

Track US2026059073A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.