Spatialized voice feedback
Abstract
Systems and methods for providing spatialized voice feedback for video conferences are provided. A computer-implemented method includes receiving a configuration for voice feedback for a client device including a spatial audio configuration including an apparent distance and direction. The method further includes receiving a first audio stream from a remote server. The method further includes receiving a second audio stream from an audio input device of the client device including a voice of a user of the client device. The method further includes playing the first audio stream on a first channel of an audio output device connected to the client device and a modified second audio stream on a second channel of the audio output device, in which the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and direction.
Claims
exact text as granted — not AI-modifiedThat which is claimed is:
1 . A computer-implemented method, comprising:
receiving, by a client device, a first configuration for voice feedback for the client device, comprising a spatial audio configuration including an apparent distance and an apparent direction; receiving a first audio stream from a remote server; receiving a second audio stream from an audio input device of the client device, the second audio stream comprising a voice of a user of the client device; playing the first audio stream on a first channel of a first audio output device connected to the client device; and playing a modified second audio stream on a second channel of the first audio output device connected to the client device, wherein the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and from the apparent direction.
2 . The computer-implemented method of claim 1 , wherein the modified second audio stream is generated using a binaural room impulse responses (“BRIR”) technique.
3 . The computer-implemented method of claim 2 , wherein the BRIR technique involves convolving the second audio stream with a BRIR corresponding to the apparent distance and the apparent direction.
4 . The computer-implemented method of claim 3 , wherein the BRIR corresponding to the apparent distance and the apparent direction is determined from a set of precomputed BRIRs.
5 . The computer-implemented method of claim 1 , where in the spatial audio configuration is generated in response to one or more inputs to a user interface comprising a spatial audio direction dial for selecting the apparent direction and a spatial audio distance slider for selecting the apparent distance.
6 . The computer-implemented method of claim 1 , wherein:
the spatial audio configuration is a three-dimensional (“3D”) spatial audio configuration, further including an apparent altitude; and the modified second audio stream is further configured to cause the user of the client device to hear the voice of the user coming from the apparent altitude.
7 . The computer-implemented method of claim 1 , wherein the first audio output device is an audio headphone including at least two earpieces, a first earpiece corresponding to the first channel and a second earpiece corresponding to the second channel.
8 . The computer-implemented method of claim 1 , wherein the remote server is a video conference provider and the client device is participating in a video conference hosted by the video conference provider, the video conference including a plurality of participating client devices.
9 . The computer-implemented method of claim 1 , further comprising:
prior to receiving the first configuration, receiving a first indication to enable voice feedback; after playing the modified second audio stream, receiving a second indication to disable voice feedback; and stopping the playing of the modified second audio stream.
10 . A non-transitory computer-readable storage medium storing processor-executable instructions configured to cause one or more processors to:
receive, by a client device, a first configuration for voice feedback for the client device, comprising a spatial audio configuration including an apparent distance and an apparent direction; receive a first audio stream from a remote server; receive a second audio stream from an audio input device of the client device, the second audio stream comprising a voice of a user of the client device; play the first audio stream on a first channel of a first audio output device connected to the client device; and play a modified second audio stream on a second channel of the first audio output device connected to the client device, wherein the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and from the apparent direction.
11 . The non-transitory computer-readable storage medium of claim 10 , wherein the modified second audio stream is generated using a binaural room impulse responses (“BRIR”) technique.
12 . The non-transitory computer-readable storage medium of claim 11 , wherein the BRIR technique involves convolving the second audio stream with a BRIR corresponding to the apparent distance and the apparent direction.
13 . The non-transitory computer-readable storage medium of claim 10 , wherein:
the spatial audio configuration is a three-dimensional (“3D”) spatial audio configuration, further including an apparent altitude; and the modified second audio stream is further configured to cause the user of the client device to hear the voice of the user coming from the apparent altitude.
14 . The non-transitory computer-readable storage medium of claim 10 , wherein the first audio output device is an audio headphone including at least two earpieces, a first earpiece corresponding to the first channel and a second earpiece corresponding to the second channel.
15 . The non-transitory computer-readable storage medium of claim 10 , comprising additional executable instructions configured to cause the one or more processors to:
prior to receiving the first configuration, receive a first indication to enable voice feedback; after playing the modified second audio stream, receive a second indication to disable voice feedback; and stop the playing of the modified second audio stream.
16 . A system comprising:
one or more non-transitory computer-readable media; and one or more processors communicatively coupled to the one or more non-transitory computer-readable media, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable media to:
receive, by a client device, a first configuration for voice feedback for the client device, comprising a spatial audio configuration including an apparent distance and an apparent direction;
receive a first audio stream from a remote server;
receive a second audio stream from an audio input device of the client device, the second audio stream comprising a voice of a user of the client device;
play the first audio stream on a first channel of a first audio output device connected to the client device; and
play a modified second audio stream on a second channel of the first audio output device connected to the client device, wherein the modified second audio stream is configured to cause the user of the client device to hear the voice of the user coming from the apparent distance and from the apparent direction.
17 . The system of claim 16 , wherein the modified second audio stream is generated using a binaural room impulse responses (“BRIR”) technique.
18 . The system of claim 17 , wherein the BRIR technique involves convolving the second audio stream with a BRIR corresponding to the apparent distance and the apparent direction.
19 . The system of claim 16 , wherein:
the spatial audio configuration is a three-dimensional (“3D”) spatial audio configuration, further including an apparent altitude; and the modified second audio stream is further configured to cause the user of the client device to hear the voice of the user coming from the apparent altitude.
20 . The system of claim 16 , wherein the one or more processors are configured to execute additional processor-executable instructions stored in the non-transitory computer-readable media to:
prior to receiving the first configuration, receive a first indication to enable voice feedback; after playing the modified second audio stream, receive a second indication to disable voice feedback; and stop the playing of the modified second audio stream.Join the waitlist — get patent alerts
Track US2026059073A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.