Eliminating spatial collisions due to estimated directions of arrival of speech
Abstract
A communication system may include, in an example, a first computing device communicatively coupled, via a network, to at least a second computing device maintained at a geographically distinct location than the first computing device; the first computing device including: an array of audio output devices and a processor to receive transmitted speech data and metadata describing an estimated direction of arrival (DOA) of speech from a plurality of speakers at an array of microphones at the second computing device and render audio at the array of audio output devices associated with the first computing device by eliminating spatial collision during rendering; said spatial collision arising due to the low angular separation of the estimated DOA of a plurality of speakers.
Claims
exact text as granted — not AI-modifiedWhat Is Claimed Is:
1 . A communication system, comprising:
an array of audio output devices; and a processor to render audio with the array of audio output devices, the audio being rendered from a signal produced by multiple sound sources at multiple locations relative to a microphone and including metadata describing an estimated direction of arrival (DOA) for each sound source; wherein the processor is to render the audio using the metadata to reduce spatial collision caused by at least two of the sound sources having an angular separation as indicated by the estimated DOA that is less than a threshold.
2 . The communication system of claim 1 , wherein the processor comprises a head-related transfer function to reduce the spatial collision when rendering the audio.
3 . The communication system of claim 1 , wherein the audio comprises speech and the multiple sound sources comprise a plurality of human speakers.
4 . The communication system of claim 1 , wherein the microphone comprises an array of individual microphones.
5 . The communication system of claim 1 , wherein the processor comprises an audio panning function to reduce the spatial collision by rendering the audio with a greater apparent angular separation between two sound sources when those two sound sources have an estimated angular separation below the threshold.
6 . The communication system of claim 1 , wherein the processor is programmed to reduce spatial collision by steering a sound field associated with respective sound sources to different spatial sound regions as reproduced by the array of audio output devices.
7 . The communication system of claim 1 , further comprising an array of microphones;
wherein the processor is further programmed to:
using microphones of the array, estimate a direction of arrival (DOA) of sound at the array of microphones; and
capture audio data of the incoming sound and associate the audio data with metadata describing the estimated DOA of that sound.
8 . A communication system, comprising:
an array of microphones; and a processor to,
using microphones of the array, estimate a direction of arrival (DOA) of sound at the array of microphones; and
capture audio data of the incoming sound and associate the audio data with metadata describing the estimated DOA of that sound.
9 . The communication system of claim 8 , wherein the incoming sound comprises speech from a plurality of human speakers at different locations.
10 . The communication system of claim 8 , further comprising:
an array of loudspeakers; the processor further to render audio with the array of loudspeakers, the audio being rendered from a signal produced by multiple sound sources at multiple locations relative to a microphone and including metadata describing an estimated direction of arrival (DOA) for each sound source; wherein the processor is to render the audio using the metadata to reduce spatial collision caused by at least two of the sound sources having an angular separation as indicated by the estimated DOA that is less than a threshold.
11 . The communication system of claim 10 , wherein the processor comprises a head-related transfer function to reduce the spatial collision when rendering the audio.
12 . The communication system of claim 10 , wherein the processor comprises an audio panning system to reduce the spatial collision by rendering the audio with a greater apparent angular separation between two sound sources when those two sound sources have an estimated angular separation below the threshold.
13 . The communication system of claim 10 , wherein the processor is programmed to reduce spatial collision due to the estimated DOAs by steering a sound field associated with respective sound sources to different spatial sound regions as reproduced by the array of loudspeakers.
14 . The communication system of claim 8 , wherein the processor is further to include in the metadata information describing a geometry of a room in which the array of microphones is disposed.
15 . The communication system of claim 8 , wherein the processor further comprises a crosstalk cancellation function to cancel crosstalk in the audio data captured with the array of microphones.
16 . A non-transitory computer-readable medium comprising instructions that, when executed by a processor of a communication system, cause the processor to
render audio data with an array of audio output devices, the audio data having been captured with an array of microphones and produced by multiple sound sources at multiple locations relative to the microphone array, the audio data including metadata describing an estimated direction of arrival (DOA) for sound from each sound source; render the audio data using the metadata so as to reduce spatial collision caused by at least two of the sound sources having an angular separation as indicated by the estimated DOA that is less than a threshold.
17 . The medium of claim 16 , wherein the instructions further comprise rules for determining an apparent location of one of the sound sources within audio rendered from the audio data, the apparent location being different from a location indicated by the DOA of sound from that sound source to reduce spatial collision.
18 . The medium of claim 17 , wherein the rules are based on a spatial resolution of human hearing.
19 . The medium of claim 16 , wherein the instructions further comprise a head-related transfer function to reduce the spatial collision when rendering the audio data.
20 . The medium of claim 16 , wherein the instructions further comprise an audio panning function to reduce the spatial collision by rendering audio from the audio data with a greater apparent angular separation between two sound sources when those two sound sources have an estimated angular separation below the threshold.Join the waitlist — get patent alerts
Track US2022201417A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.