USRE50627EActiveUtility
Wired and wireless microphone arrays
Est. expiryJun 18, 2032(~5.9 yrs left)· nominal 20-yr term from priority
H04R 2201/401H04R 29/005H04R 3/005H04R 3/002
50
PatentIndex Score
0
Cited by
72
References
41
Claims
Abstract
An acoustic noise canceling microphone arrangement and processor that uses a principal microphone and other microphones that may be incidentally or deliberately located in the vicinity of the principal microphone in order to derive an audio signal of enhanced signal-to-background noise ratio. In one implementation, the principal and incidental microphones comprise the microphone built into a mobile phone and the microphone built into a Bluetooth headset.
Claims
exact text as granted — not AI-modifiedWe claim:
1. A system and apparatus for dynamically improving the ratio of a wanted speech signal from a principal speaker to random background noise that is not known or characterized a priori, comprising:
a principal microphone configured to be worn by said principal speaker and configured to produce a first principal audio signal containing a first sampling of the wanted speech plus unwanted background noise that has not previously been measured or characterized;
at least one incidental microphone located remotely from said principal microphone by several acoustic wavelengths at a mid-band audio frequency, and each said at least one incidental microphone configured to produce a second respective incidental audio signal containing a second sampling of at least said unwanted background noise that has not previously been measured or characterized; and
a signal processor configured to dynamically jointly process said first principal audio and second at least one said incidental audio signalssignal, without reference to noise profiles or filters constructed in advance, by
receiving and processing said first principal audio signal to determine a first set of individual spectral components at a set of predetermined frequencies;
receiving and processing the at least one said second incidental audio signal to determine at least one or more additional sets set of individual spectral components at said set of predetermined frequencies; and
dynamically combining corresponding spectral components from said first set and the at least one or more of said additional sets set to obtain a combined set of spectral components in which unwanted background noise components are reduced compared to wanted speech components; and
generating an output audio waveform solely from the combined set of spectral components, without filtering or suppressing noise by reference to a predetermined noise profile, in which the ratio of the wanted speech to unwanted noise is greater than the corresponding ratio for either the first principal or the second any of the at least one said incidental audio signals alone.
2. The system and apparatus of claim 1 in which said principal microphone is part of a Bluetooth wireless headset and said at least one incidental microphone is part of a Bluetooth-equipped communication device in wireless communication with said Bluetooth headset.
3. The system and apparatus of claim 1 , further comprising additional microphones producing additional audio signals containing different samplings of said wanted speech signal and unwanted background noise, and wherein the signal processor is configured to receive all of the first principal, second at least one said incidental, and said additional audio signals, and to derive therefrom a derived output signal to dynamically jointly process also said additional audio signals, wherein, in the output audio waveform generated, the ratio of said wanted signal speech to unwanted background noise is greater than the corresponding ratio for any of the audio signals alone.
4. The system and apparatus of claim 1 , further comprising additional microphones producing additional audio signals containing different samplings of said wanted speech signal and unwanted background noise, and wherein the signal processor is configured to receive all of the first principal, second at least one said incidental, and said additional audio signals and to derive a derived output signal by processing to dynamically jointly process the first principal and at least one incidental audio signals jointly with a selected one of the second and said additional audio signal signals, wherein the ratio of said wanted signal speech to unwanted background noise in the derived signal output audio waveform generated is greater than the corresponding ratio for any of the audio signals alone.
5. The system and apparatus of claim 1 in which said joint processing comprises time-domain to spectral domain converters conversion for separating said first principal and second at least one said incidental audio signals into said spectral components, a spectral combiner combining for performing weighted combining of corresponding said spectral components to produce a said combined set of spectral domain signal components, and a spectral domain to time domain converter conversion to convert said combined set of spectral domain signal components to generate said derived output signal audio waveform.
6. A system and apparatus for dynamically enhancing speech communications between a first multiplicity of speakers three or more persons in the presence of random acoustic background noise that is not known or characterized a priori, comprising:
a second multiplicity of microphones arranged such that for each of said first multiplicity of speakers, at least one of the second multiplicity of microphones is a principal microphone associated with that speaker, the second multiplicity of microphones producing a corresponding number of audio output signals containing different combinations of wanted speech anda plurality of at least three microphones arranged such that at least one microphone is associated with each said person, wherein each microphone outputs an audio signal containing speech if the associated person is speaking, or acoustic background noise alone that has not previously been measured or characterized; and
a signal processor configured to dynamically process jointly an audio output signal from a principle microphone containing speech along with one two or more other said audio output signals containing acoustic background noise in order to derive a derived output signal solely from the audio signals and without reference to noise profiles or filters constructed in advance, in which the a ratio of the speech signal from the principle microphone to unwanted acoustic background noise is greater than the a corresponding ratio for any one of said the audio output signals signal containing speech or the two or more said audio signals containing acoustic background noise alone;
wherein the joint processing of the audio output signal from the principle microphone containing speech and one the two or more other said audio output signals containing acoustic background noise includes, on a frame-block basis:
estimating a signal an inverse noise correlation matrix without reliance on stored statistics;
for each audio signal,
distinguishing between noise with whether speech present and noise without speech present,
updating the signal inverse noise spatial correlation matrix only if speech is present, and
calculating a frequency response from the updated signal inverse noise spatial correlation matrix if speech is present, or from a non-updated inverse noise spatial correlation matrix if speech is not present;
dynamically jointly processing the frequency responses response for each the audio signal containing speech and two or more audio signals containing acoustic background noise to derive an output signal in the frequency domain solely from the audio output signals and without reference to noise profiles or filters constructed in advance; and
converting the derived output signal to the a time domain.
7. The system and apparatus of claim 6 in which the an audio output signal of at least one of said multiplicity of plurality of at least three microphones is conveyed to said signal processor by a wireless link using any of a Bluetooth radio frequency link; a WiFi radio frequency link; a modulated Infra Red infrared link; an analog frequency-modulated link; a digital wireless link; a modulated visible light link; an inductively-coupled link and an electrostatically-coupled link.
8. The system and apparatus of claim 6 configured for a lecture hall environment in which said first multiplicity of speakers may three or more persons comprises a first group of speakers on stage and a second group speakers in the an audience, and said second multiplicity of plurality of at least three microphones comprises any combination of wireless microphones, lapel microphones, wireless headsets, fixed microphones, and roaming microphones.
9. The system and apparatus of claim 6 configured for use on the a flight deck of an aircraft, in which said second multiplicity of plurality of at least three microphones comprises the headsets provided for at least two crew members.
10. A system and apparatus for improving the a speech quality of conference calls using a telephone network, comprising:
a first conference phone installed at a first location and configured to serve a first group containing at least one intermittent speaker;
at least one second conference phone installed at a second location and configured to serve a second group containing at least one second intermittent speaker, the first phone and at least one second conference phones phone being in mutual communication via a telephone network;
at least two three microphones at at least one of said first or at least one second location configured to produce corresponding audio output signals containing respective samplings of a wanted speech signal and unwanted background noise;
a signal processor configured to receive said audio output signals from said at least two three microphones and to dynamically jointly process the at least two audio output signals to derive therefrom, solely from the audio output signals and without reference to noise profiles or filters constructed in advance, a derived output signal in which the a ratio of the wanted speech signal to unwanted background noise is greater than the a corresponding ratio for the audio output signal from any one alone of said at least two three microphones, said derived audio output signal from the signal processor being transmitted via said telephone network from the a location of the at least two three microphones to all other locations participating in the conference mutual communication via the telephone network;
wherein the joint processing of the audio output signal from the principle microphone and one or more other said audio output signals containing respective samplings of a wanted speech signal and unwanted background noise includes, on a frame-block basis:
estimating a signal correlation matrix without reliance on stored statistics;
estimating an inverse noise spatial correlation matrix;
for each audio signal,
distinguishing between noise with speech present and noise without speech present,
updating the signal correlation matrix only if speech is present, and
calculating a frequency response from the updated signal correlation matrix and the inverse noise spatial correlation matrix if speech is present, or from a non-updated signal correlation matrix and the inverse noise spatial correlation matrix if speech is not present;
dynamically jointly processing the frequency responses for each audio signal to derive an output signal in the a frequency domain solely from the audio signals and without reference to noise profiles or filters constructed in advance; and
converting the derived output signal to the a time domain.
11. The system and apparatus of claim 10 in which said at least two three microphones comprises any of one or more microphones associated with said first conference phone or said at least one second conference phone and connected thereto; any headset or lapel microphones worn by any person; any microphone contained by or connected to a laptop computer by wire or wireless means and any other fixed or hand-held microphones.
12. The system and apparatus of claim 10 in which said signal processor is located within said first conference phone or said at least one second conference phone, and the first conference phone or said at least one second conference phone is configured to receive the audio signals from said at least two three microphones using any of a wired connection; a wireless connection, or a connection to a server that forwards audio signals received at the server from any microphone.
13. The system and apparatus of claim 10 in which said signal processor is implemented in software on a server, the server being configured to receive audio signals from said at least two three microphones and to derive said derived output signal.
14. A method for improving the a signal to noise ratio of an audio signal received from a microphone associated with a principal active speaker, comprising the steps of:
providing a plurality of microphones;
associating at least one microphone of the plurality of microphones with each of a number of potential speakers;
determining the microphone of the plurality of microphones that is normally associated with the principal active speaker from among the number of potential speakers;
activating or maintaining in an active state at least one other microphone of the plurality of microphones that is normally associated with a speaker from among the number of potential speakers other than the principal active speaker; and
jointly processing, in a frequency domain and without performing beamforming, in a digital signal processor, the audio signals received from the microphone of the plurality of microphones that is normally associated with the principal active speaker and together with audio signals received from said at least one other microphone of the plurality of microphones in order to derive a processed signal in which the a ratio of the a wanted speech signal from the principal active speaker to background noise is greater than from any one microphone of the plurality of microphones alone.
15. The method of claim 14 in which the step of determining the microphone of the plurality of microphones that is associated with the principal active speaker is based on the a state of a press-to-talk switch associated with the microphone.
16. The method of claim 14 in which the step of determining the microphone of the plurality of microphones that is associated with the principal active speaker is based on an indication from a Voice Activity Detector associated with the microphone.
17. The method of claim 14 wherein jointly processing the audio signals received from the microphone of the plurality of microphones that is associated with the principal active speaker and said at least one other microphone of the plurality of microphones comprises:
decomposing all the such audio signals into a set of narrowband constituent components using a windowed Fast Fourier Transform;
processing overlapping blocks of signals, wherein the an overlap of a windowing function adds to unity, and applying frequency domain filtering on a frame-block basis;
estimating a signal correlation matrix and a an inverse noise spatial correlation matrix, based on a recursive linear squares algorithm modified for processing in a frequency domain, for each frame;
using voice activity detection on each audio signal to distinguish between noise with speech present and noise without speech present;
for each audio signal in each frame, updating the signal correlation matrix only if speech is present, and updating the inverse noise spatial correlation matrix only if speech is not detected present;
calculating Green's function for each frame from the updated signal correlation matrix if speech is present, or from a non-updated signal correlation matrix if speech is not present;
calculating a frequency response for each audio signal from the calculated Green's function and the updated signal inverse noise correlation matrix if speech is not present, or from a non-updated inverse noise correlation matrix if speech is present;
calculating an output signal in the frequency domain from the Green's function and frequency responses; and
converting the output signal to the a time domain using inverse Fast Fourier Transform.
18. The method of claim 17 , wherein the noise spatial correlation matrix is calculated using a recursive linear squares algorithm modified for processing in the frequency domain.
19. The method of claim 17 , further comprising calculating a power spectral density of the output signal if speech is detected, prior to the converting using the inverse Fast Fourier Transform.
20. A Press-To-Talk (PTT) communication system comprising:
at least two communication terminals, each terminal including a pressel switch used by an operator of the terminal to indicate active speech; and a signal processor operative to
continuously receive the state of the pressel switch from each terminal;
continuously receive an audio signal from each terminal, regardless of the state of the pressel switch;
determine, from the states of all pressel switches, a currently active speaker;
jointly process audio signals from the currently active speaker's terminal and at least one other terminal to derive an output audio signal in which the ratio of speech by the currently active speaker to background noise is greater than such ratio derived from any one terminal alone; and
output the derived output audio signal to at least one terminal.
21. The system and apparatus of claim 1 wherein dynamically jointly process processing said first principal audio signal and second the at least one said incidental audio signalssignal, without reference to noise profiles or filters constructed in advance, further comprises processing the audio signals under the a constraint that the a spectrum of the wanted speech is substantially unchanged.
22. The system and apparatus of claim 10 wherein the joint processing of the audio output signal from the principle microphone and one or more other said audio output signals containing respective samplings of a wanted speech signal and unwanted background noise comprises joint processing under the a constraint that the a spectrum of the wanted speech signal is substantially unchanged.
23. The method of claim 14 wherein jointly processing the audio signals received from the microphone associated with the principal active speaker and said at least one other microphone comprises jointly processing the audio signals under the a constraint that the a spectrum of the wanted speech signal from the principal active speaker is substantially unchanged.
24. The system of claim 1 in which said principal microphone is part of a Bluetooth wireless device and said at least one incidental microphone is part of a Bluetooth-equipped communication device in which wireless communication with said Bluetooth device.
25. A system for dynamically improving a ratio of a wanted audio signal to unwanted noise that is not known or characterized a priori, comprising:
a principal microphone configured to produce a principal audio signal containing a sampling of the wanted audio signal and unwanted noise that has not previously been measured or characterized; at least one incidental microphone located remotely from said principal microphone by several audio wavelengths at a mid-band audio frequency and configured to produce at least one corresponding incidental audio signal containing a sampling of said unwanted noise that has not previously been measured or characterized; and a signal processor configured to dynamically jointly process said principal audio signal and said at least one incidental audio signal, without reference to profiles or filters constructed in advance with respect to said unwanted noise,
by receiving and processing said principal audio signal to determine a first set of individual spectral components at a set of predetermined frequencies;
receiving and processing said at least one incidental audio signal to determine one or more additional sets of individual spectral components at said set or predetermined frequencies; and
dynamically combining corresponding spectral components from said first set and one or more of said additional sets to obtain a combined set of spectral components in which unwanted noise components are reduced compared to wanted audio signal components; and
generating an output audio waveform solely from the combined set of spectral components, without filtering or suppressing unwanted noise by reference to a predetermined noise profile, in which the ratio of the wanted audio signal to the unwanted noise is greater than a corresponding ratio for either said principal microphone or any of said at least one incidental microphone alone.
26. The system of claim 1 in which said principal microphone is part of a first wireless communications device and said at least one incidental microphone is a part of a second wireless communications device in wireless communication with said first wireless communications device.
27. The system of claim 1 wherein dynamically combining corresponding spectral components from said first set and the at least one said additional set comprises combining the corresponding spectral components by minimizing said background noise components under a nonlinear equality constraint leaving characteristics of the wanted speech components nominally undisturbed or having an a priori desired amplitude and phase distortion.
28. The system of claim 1 wherein positions of said principal microphone and said at least one incidental microphone are arbitrary relative to each other, other than said at least one incidental microphone being located remotely from said principal microphone by several acoustic wavelengths at a mid-band audio frequency.
29. The system of claim 1 wherein a position of at least one of said at least one incidental microphones is changing relative to said principal microphone.
30. The system of claim 6 wherein positions of the plurality of at least three microphones are arbitrary relative to each other.
31. The system of claim 6 wherein a position of at least one of the plurality of at least three microphones is changing relative to another of the plurality of at least three microphones.
32. The method of claim 14 wherein jointly processing audio signals received from the microphone of the plurality of microphones that is normally associated with the principal active speaker together with audio signals received from said at least one other microphone of the plurality of microphones comprises jointly processing the audio signals subject to a nonlinear equality constraint that represents a predetermined degree of degradation of the speech.
33. The method of claim 14 wherein jointly processing audio signals received from the microphone associated with the principal active speaker and said at least one other microphone comprises jointly processing the audio signals to minimize the background noise under a nonlinear equality constraint leaving the characteristics of the wanted speech signal nominally undisturbed or having a predetermined amplitude and phase distortion representing a degree of degradation of the wanted speech signal.
34. The system of claim 25 in which said principal microphone is part of a first wireless communications device and said at least one incidental microphone is a part of a second wireless communications device in wireless communication with said first wireless communications device.
35. The system of claim 25 , further comprising additional microphones producing additional audio signals containing different samplings of said wanted audio signal and unwanted background noise, and wherein the signal processor is configured to receive all of the principal, at least one incidental, and said additional audio signals, and to dynamically jointly process also said additional audio signals, wherein, in the output audio waveform generated, the ratio of said wanted audio signal to unwanted background noise is greater than the corresponding ratio for any of the audio signals alone.
36. The system of claim 25 , further comprising additional microphones producing additional audio signals containing different samplings of said wanted speech signal and unwanted background noise, and wherein the signal processor is configured to receive all of the principal, at least one incidental, and additional audio signals and to dynamically jointly process the principal and at least one incidental audio signals jointly with a selected one of the additional audio signals, wherein the ratio of said wanted audio signal to unwanted background noise in the output audio waveform generated is greater than the corresponding ratio for any of the audio signals alone.
37. The system of claim 25 in which said joint processing comprises time-domain to spectral domain conversion for separating said principal and at least one said incidental audio signal into said spectral components, spectral combining for performing weighted combining of corresponding said spectral components to produce said combined set of spectral components, and spectral domain to time domain conversion to convert said combined set of spectral components to generate said output audio waveform.
38. The system of claim 25 wherein dynamically combining corresponding spectral components from said first set and one or more of said additional sets comprises combining the corresponding spectral components by minimizing said background noise under a nonlinear equality constraint leaving characteristics of the wanted audio signal nominally undisturbed or having a predetermined amplitude and phase distortion representing a degree of degradation of the wanted audio signal.
39. The system of claim 25 wherein positions of said principal microphone and said at least one incidental microphone are arbitrary relative to each other, other than said at least one incidental microphone being located remotely from said principal microphone by several acoustic wavelengths at a mid-band audio frequency.
40. The system of claim 25 wherein a position of at least one of said at least one incidental microphones is changing relative to said principal microphone.
41. A system for dynamically enhancing speech communications between two people in the presence of random acoustic background noise that is not known or characterized a priori, comprising:
two microphones arranged such that one microphone is associated with each said person, wherein each microphone ouptuts an audio signal containing speech if the associated person is speaking, or an audio signal containing acoustic background noise alone that has not previously been measured or characterized; and a signal processor configured to dynamically process jointly an audio signal containing speech along with one audio signal containing acoustic background noise in order to derive a derived an output signal, without reference to noise profiles or filters constructed in advance, in which a ratio of speech to acoustic background noise is greater than a corresponding ratio for the audio signal containing speech alone, or for the one audio signal containing acoustic background noise alone; wherein the joint processing of the audio output signal containing speech and the one audio signal containing acoustic background noise includes, on a frame-block basis:
estimating a noise correlation matrix without reliance on stored statistics;
for each audio signal,
distinguishing whether speech is present,
updating the noise spatial correlation matrix only if speech is present, and
calculating a frequency response from the updated noise spatial correlation matrix if speech is present, or from a non-updated spatial correlation matrix if speech is not present;
dynamically jointly processing the frequency responses for the audio signal containing speech and the one audio signal containing acoustic background noise to derive an output signal in a frequency domain without reference to noise profiles or filters constructed in advance; and
converting the derived output signal to a time domain.Join the waitlist — get patent alerts
Track USRE50627E — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.