Spectral and Spatial Modification of Noise Captured During Teleconferencing
Abstract
In some embodiments, a method for modifying noise captured at endpoints of a teleconferencing system, including steps of capturing noise at each endpoint, and modifying the captured noise to generate modified noise having a frequency-amplitude spectrum which matches a target spectrum and a spatial property set which matches a target spatial property set. In other embodiments, a teleconferencing method including steps of: at endpoints of a teleconferencing system, determining audio frames indicative of audio captured at each endpoint, each of a subset of the frames indicative of noise but not a significant level of speech; and at each endpoint, generating modified frames indicative of modified noise having a frequency-amplitude spectrum which matches a target spectrum and a spatial property set which matches a target spatial property set, and generating encoded audio including by encoding the modified frames. Other aspects are systems configured to perform any embodiment of the method.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for modifying noise captured during a conference at each endpoint of a set of at least two endpoints of a teleconferencing system, said method including the steps of:
(a) generating first noise samples indicative of noise captured during the conference at a first one of the endpoints and second noise samples indicative of noise captured during the conference at a second one of the endpoints; and (b) modifying the first noise samples to generate first modified noise samples indicative of modified noise having a frequency-amplitude spectrum which at least substantially matches a target spectrum, and at least one spatial property which at least substantially matches at least one target spatial property, and modifying the second noise samples to generate second modified noise samples indicative of modified noise having a frequency-amplitude spectrum which at least substantially matches the target spectrum, and at least one spatial property which at least substantially matches the target spatial property.
2 . The method of claim 1 , wherein the first noise samples are a first subset of a set of frames of audio samples captured during the conference, each frame of a second subset of the set of frames is indicative of speech uttered by a conference participant, each frame of the first subset is indicative of noise but not a significant level of speech, and the first modified noise samples are generated by modifying the first subset of the set of frames.
3 . The method of claim 1 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the first noise samples, and wherein step (b) includes the steps of:
for said each band, determining a covariance matrix indicative of at least one spatial property of the noise in the band indicated by the first noise samples; generating intermediate samples by applying band-specific gain to said each band of the first noise samples; and processing the intermediate samples to generate the first modified noise samples, including by applying a band-specific warping matrix to each band of the intermediate samples, wherein the warping matrix for each band is determined by the covariance matrix for the band and the at least one target spatial property for said band.
4 . The method of claim 1 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the first noise samples, and wherein step (b) includes the steps of:
for said each band, determining a data structure indicative of at least one spatial property of the noise in the band indicated by the first noise samples; generating intermediate samples by applying band-specific gain to said each band of the first noise samples; and processing the intermediate samples to generate the first modified noise samples, including by applying a band-specific warping matrix to each band of the intermediate samples, wherein the warping matrix for each band is determined by the data structure for the band and the at least one target spatial property for said band.
5 . The method of claim 1 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the first noise samples, and wherein step (b) includes the steps of:
for said each band, determining a covariance matrix indicative of at least one spatial property of the noise in the band indicated by the first noise samples; and processing the first noise samples to generate the first modified noise samples, including by applying a band-specific spectral modification and warping matrix to each band of the first noise samples, wherein the spectral modification and warping matrix for each band is determined by the covariance matrix for the band, the at least one target spatial property for said band, and the target spectrum.
6 . The method of claim 1 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the first noise samples, and wherein step (b) includes the steps of:
for said each band, determining a data structure indicative of at least one spatial property of the noise in the band indicated by the first noise samples; and processing the first noise samples to generate the first modified noise samples, including by applying a band-specific spectral modification and warping matrix to each band of the first noise samples, wherein the spectral modification and warping matrix for each band is determined by the data structure for the band, the at least one target spatial property for said band, and the target spectrum.
7 . The method of claim 1 , wherein each of the endpoints is a telephone system.
8 . A teleconferencing method, including the steps of:
(a) at each endpoint of a set of at least two endpoints of a teleconferencing system, determining a sequence of audio frames indicative of audio captured at the endpoint during a conference, wherein each frame of a first subset of the frames is indicative of speech uttered by a conference participant, each frame of a second subset of the frames is indicative of noise but not a significant level of speech; (b) at said each endpoint, generating modified frames by modifying each frame of the second subset of the frames, such that each of the modified frames is indicative of modified noise having a frequency-amplitude spectrum which at least substantially matches a target spectrum, and a spatial property set which at least substantially matches a target spatial property set; and (c) at said each endpoint, generating encoded audio including by encoding the modified frames and encoding each frame of the first subset of the frames.
9 . The method of claim 8 , also including the steps of:
transmitting the encoded audio generated at said each endpoint to a server of the teleconferencing system; and at the server, generating conference audio indicative of a mix or sequence of audio captured at different ones of the endpoints.
10 . The method of claim 8 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein step (b) includes the steps of:
for said each band of noise indicated by each frame of the second subset of the frames, determining a covariance matrix indicative of at least one spatial property of the noise in the band indicated by said frame; generating intermediate frames by applying band-specific gain to the second subset of the frames; and processing the intermediate frames to generate the modified frames, including by applying a band-specific warping matrix to each band of each of the intermediate frames, wherein each said warping matrix is determined by the covariance matrix for noise in the band indicated by one said frame of the second subset of the frames and the at least one target spatial property for said band.
11 . The method of claim 8 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein step (b) includes the steps of:
for said each band of noise indicated by each frame of the second subset of the frames, determining a data structure indicative of at least one spatial property of the noise in the band indicated by said frame; generating intermediate frames by applying band-specific gain to the second subset of the frames; and processing the intermediate frames to generate the modified frames, including by applying a band-specific warping matrix to each band of each of the intermediate frames, wherein each said warping matrix is determined by the data structure for noise in the band indicated by one said frame of the second subset of the frames and the at least one target spatial property for said band.
12 . The method of claim 8 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein step (b) includes the steps of:
for said each band of noise indicated by each frame of the second subset of the frames, determining a covariance matrix indicative of at least one spatial property of the noise in the band indicated by said frame; and processing each frame of the second subset of the frames to generate the modified frames, including by applying a band-specific spectral modification and warping matrix to each band of said each frame of the second subset of the frames, wherein the spectral modification and warping matrix for each band of said each frame is determined by the covariance matrix for the noise in the band indicated by said frame, the at least one target spatial property for said band, and the target spectrum.
13 . The method of claim 8 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein step (b) includes the steps of:
for said each band of noise indicated by each frame of the second subset of the frames, determining a data structure indicative of at least one spatial property of the noise in the band indicated by said frame; and processing each frame of the second subset of the frames to generate the modified frames, including by applying a band-specific spectral modification and warping matrix to each band of said each frame of the second subset of the frames, wherein the spectral modification and warping matrix for each band of said each frame is determined by the data structure for the noise in the band indicated by said frame, the at least one target spatial property for said band, and the target spectrum.
14 . The method of claim 8 , wherein each of the endpoints is a telephone system.
15 . A system configured for use as a teleconferencing system endpoint, including:
a microphone array; a first subsystem coupled and configured to generate noise samples indicative of noise captured during a conference by the microphone array; and a second subsystem coupled and configured to modify the noise samples to generate modified noise samples indicative of modified noise having a frequency-amplitude spectrum which at least substantially matches a target spectrum, and at least one spatial property which at least substantially matches at least one target spatial property.
16 . The system of claim 15 , wherein the noise samples are a first subset of a set of frames of audio samples captured during the conference, each frame of a second subset of the set of frames is indicative of speech uttered by a conference participant, each frame of the first subset is indicative of noise but not a significant level of speech, and the second subsystem is configured to generate the modified noise samples by modifying the first subset of the set of frames.
17 . The system of claim 15 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the noise samples, and wherein the second subsystem is coupled and configured to:
determine, for said each band, a covariance matrix indicative of at least one spatial property of the noise in the band indicated by the noise samples; generate intermediate samples by applying band-specific gain to said each band of the noise samples; and process the intermediate samples to generate the modified noise samples, including by applying a band-specific warping matrix to each band of the intermediate samples, wherein the warping matrix for each band is determined by the covariance matrix for the band and the at least one target spatial property for said band.
18 . The system of claim 15 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the noise samples, and wherein the second subsystem is coupled and configured to:
determine, for said each band, a data structure indicative of at least one spatial property of the noise in the band indicated by the noise samples; generate intermediate samples by applying band-specific gain to said each band of the noise samples; and process the intermediate samples to generate the modified noise samples, including by applying a band-specific warping matrix to each band of the intermediate samples, wherein the warping matrix for each band is determined by the data structure for the band and the at least one target spatial property for said band.
19 . The system of claim 15 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the noise samples, and wherein the second subsystem is coupled and configured to:
determine, for said each band, a covariance matrix indicative of at least one spatial property of the noise in the band indicated by the noise samples; and process the noise samples to generate the modified noise samples, including by applying a band-specific spectral modification and warping matrix to each band of the noise samples, wherein the spectral modification and warping matrix for each band is determined by the covariance matrix for the band, the at least one target spatial property for said band, and the target spectrum.
20 . The system of claim 15 , wherein the at least one target spatial property includes at least one target spatial property for each band of at least one frequency band of the noise samples, and wherein the second subsystem is coupled and configured to:
determine, for said each band, a data structure indicative of at least one spatial property of the noise in the band indicated by the noise samples; and process the noise samples to generate the modified noise samples, including by applying a band-specific spectral modification and warping matrix to each band of the noise samples, wherein the spectral modification and warping matrix for each band is determined by the data structure for the band, the at least one target spatial property for said band, and the target spectrum.
21 . The system of claim 15 , wherein said system is a telephone system including a processor programmed to implement the first subsystem and the second subsystem.
22 . The system of claim 15 , wherein said system is a telephone system including a digital signal processor programmed to implement the first subsystem and the second subsystem.
23 . A teleconferencing system, including:
a link; and at least two endpoints coupled to the link,
wherein each of the endpoints includes:
a microphone array;
a first subsystem coupled and configured to determine a sequence of audio frames indicative of audio captured using the microphone array during a conference, wherein each frame of a first subset of the frames is indicative of speech uttered by a conference participant, each frame of a second subset of the frames is indicative of noise but not a significant level of speech;
a second subsystem coupled and configured to generate modified frames by modifying each frame of the second subset of the frames, such that each of the modified frames is indicative of modified noise having a frequency-amplitude spectrum which at least substantially matches a target spectrum, and a spatial property set which at least substantially matches a target spatial property set; and
an encoding subsystem, coupled and configured to generate encoded audio including by encoding the modified frames and encoding each frame of the first subset of the frames.
24 . The system of claim 23 , also including:
a server coupled to the link, wherein each of the endpoints is configured to transmit the encoded audio generated therein via the link to the server, and the server is configured to generate conference audio indicative of a mix or sequence of audio captured at different ones of the endpoints.
25 . The system of claim 23 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein the second subsystem of each of the endpoints includes:
a subsystem, coupled and configured to determine, for said each band of noise indicated by each frame of the second subset of the frames, a covariance matrix indicative of at least one spatial property of the noise in the band indicated by said frame; a subsystem, coupled and configured to generate intermediate frames by applying band-specific gain to the second subset of the frames; and a subsystem, coupled and configured to process the intermediate frames to generate the modified frames, including by applying a band-specific warping matrix to each band of each of the intermediate frames, wherein each said warping matrix is determined by the covariance matrix for noise in the band indicated by one said frame of the second subset of the frames and the at least one target spatial property for said band.
26 . The system of claim 23 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein the second subsystem of each of the endpoints includes:
a subsystem, coupled and configured to determine, for said each band of noise indicated by each frame of the second subset of the frames, a data structure indicative of at least one spatial property of the noise in the band indicated by said frame; a subsystem, coupled and configured to generate intermediate frames by applying band-specific gain to the second subset of the frames; and a subsystem, coupled and configured to process the intermediate frames to generate the modified frames, including by applying a band-specific warping matrix to each band of each of the intermediate frames, wherein each said warping matrix is determined by the data structure for noise in the band indicated by one said frame of the second subset of the frames and the at least one target spatial property for said band.
27 . The system of claim 23 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein the second subsystem of each of the endpoints includes:
a subsystem, coupled and configured to determine, for said each band of noise indicated by each frame of the second subset of the frames, a covariance matrix indicative of at least one spatial property of the noise in the band indicated by said frame; and a subsystem, coupled and configured to process each frame of the second subset of the frames to generate the modified frames, including by applying a band-specific spectral modification and warping matrix to each band of said each frame of the second subset of the frames, wherein the spectral modification and warping matrix for each band of said each frame is determined by the covariance matrix for the noise in the band indicated by said frame, the at least one target spatial property for said band, and the target spectrum.
28 . The system of claim 23 , wherein the target spatial property set includes at least one target spatial property for each band of at least one frequency band of the noise, and wherein the second subsystem of each of the endpoints includes:
a subsystem, coupled and configured to determine, for said each band of noise indicated by each frame of the second subset of the frames, a data structure indicative of at least one spatial property of the noise in the band indicated by said frame; and a subsystem, coupled and configured to process each frame of the second subset of the frames to generate the modified frames, including by applying a band-specific spectral modification and warping matrix to each band of said each frame of the second subset of the frames, wherein the spectral modification and warping matrix for each band of said each frame is determined by the data structure for the noise in the band indicated by said frame, the at least one target spatial property for said band, and the target spectrum.
29 . The system of claim 23 , wherein each of the endpoints is a telephone system including a processor programmed to implement the first subsystem and the second subsystem.
30 . The system of claim 23 , wherein said system is a telephone system including a digital signal processor programmed to implement the first subsystem and the second subsystem.Join the waitlist — get patent alerts
Track US2014278380A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.