Real-time jitter control and packet-loss concealment in an audio signal
Abstract
An “adaptive audio playback controller” operates by decoding and reading received packets of an audio signal into a signal buffer. Samples of the decoded audio signal are then played out of the signal buffer according to the needs of a player device. Jitter control and packet loss concealment are accomplished by continuously analyzing buffer content in real-time, and determining whether to provide unmodified playback from the buffer contents, whether to compress buffer content, stretch buffer content, or whether to provide for packet loss concealment for overly delayed or lost packets as a function of buffer content. Further, the adaptive audio playback controller also determines where to stretch or compress particular frames or signal segments in the signal buffer, and how much to stretch or compress such segments in order to optimize perceived playback quality.
Claims
exact text as granted — not AI-modified1 - 21 . (canceled)
22 . A method for adaptive playback of received frames of an audio signal transmitted across a packet-based network, comprising using a computing device to:
receive a packetized audio signal broadcast across a packet-based network; decode each received packet and store the resulting decoded signal frame in a signal buffer; output a current packet in the case where the current packet has been received across the packet-based network; instantiate a mute mode whereby a playback of the audio signal is at least partially muted when a maximum delay time for receiving the current packet has been exceeded, and the current packet has not been received; instantiate a packet loss concealment mode whereby the playback of the audio signal is modified for reducing audible artifacts resulting from one or more lost packets when a current buffer content has been previously temporally stretched, the current packet has not yet been received, and a packet subsequent to the current packet has already been received.
23 . The method of claim 22 further comprising analyzing the content of the signal buffer for determining a current length of the contents of the signal buffer.
24 . The method of claim 23 further comprising stretching and outputting one or more decoded frames from the signal buffer when the current length of the contents of the signal buffer is less than a predetermined minimum buffer size.
25 . The method of claim 24 wherein the predetermined minimum buffer size is optimized to compensate for clock drift between an encoder and a decoder.
26 . The method of claim 23 further comprising compressing and outputting one or more decoded frames from the signal buffer when the current length of the contents of the signal buffer is greater than a predetermined maximum buffer size.
27 . The method of claim 24 wherein the predetermined maximum buffer size is optimized to compensate for clock drift between an encoder and a decoder.
28 . The method of claim 22 wherein modification of the playback of the audio signal is in the packet loss concealment mode comprises:
computing an average energy for a frame in the signal buffer immediately preceding the current packet that has not yet been received; computing an average energy for a frame in the signal buffer immediately succeeding the current packet that has not yet been received; and determining a target frame size for both the preceding and succeeding frames as a function of the ratio of the of the average energy of the succeeding frame to the preceding frame.
29 . The method of claim 28 wherein determining a target frame size for both the preceding and succeeding frames further comprises stretching the succeeding frame and the preceding frames by an amount that is inversely proportional to the ratio of the average energy.
30 . The method of claim 29 wherein instantiating the mute mode comprises generating and providing playback of a comfort noise signal to replace lost packets, said comfort noise signal being generated from at least one signal frame stored in a silence buffer, said signal frame having been determined to represent nominal background noise.
31 . The method of claim 30 further comprising periodically replacing the signal frames in the silence buffer as a function of a computed energy of those frames.
32 . The method of claim 30 wherein generating the comfort noise signal from the at least one signal frame stored in a silence buffer comprises:
automatically computing the FFT of the at least one signal frame stored in the silence buffer; introducing a random rotation of the phase into the FFT coefficients; computing the inverse FFT for each segment, thereby creating the at least one synthetic silence segment; and providing the at least one silence segment for playback as the comfort noise signal.
33 - 50 . (canceled)
51 . A physical computer-readable media having computer executable instructions stored thereon for adaptive playback of received frames of an audio signal transmitted across a packet-based network, said computer-executable instructions comprising:
receiving a packetized audio signal broadcast across a packet-based network; decoding each received packet and store the resulting decoded signal frame in a signal buffer; outputting a current packet in the case where the current packet has been received across the packet-based network; instantiating a mute mode whereby a playback of the audio signal is at least partially muted when a maximum delay time for receiving the current packet has been exceeded, and the current packet has not been received; and instantiating a packet loss concealment mode whereby the playback of the audio signal is modified for reducing audible artifacts resulting from one or more lost packets when a current buffer content has been previously temporally stretched, the current packet has not yet been received, and a packet subsequent to the current packet has already been received.
52 . The computer-readable media of claim 51 further comprising instructions for analyzing the content of the signal buffer for determining a current length of the contents of the signal buffer.
53 . The computer-readable media of claim 51 further comprising instructions for stretching and outputting one or more decoded frames from the signal buffer when the current length of the contents of the signal buffer is less than a predetermined minimum buffer size.
54 . The computer-readable media of claim 51 further comprising instructions for compressing and outputting one or more decoded frames from the signal buffer when the current length of the contents of the signal buffer is greater than a predetermined maximum buffer size.
55 . A system for providing adaptive playback of received frames of an audio signal transmitted across a packet-based network, comprising using a computing device for:
receiving a packetized audio signal broadcast across a packet-based network; decoding each received packet and store the resulting decoded signal frame in a signal buffer; outputting a current packet in the case where the current packet has been received across the packet-based network; instantiating a mute mode whereby a playback of the audio signal is at least partially muted when a maximum delay time for receiving the current packet has been exceeded, and the current packet has not been received; and instantiating a packet loss concealment mode whereby the playback of the audio signal is modified for reducing audible artifacts resulting from one or more lost packets when a current buffer content has been previously temporally stretched, the current packet has not yet been received, and a packet subsequent to the current packet has already been received.
56 . The system claim 55 further comprising analyzing the content of the signal buffer for determining a current length of the contents of the signal buffer.
57 . The system claim 55 further comprising stretching and outputting one or more decoded frames from the signal buffer when the current length of the contents of the signal buffer is less than a predetermined minimum buffer size.
58 . The system of claim 57 wherein the predetermined minimum buffer size is optimized to compensate for clock drift between an encoder and a decoder.
59 . The system claim 55 further comprising compressing and outputting one or more decoded frames from the signal buffer when the current length of the contents of the signal buffer is greater than a predetermined maximum buffer size.Join the waitlist — get patent alerts
Track US2009304032A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.