US9912617B2ActiveUtilityA1

Method and apparatus for voice communication based on voice activity detection

Assignee: DOLBY LABORATORIES LICENSING CORPPriority: Mar 23, 2012Filed: Jan 4, 2017Granted: Mar 6, 2018
Est. expiryMar 23, 2032(~5.7 yrs left)· nominal 20-yr term from priority
H04L 65/1066H04L 49/90G10L 19/167H04M 3/569G10L 25/78H04L 49/9023
48
PatentIndex Score
0
Cited by
25
References
6
Claims

Abstract

Voice communication method and apparatus and method and apparatus for operating jitter buffer are described. Audio blocks are acquired in sequence. Each of the audio blocks includes one or more audio frames. Voice activity detection is performed on the audio blocks. In response to deciding voice onset for a present one of the audio blocks, a subsequence of the sequence of the acquired audio blocks is retrieved. The subsequence precedes the present audio block immediately. The subsequence has a predetermined length and non-voice is decided for each audio block in the subsequence. The present audio block and the audio blocks in the subsequence are transmitted to a receiving party. The audio blocks in the subsequence are identified as reprocessed audio blocks. In response to deciding non-voice for the present audio block, the present audio block is cached.

Claims

exact text as granted — not AI-modified
We claim: 
     
       1. A method of operating at least one jitter buffer having the same length in voice communication to improve voice quality, comprising:
 receiving at least one audio block from two or more transmitting parties; 
 for each of the at least one received audio block,
 in response to receiving an audio block identified as a reprocessed audio block providing notification that these audio blocks are different from the present audio block and reprocessed as including voice,
 determining whether the received audio block is timed out relative to a time range of a jitter buffer associated with the transmitting party; 
 in response to determining that the received audio block is timed out,
 in response to determining that the received audio block corresponds to the same time as the entry moved out from the corresponding jitter buffer but waiting for mixing, updating the entry with the received audio block; and 
 in response to determining that the received audio block does not correspond to the same time as the entry moved out from the corresponding jitter buffer but waiting for mixing, transmitting the received audio block to a destination 
 
 in response to determining that the received audio block is not timed out, updating an entry corresponding to the same time as the received audio block in the jitter buffer, with the received audio block; and 
 
 in response to receiving an audio block not identified as a reprocessed audio block, filling the received audio block into an entry corresponding to the same time as the received audio block in the jitter buffer associated with the transmitting party. 
 
 
     
     
       2. The method according to  claim 1 , wherein the updating of the entry with the received audio block comprises:
 in response to determining that the entry includes an audio block, replacing the included audio block with the received audio block; and 
 in response to determining that the entry is empty, filling the received audio block into the entry. 
 
     
     
       3. The method according to  claim 1 , wherein the updating of the entry with the received audio block comprises:
 in response to determining that the entry includes an audio block generated by a source different from that of the received audio block, mixing the received audio block and the included audio block into one audio block and replacing the included audio block with the mixed audio block; 
 in response to determining that the entry includes an audio block generated by the same source as the received audio block, replacing the included audio block with the received audio block; and 
 in response to determining that the entry is empty, filling the received audio block into the entry, and 
 the method further comprises: 
 feeding a voice processor with audio blocks in the jitter buffer corresponding to the voice processor. 
 
     
     
       4. An apparatus for use in voice communication, comprising:
 at least one jitter buffer having the same length, configured to temporarily storing audio blocks received in the voice communication; 
 a receiver configured to receive at least one audio block from two or more transmitting parties; 
 a transmitter; and 
 a jitter buffer controller comprising at least one processor and one memory configured to, for each of the at least one received audio block,
 in response to receiving an audio block identified as a reprocessed audio block providing notification that these audio blocks are different from the present audio block and reprocessed as including voice,
 determine whether the received audio block is timed out relative to a time range of a jitter buffer associated with the transmitting party; 
 in response to determining that the received audio block is timed out,
 in response to determining that the received audio block corresponds to the same time as the entry moved out from the corresponding jitter buffer but waiting for mixing, update the entry with the received audio block; and 
 in response to determining that the received audio block does not correspond to the same time as the entry moved out from the corresponding jitter buffer but waiting for mixing, transmit the received audio block to a destination through the transmitter, 
 
 in response to determining that the received audio block is not timed out, update an entry corresponding to the same time as the received audio block in the jitter buffer, with the received audio block; and 
 
 in response to receiving an audio block not identified as a reprocessed audio block, fill the received audio block into an entry corresponding to the same time as the received audio block in the jitter buffer associated with the transmitting party. 
 
 
     
     
       5. The apparatus according to  claim 4 , wherein the updating of the entry with the received audio block comprises:
 in response to determining that the entry includes an audio block, replacing the included audio block with the received audio block; and 
 in response to determining that the entry is empty, filling the received audio block into the entry. 
 
     
     
       6. The apparatus according to  claim 4 , further comprising:
 a voice processor fed with audio blocks in the corresponding jitter buffer, and 
 the updating of the entry with the received audio block comprises: 
 in response to determining that the entry includes an audio block generated by a source different from that of the received audio block, mixing the received audio block and the included audio block into one audio block and replacing the included audio block with the mixed audio block; 
 in response to determining that the entry includes an audio block generated by the same source as the received audio block, replacing the included audio block with the received audio block; and 
 in response to determining that the entry is empty, filling the received audio block into the entry.

Join the waitlist — get patent alerts

Track US9912617B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.