Digital speech processor using arbitrary excitation coding
Abstract
An arrangement for processing a speech message which uses arbitrary value codes to form time frame excitation signals. The arbitrary value codes, e.g., random numbers, are stored as well as signals indexing the codes and transform domain signals corresponding to the arbitrary codes are generated. The speech message is partitioned into time frame interval speech patterns and a first signal representative of the transform domain speech pattern of each successive time frame interval is formed responsive to the partitioned speech message. A plurality of second signals representative of time frame interval patterns corresponding to the transform code signals are generated responsive to said set of transform signals. One of the arbitrary code signals is selected jointly responsive to the first and second signals of each successive time interval to represent the time frame speech signal excitation, and the index signal corresponding to said selected arbitrary code signal is output. A replica of the speech message is formed from the arbitrary codes by concatenating a sequence of said arbitrary codes identified by the output index signals.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. Apparatus for encoding speech comprising means (330) for storing a set of signals each representative of a random code and a set of index signals each identifying one of the random codes; means (203 through 247 except 225 and 245) for partitioning the speech into successive time frame interval portions and for forming a time-domain signal representative of the portion of speech in each successive time frame interval; means (225, 245, 250) for generating at least one transform domain signal from each such time-domain signal; means (305) responsive to each random code signal for generating a transform domain code signal corresponding thereto, via the same type of transformation as in the aforesaid means for generating a transform domain signal; means (315 and 320, or 501 through 520 and 320) for cross-correlating transform domain signals for each time frame interval with each of said transform domain code signals to select one of the transform domain code signals as yielding minimum error or maximum similarity as a representative of the speech portion in the time-frame interval; and means (325) for outputting the index signal corresponding to the random code signal corresponding to the selected transform domain code signal.
2. Apparatus for encoding speech of the type claimed in claim 1 in which the means for forming a time domain signal comprises means for forming said signal as representative of the predictive parameters of the portion of speech in each successive time frame interval; the means for generating at least one trnsform domain signal comprises means for generating a transform domain signal representative of the predictive parameters from said time domain signal representative of the predictive parameters; and the means for generating at least one transform domain signal further comprises means (225, 245) for generating a transform domain signal representative of predictive characteristics for said portion of speech; the means for cross-correlating includes means responsive to the predictive characteristics representative signal for forming a signal (γ) representative of the relative scaling of the transform domain code signal with respect to a transform domain signal representative of the predictive parameters for each time frame interval; and the outputting means comprises means for outputting the relative scaling signal and the signal representative of the predictive parameters.
3. Apparatus for encoding speech of the type claimed in claim 2, in which the means for forming a time domain signal as representative of the portion of speech in each successive time frame interval comprises means (209, 213, 215) for generating a set of signals representative of the predictive parameters of the speech in each successive time frame interval; means (207, 211) for forming a signal representative of the predictive residual for the speech in each successive time frame interval; and means (217, 227, 222, 235, 240, 247) responsive to the predictive residual generating means and to the predictive parameter signal generating means for removing the contribution attributable to speech from the previous time frame.
4. Apparatus for encoding speech of the type claimed in claim 3, in which the means for partitioning and forming a time domain signal, further includes means (220, 230), responsive to the predictive residual generating means, for producing pitch predictive parameters including contributions of previous frames; and the combining means of the outputting means is responsive to said means for producing pitch predictive parameters.
5. Apparatus for encoding speech of the type claimed in either of claims 2 or 3 in which the cross-correlating means comprises means (501) for cross-correlating all three of said predictive-parameter-representative transform domain signal, said transform domain signal representative of the relative scaling for the portion of speech, and said transform domain code signal; means (505, 510, 515, 520) responsive to the output of the means for cross-correlating specifically and to one or more of the three signals for producing the relative scaling signal (γ) and for producing a cross-correlation error signal (E.sub.(k)).
6. Apparatus for encoding speech comprising means (330) for storing a set of signals each representative of a random code and set of index signals each identifying one of the random codes; means (203 through 247 except 225 and 245) for partitioning the speech into successive time frame interval portions and for forming a time-domain signal representative of the portion of speech in each successive time frame interval; means (225, 245, 250) for generating at least one transform domain signal from each such time-domain signal; means (305) responsive to each random code signal for generating a transform domain code signal corresponding thereto, via the same type of transformation as in the aforesaid means for generating a transform domain signal; means (315 and 320 or 501 through 520 and 320) for responding in a comparative fashion to transform domain signals for each time frame interval and, for each such signal, to each of said transform domain code signals to select one of the transform domain code signals as yielding minimum error or maximum similarity as a representative of the speech portion in the time frame interval; and means (325) for outputting the index signal corresponding to the random code signal corresponding to the selected transform domain code signal.
7. A method for encoding speech comprising the steps of storing a set of signals each representative of a random code and a set of index signals each identifying one of the random codes; partitioning the speech into successive time frame interval portions; forming a time-domain signal representative of the portion of speech in each successive time frame interval; generating at least one transform domain signal from each such time-domain signal; generating a transform domain code signal responsive to each random code signal, via the same type of transformation as in the aforesaid steps of generating a transform domain signal; cross-correlating transform domain signals for each time frame interval with each of said transform domain code signals to select one of the transform domain code signals as yielding minimum error or maximum similarly as a representative of the speech portion in the time-frame interval; and outputting the index signal corresponding to the random code signal corresponding to the selected transform domain code signal.
8. A method for encoding speech of the type claimed in claim 7 in which the step of forming a time domain signal comprises the step of forming said signal as representative of the predictive parameters of the portion of speech in each successive time frame interval; the step of generating at least one transform domain signal comprises generating a transform domain signal representative of the predictive parameter from said time domain signal representative of the predictive parameters; and the step of generating at least one transform domain signal further comprises step of generating a transform domain signal representative of predictive characteristics for said portion of speech; the step of cross-correlating includes the step of forming a signal (γ) representative of the relative scaling of the transform domain code signal with respect to a transform domain signal representative of the predictive parameters for each time frame interval in response to the representative signal representative of the energy predictive characteristics; and the outputting means comprises means for outputting the relative scaling signal and the signal representative of the predictive parameters.
9. A method for encoding speech of the type claimed in claim 8, in which the step of forming a time domain signal as representative of the pattern of the portion of speech in each successive time frame interval comprises generating a set of signals representative of the predictive parameters of the speech in each successive time frame interval; forming a signal representative of the predictive residual for the speech in each successive time frame interval; and removing the contribution attributable to speech from the previous time frame in response to the predictive residual generating means and to the predictive parameter signal generating means.
10. Apparatus for encoding speech of the type claimed in claim 9, in which the partitioning step and the step of forming a time domain signal includes producing pitch predictive parameters including contributions of previous frames in response to the predictive residual representative signal; and the combining step also combines said pitch predictive parameters.
11. A method for encoding speech of the type claimed in either of claims 8 or 9 in which the cross-correlating step comprises specifically cross-correlating all three of said predictive-parameter-representative transform domain signal, said transform domain signal representative of the relative scaling for the portion of speech, and said transform domain code signal; applying the output of the specifically cross-correlating step and one or more of the three signals to produce the relative scaling signal (γ) and a cross-correlation error signal (E.sub.(k)).
12. A method for encoding speech comprising storing a set of signals each representative of a random code and a set of index signals each identifying one of the random codes; partitioning the speech into successive time frame interval portions; forming a time-domain signal representative of the portion of speech in each successive time frame interval; generating at least one transform domain signal from each such time-domain signal; generating a transform domain code signal responsive to each random code signal via the same type of transformation as in the aforesaid step of generating a transform domain signal; responding in a comparative fashion to transform domain signals for each time frame interval and, for each such signal, to each of said transform domain code signals to select one of the transform domain code signals as yielding minimum error or maximum similarity as a representative of the speech portion in the time frame interval; and outputting the index signal corresponding to the random code signal corresponding to the selected transform.
13. Apparatus for providing a speech message comprising means for receiving a sequence of speech message signals for the successive time intervals of the speech message, each time interval speech message signal including a set of transform-domain-coded signals representative of the time interval portion of the speech message, at least a portion of which are index signals corresponding to a known set of random codes means for storing said known set of random codes in one-for-one association with the corresponding index signals means for generating said random codes for each of the set of index signals, and means for controlling speech wave generation for said time interval in response to said generated random codes.
14. Apparatus of the type claimed in claim 13 in which the storing means comprises means for storing the random codes sequentially so that a first portion of each succeeding one is derived from the latter portion of the preceding one.
15. A method for producing a speech message comprising receiving a sequence of speech message signals for the successive time intervals of the speech message, each time interval speech message signal including a set of transform-domain-coded signals representative of the time interval portion of the speech message, at least a portion of which are index signals corresponding to a known set of random codes; storing said known set of random codes in one-for-one association with the corresponding index signals; generating said codes sequentially for each of the set of index signals; and controlling speech wave generation for said time interval in response to said sequentially generated random codes. .Iadd.
16. Apparatus for processing input speech signals occurring over one or more successive time frame intervals comprising means (e.g., 203, 205) for forming first signals representative of said speech signals for each of said time frame intervals, means (e.g., 305) for providing a set of transform domain representations of a corresponding set of codes each of said codes being representative of possible speech signals over a time frame interval, and for providing a set of index signals, each identifying one of said codes, means (e.g., 315) for generating a set of similarity signals corresponding to the similarity of said first signals to each of said set of transform domain representations, means (e.g., 320) for generating index signals corresponding to codes having transform domain representations giving rise to similarity signals meeting a predetermined criterion. .Iaddend. .Iadd.17. Apparatus according to claim 16, wherein said means for forming first signals comprises means (e.g., 215) for forming a perceptually weighted representation of said input speech signals. .Iaddend. .Iadd.18. Apparatus according to claim 16, wherein said means for providing a set of transform domain representations comprises memory means (e.g., 305) for storing a set of transform domain signals. .Iaddend. .Iadd.19. Apparatus according to claim 16, wherein said means for forming first signals comprises means (e.g., 207, 209, 211, 215, 217, 222, 227, 240 and 247) for reducing for each time frame interval any contribution to said first signal arising from speech signals occurring during another time frame interval. .Iaddend. .Iadd.20. The method of processing input speech signals occurring over one or more successive time frame intervals comprising forming first signals representative of said speech signals for each of said time frame intervals, providing a set of transform domain representations of a corresponding set of codes each of said codes being representative of possible speech signals over a time frame interval, and for providing a set of index signals, each identifying one of said codes, generating a set of similarity signals corresponding to the similarity of said first signals to each of said set of transform domain representations, generating index signals corresponding to codes having transform domain representations giving rise to similarity signals meeting a predetermined
criterion. .Iaddend. .Iadd.21. The method according to claim 20, wherein said step of forming first signals comprises forming a perceptually weighted representation of said input speech signals. .Iaddend. .Iadd.22. The method according to claim 20, wherein said step of providing a set of transform domain representations comprises accessing a set of stored transform domain signals. .Iaddend. .Iadd.23. The method according to claim 20, wherein said forming first signals comprises reducing for each time frame interval any contribution to said first signal arising from speech signals occurring during another time frame interval. .Iaddend.Join the waitlist — get patent alerts
Track USRE34247E — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.