Single-sided speech quality measurement
Abstract
A non-intrusive speech quality estimation technique is based on statistical or probability models such as Gaussian Mixture Models (“GMMs”). Perceptual features are extracted from the received speech signal and assessed by an artificial reference model formed using statistical models. The models characterize the statistical behavior of speech features. Consistency measures between the input speech features and the models are calculated to form indicators of speech quality. The consistency values are mapped to a speech quality score using a mapping optimized using machine learning algorithms, such as Multivariate Adaptive Regression Splines (“MARS”). The technique provides competitive or better quality estimates relative to known techniques while having lower computational complexity.
Claims
exact text as granted — not AI-modified1 . A single-ended speech quality measurement method comprising the steps of:
extracting perceptual features from a speech signal; assessing the perceptual features with statistical models to form indicators of speech quality; and employing the indicators of speech quality to produce a speech quality score.
2 . The method of claim 1 including the further step of extracting the perceptual features from the received speech signal frame-by-frame.
3 . The method of claim 2 further including the step of classifying the frames by signal contents.
4 . The method of claim 3 including the further step of separately modeling the probability distribution of the features for each frame class and different classes of speech signals with statistical models.
5 . The method of claim 4 including the further step of calculating a consistency measure indicative of speech quality for each class separately with a plurality of statistical models.
6 . The method of claim 5 including the further step of employing the consistency measures to obtain an estimate of subjective scores.
7 . The method of claim 6 including the further step of mapping the consistency measures to a speech quality score using a mapping, such as Multivariate Adaptive Regression Splines.
8 . The method of claim 1 wherein the perceptual features are assessed with Gaussian Mixture Models to form indicators of speech quality.
9 . The method of claim 4 wherein the classes include voiced, unvoiced, and inactive.
10 . Apparatus operable to provide a single-ended speech quality measurement, comprising:
a feature extraction module operable to extract perceptual features from a received speech signal; a statistical reference model and consistency calculation module operable in response to output from the feature extraction module to assess the perceptual features to form indicators of speech quality; and a scoring module operable to employ the indicators of speech quality to produce a speech quality score.
11 . The apparatus of claim 10 wherein the feature extraction module is further operable to extract the perceptual features from the received speech signal frame-by-frame.
12 . The apparatus of claim 11 further including a time segmentation module operable to classify the frames by signal contents.
13 . The apparatus of claim 12 wherein the consistency calculation module is further operable to separately model the probability distribution of the features for each class and different classes of speech signals with the statistical models.
14 . The apparatus of claim 13 wherein the consistency calculation module is further operable to calculate a consistency measure indicative of speech quality for each class separately with a plurality of Gaussian Mixture Models.
15 . The apparatus of claim 14 further including a mapping module operable to employ the consistency measures to obtain an estimate of subjective scores.
16 . The apparatus of claim 15 wherein the mapping module employs a mapping, such as one optimized using Multivariate Adaptive Regression Splines.
17 . The apparatus of claim 10 wherein the statistical reference model includes Gaussian Mixture Models.
18 . The apparatus of claim 13 wherein the classes include voiced, unvoiced, and inactive.Join the waitlist — get patent alerts
Track US2007203694A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.