Method and apparatus for evaluating speech quality
Abstract
In an embodiment a method for evaluating speech quality includes receiving, by a computing device, synthesized speech comprising one or more frames, determining, by the computing device, a latent representation corresponding to each frame, clustering, by the computing device, each latent representation and then mapping a center point of each cluster to an index to determine a centroid index sequence, determining, by the computing device, an embedding sequence by replacing each index of the centroid index sequence with an embedding corresponding to each index, and then masking one or more of embeddings, determining, by the computing device, a predicted index sequence that reconstructs a masked embedding based on a masked embedding sequence, and determining, by the computing device, an evaluation score based on a difference between the centroid index sequence and the predicted index sequence.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for evaluating speech quality, the method comprising:
receiving, by a computing device, synthesized speech comprising one or more frames; determining, by the computing device, a latent representation corresponding to each frame; clustering, by the computing device, each latent representation and then mapping a center point of each cluster to an index to determine a centroid index sequence; determining, by the computing device, an embedding sequence by replacing each index of the centroid index sequence with an embedding corresponding to each index, and then masking one or more of embeddings; determining, by the computing device, a predicted index sequence that reconstructs a masked embedding based on a masked embedding sequence; and determining, by the computing device, an evaluation score based on a difference between the centroid index sequence and the predicted index sequence.
2 . The method of claim 1 , wherein the latent representation is extracted using a machine learning model that has pre-trained referenced speech using a self-supervised learning method.
3 . The method of claim 1 , wherein clustering uses a k-means clustering algorithm.
4 . The method of claim 1 , wherein masking is performed by randomly selecting one embedding from the embedding sequences and replacing five consecutive embeddings starting from the selected embedding with the one embedding.
5 . The method of claim 1 , wherein the predicted index sequence is estimated using a machine learning model that has pre-trained a feature distribution of referenced speech using an unsupervised learning method.
6 . The method of claim 1 , wherein determining the evaluation score comprises determining a loss value between the centroid index sequence and the predicted index sequence using a cross-entropy loss function.
7 . The method of claim 1 , further comprising determining an integrated score that evaluates multidimensional speech quality by integrating a plurality of evaluation scores determined based on latent representations corresponding to different acoustic features.
8 . The method of claim 1 , wherein determining the latent representation corresponding to each frame comprises extracting acoustic information from a sixth layer and linguistic information from a twenty-third layer of a transformer encoder.
9 . The method of claim 1 , wherein clustering each latent representation comprises discretizing the latent representation into 1024 clusters.
10 . A method for training a speech quality evaluation method, the method comprising:
receiving, by a computing device, referenced speech comprising one or more frames; determining, by the computing device, a latent representation corresponding to each frame; clustering, by the computing device, each latent representation, and then mapping a center point of each cluster to an index to determine a centroid index sequence; determining, by the computing device, an embedding sequence by replacing each index of the centroid index sequence with an embedding corresponding to each index, and then masking one or more of the embeddings; determining, by the computing device, a predicted index sequence that reconstructs a masked embedding based on a masked embedding sequence, using a machine learning model that has pre-trained a feature distribution of the referenced speech using an unsupervised learning method; determining, by the computing device, a loss value between the centroid index sequence and the predicted index sequence; and updating, by the computing device, a parameter of the machine learning model based on the loss value.
11 . An apparatus comprising:
one or more processors; and at least one memory storing a program including program instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving synthesized speech comprising one or more frames; determining a latent representation corresponding to each frame; clustering each latent representation, and then mapping a center point of each cluster to an index to determine a centroid index sequence; determining an embedding sequence by replacing each index of the centroid index sequence with an embedding corresponding to each index, and then masking one or more of embeddings; determining a predicted index sequence that reconstructs a masked embedding based on a masked embedding sequence; and determining an evaluation score based on a difference between the centroid index sequence and the predicted index sequence.
12 . The apparatus of claim 11 , wherein the latent representation is extracted using a machine learning model that has pre-trained referenced speech using a self-supervised learning method.
13 . The apparatus of claim 11 , wherein clustering uses a k-means clustering algorithm.
14 . The apparatus of claim 11 , wherein masking is performed by randomly selecting one embedding from the embedding sequences and replacing five consecutive embeddings starting from the selected embedding with the one embedding.
15 . The apparatus of claim 11 , wherein the predicted index sequence is estimated using a machine learning model that has pre-trained a feature distribution of referenced speech using an unsupervised learning method.
16 . The apparatus of claim 11 , wherein determining the evaluation score comprises determining a loss value between the centroid index sequence and the predicted index sequence using a cross-entropy loss function.
17 . The apparatus of claim 11 , wherein the operations further comprises determining an integrated score that evaluates multidimensional speech quality by integrating a plurality of evaluation scores determined based on latent representations corresponding to different acoustic features.
18 . The apparatus of claim 11 , wherein determining the latent representation corresponding to each frame comprises extracting acoustic information from a sixth layer and linguistic information from a twenty-third layer of a transformer encoder.
19 . The apparatus of claim 11 , wherein clustering each latent representation comprises discretizing the latent representation into 1024 clusters.
20 . An apparatus comprising:
one or more processors; and at least one memory storing a program including program instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving referenced speech including one or more frames; determining a latent representation corresponding to each frame; clustering each latent representation, and then mapping a center point of each cluster to an index to determine a centroid index sequence; determining an embedding sequence by replacing each index of the centroid index sequence with an embedding corresponding to each index, and then masking one or more of the embeddings; determining a predicted index sequence that reconstructs a masked embedding based on a masked embedding sequence, using a machine learning model that has pre-trained a feature distribution of the referenced speech using an unsupervised learning method; determining a loss value between the centroid index sequence and the predicted index sequence; and updating a parameter of the machine learning model based on the loss value.Join the waitlist — get patent alerts
Track US2026065899A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.