Speaker recognition system and method of using the same
Abstract
A speaker recognition system includes a non-transitory computer readable medium configured to store instructions. The speaker recognition system further includes a processor connected to the non-transitory computer readable medium. The processor is configured to execute the instructions for extracting acoustic features from each frame of a plurality of frames in input speech data. The processor is configured to execute the instructions for calculating a saliency value for each frame of the plurality of frames using a first neural network (NN) based on the extracted acoustic features, wherein the first NN is a trained NN using speaker posteriors. The processor is configured to execute the instructions for extracting a speaker feature using the saliency value for each frame of the plurality of frames.
Claims
exact text as granted — not AI-modified1 . A speaker recognition system comprising:
a non-transitory computer readable medium configured to store instructions; and a processor connected to the non-transitory computer readable medium, wherein the processor is configured to execute the instructions for:
extracting acoustic features from each frame of a plurality of frames in input speech data; and
extracting a speaker feature using an output of a first neural network (NN), wherein the output of the first NN is obtained by inputting the extracted acoustic features into the first NN, and the first NN is a trained NN using speaker posteriors.
2 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for extracting the speaker feature using a weighted pooling process implemented using the saliency value for each frame of the plurality of frames.
3 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for training the first NN using the speaker posteriors.
4 . The speaker recognition system of claim 3 , wherein the processor is configured to execute the instructions for generating the speaker posteriors using training data and speaker identification information.
5 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for calculating the saliency value for each frame of the plurality of frames based on a gradient of the speaker posterior for each frame of the plurality of frames based on the extracted acoustic features.
6 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for calculating the saliency value for each frame of the plurality of frames using a first node of the first NN and a second node of the first NN, wherein a first frame of the plurality of frames output at the first node indicates the first frame has more useful information than a second frame of the plurality of frames output at the second node.
7 . The speaker recognition system of claim 6 , wherein the processor is configured to execute the instructions for calculating the saliency value for each frame of the plurality of frames based on a gradient of the speaker posterior for each frame of the plurality of frames output at the first node of the first NN based on the extracted acoustic features.
8 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for outputting an identity of a speaker of the input speech data based on the extracted speaker feature.
9 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for matching a speaker of the input speech data to a stored speaker identification based on the extracted speaker feature.
10 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for permitting access to a computer system in response to the extracted speaker feature matching an authorized user.
11 . A speaker recognition method comprising:
receiving input speech data; extracting acoustic features from each frame of a plurality of frames in the input speech data; and extracting a speaker feature using an output of a first neural network (NN), wherein the output of the first NN is obtained by inputting the extracted acoustic features into the first NN, and the first NN is a trained NN using speaker posteriors.
12 .- 20 . (canceled)
21 . The speaker recognition system of claim 1 , wherein the processor is configured to execute the instructions for calculating a saliency value for each frame of the plurality of frames using the output of the first NN and extracting the speaker feature using the saliency value for each frame.
22 . The speaker recognition system of claim 1 , wherein the first NN is a trained NN using speaker posteriors categorized based on a predetermined threshold value.Join the waitlist — get patent alerts
Track US2022130397A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.