Speaker recognition from telephone calls
Abstract
The present invention relates to a method for speaker recognition, comprising the steps of obtaining and storing speaker information for at least one target speaker; obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker; classifying the speech samples according to at least one unknown speaker thereby providing speaker-dependent classes of speech samples; extracting speaker information for the speech samples of each of the speaker-dependent classes of speech samples; combining the extracted speaker information for each of the speaker-dependent classes of speech samples; comparing the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for at least one target speaker to obtain at least one comparison result; and determining whether at least one unknown speaker is identical with at least one target speaker based on at least one comparison result.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for speaker recognition, comprising the steps of
obtaining and storing, in a database on a computer, speaker information for at least one target speaker; obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker; classifying, using software stored and operating on the computer, the speech samples according to the at least one unknown speaker thereby providing one, two or more speaker-dependent classes of speech samples; extracting, using software stored and operating on the computer, speaker information for the speech samples of each of the speaker-dependent classes of speech samples; combining, using software stored and operating on the computer, the extracted speaker information for each of the speaker-dependent classes of speech samples; comparing, using software stored and operating on the computer, the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and determining, using software stored and operating on the computer, whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.
2 . The method of claim 1 , further comprising grouping of the telephone calls according to the telephone numbers of the telephone calls.
3 . The method of claim 2 , wherein the speaker information for the at least one target speaker are obtained by obtaining a plurality of speech samples of the at least one target speaker.
4 . The method of claim 3 , wherein at least one of the plurality of speech samples of the at least one target speaker is obtained from a telephone call of the at least one target speaker.
5 . The method of claim 4 , wherein the speech samples according to the at least one unknown speaker are classified by a speaker clustering technique, in particular, by Agglomerative Hierarchical Clustering.
6 . The method of claim 5 , wherein the speaker clustering technique is based on a Gaussian Mixture Model and a Gaussian Mixture Model metric.
7 . The method of claim 6 , wherein the speaker clustering technique employs a Joint Factor Analysis.
8 . The method of claim 1 , wherein combining the extracted speaker information for each of the speaker-dependent classes of speech samples comprises generating for a particular class a combined Gaussian Mixture Model from the extracted speaker information of the speech samples of that class.
9 . The method of claim 6 , wherein the combined Gaussian Mixture Model is generated from Gaussian Mixture Models of the speech samples of that class.
10 . The method of claim 1 , wherein combining the extracted speaker information for each of the speaker-dependent classes of speech samples comprises combining feature vectors obtained for one or more speech samples of a speaker-dependent class with feature vectors of one or more other speech samples of the same speaker-dependent class, in particular, by summation of at least some of the feature vectors, more particularly, comprising adding a feature vector of one speech sample of the speaker-dependent class and another feature vector of another speech sample of the speaker-dependent class, if they are close to each other within predetermined limits.
11 . A computer program product, comprising one or more computer readable media having computer-executable instructions for performing steps of the method according to one of the preceding claims when run on a computer.
12 . A system for performing speaker recognition, comprising: a database stored and operating on a computer and configured to store speaker
information for a target speaker; software means stored and operating on the computer and configured to classify speech samples of telephone calls according to at least one unknown speaker thereby providing one, two or more speaker-dependent classes of speech samples; software means stored and operating on the computer and configured to extract speaker information for the speech samples of each of the speaker-dependent classes of speech samples; software means stored and operating on the computer and configured to combine the extracted speaker information for each of the speaker-dependent classes of speech samples; software means stored and operating on the computer and configured to compare the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and software means stored and operating on the computer and configured to determine whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.
13 . The system of claim 12 , further comprising software means stored and operating on the computer and configured to receive telephone calls from at least one unknown speaker.
14 . The system of claim 12 , further comprising software means stored and operating on the computer and configured to group the telephone calls according to the telephone numbers of the telephone calls.Join the waitlist — get patent alerts
Track US2016019897A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.