US2016019897A1PendingUtilityA1

Speaker recognition from telephone calls

Assignee: AGNITIO SLPriority: Nov 12, 2009Filed: May 22, 2015Published: Jan 21, 2016
Est. expiryNov 12, 2029(~3.3 yrs left)· nominal 20-yr term from priority
G10L 17/02
42
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention relates to a method for speaker recognition, comprising the steps of obtaining and storing speaker information for at least one target speaker; obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker; classifying the speech samples according to at least one unknown speaker thereby providing speaker-dependent classes of speech samples; extracting speaker information for the speech samples of each of the speaker-dependent classes of speech samples; combining the extracted speaker information for each of the speaker-dependent classes of speech samples; comparing the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for at least one target speaker to obtain at least one comparison result; and determining whether at least one unknown speaker is identical with at least one target speaker based on at least one comparison result.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for speaker recognition, comprising the steps of
 obtaining and storing, in a database on a computer, speaker information for at least one target speaker;   obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker;   classifying, using software stored and operating on the computer, the speech samples according to the at least one unknown speaker thereby providing one, two or more speaker-dependent classes of speech samples;   extracting, using software stored and operating on the computer, speaker information for the speech samples of each of the speaker-dependent classes of speech samples;   combining, using software stored and operating on the computer, the extracted speaker information for each of the speaker-dependent classes of speech samples;   comparing, using software stored and operating on the computer, the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and   determining, using software stored and operating on the computer, whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.   
     
     
         2 . The method of  claim 1 , further comprising grouping of the telephone calls according to the telephone numbers of the telephone calls. 
     
     
         3 . The method of  claim 2 , wherein the speaker information for the at least one target speaker are obtained by obtaining a plurality of speech samples of the at least one target speaker. 
     
     
         4 . The method of  claim 3 , wherein at least one of the plurality of speech samples of the at least one target speaker is obtained from a telephone call of the at least one target speaker. 
     
     
         5 . The method of  claim 4 , wherein the speech samples according to the at least one unknown speaker are classified by a speaker clustering technique, in particular, by Agglomerative Hierarchical Clustering. 
     
     
         6 . The method of  claim 5 , wherein the speaker clustering technique is based on a Gaussian Mixture Model and a Gaussian Mixture Model metric. 
     
     
         7 . The method of  claim 6 , wherein the speaker clustering technique employs a Joint Factor Analysis. 
     
     
         8 . The method of  claim 1 , wherein combining the extracted speaker information for each of the speaker-dependent classes of speech samples comprises generating for a particular class a combined Gaussian Mixture Model from the extracted speaker information of the speech samples of that class. 
     
     
         9 . The method of  claim 6 , wherein the combined Gaussian Mixture Model is generated from Gaussian Mixture Models of the speech samples of that class. 
     
     
         10 . The method of  claim 1 , wherein combining the extracted speaker information for each of the speaker-dependent classes of speech samples comprises combining feature vectors obtained for one or more speech samples of a speaker-dependent class with feature vectors of one or more other speech samples of the same speaker-dependent class, in particular, by summation of at least some of the feature vectors, more particularly, comprising adding a feature vector of one speech sample of the speaker-dependent class and another feature vector of another speech sample of the speaker-dependent class, if they are close to each other within predetermined limits. 
     
     
         11 . A computer program product, comprising one or more computer readable media having computer-executable instructions for performing steps of the method according to one of the preceding claims when run on a computer. 
     
     
         12 . A system for performing speaker recognition, comprising: a database stored and operating on a computer and configured to store speaker
 information for a target speaker;   software means stored and operating on the computer and configured to classify speech samples of telephone calls according to at least one unknown speaker thereby providing one, two or more speaker-dependent classes of speech samples;   software means stored and operating on the computer and configured to extract speaker information for the speech samples of each of the speaker-dependent classes of speech samples;   software means stored and operating on the computer and configured to combine the extracted speaker information for each of the speaker-dependent classes of speech samples;   software means stored and operating on the computer and configured to compare the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and   software means stored and operating on the computer and configured to determine whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.   
     
     
         13 . The system of  claim 12 , further comprising software means stored and operating on the computer and configured to receive telephone calls from at least one unknown speaker. 
     
     
         14 . The system of  claim 12 , further comprising software means stored and operating on the computer and configured to group the telephone calls according to the telephone numbers of the telephone calls.

Join the waitlist — get patent alerts

Track US2016019897A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.