System and method for identification of a speaker by phonograms of spontaneous oral speech and by using formant equalization
Abstract
A system and method for identification of a speaker by phonograms of oral speech is disclosed. Similarity between a first phonogram of the speaker and a second, or sample, phonogram is evaluated by matching formant frequencies in referential utterances of a speech signal, where the utterances for comparison are selected from the first phonogram and the second phonogram. Referential utterances of speech signals are selected from the first phonogram and the second phonogram, where the referential utterances include formant paths of at least three formant frequencies. The selected referential utterances including at least two identical formant frequencies are compared therebetween. Similarity of the compared referential utterances from matching other formant frequencies is evaluated, where similarity of the phonograms is determined from evaluation of similarity of all the compared referential utterances.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for identification of a speaker by phonograms of oral speech, the method comprising:
evaluating, via a computing device, similarity between a first phonogram of the speaker and a second, or sample, phonogram by matching formant frequencies in referential utterances of a speech signal, wherein the utterances for comparison are selected from the first phonogram and the second phonogram; selecting referential utterances of speech signals from the first phonogram and the second phonogram, wherein the referential utterances comprise formant paths of at least three formant frequencies; comparing, via the computing device, therebetween the selected referential utterances comprising at least two identical formant frequencies; and evaluating, via the computing device, similarity of the compared referential utterances from matching other formant frequencies, wherein similarity of the phonograms is determined from evaluation of similarity of all the compared referential utterances.
2 . The method according to claim 1 , wherein the formant frequencies in each of the selected referential utterance are calculated as average values for fixed time intervals in which the formant frequencies are relatively constant.
3 . The method according to claim 1 , wherein the selection and comparison are conducted in relation to referential utterances comprising the same frequency values for the first two formants within the given typical variability limits of formant frequency values for the corresponding type of vowel phonemes in a given language.
4 . The method according to claim 1 , wherein the selection from the phonograms for comparison is conducted in relation to at least two referential utterances of a speech signal related to sounds articulated as differently as possible with maximum and minimum frequency values for the first and the second formants in the given phonogram.
5 . The method according to claim 1 , wherein,
before calculating values of formant frequencies, subjecting a power spectrum of the speech signal for each phonogram to inverse filtering, wherein the time average for each frequency component of the power spectrum is calculated, at least for particular utterances of the phonogram, and then the original value of the power spectrum of the phonogram for each frequency component of the spectrum is divided by its inverse mean value.
6 . The method according to claim 1 , wherein,
before calculating values of formant frequencies, subjecting a power spectrum of a speech signal for each phonogram to inverse filtering, wherein the time average for each frequency component of the power spectrum is calculated, at least for individual utterances of the phonogram, and then a logarithm of the spectra is taken, and the average value logarithm of the phonogram signal power spectrum for each frequency component is subtracted from its original value logarithm.
7 . A system for identification of a speaker by phonograms of oral speech, the system comprising:
a computer memory configured to store digital audio signal files representative of a plurality of phonograms converted into digital form; a computing device configured to:
evaluate similarity between a first phonogram of the speaker and a second phonogram by matching formant frequencies in referential utterances of a speech signal, wherein the utterances for comparison are selected from the first phonogram and the second phonogram;
select referential utterances of speech signals from the first phonogram and the second phonogram, wherein the referential utterances comprise formant paths of at least three formant frequencies;
compare therebetween the selected referential utterances comprising at least two identical formant frequencies; and
evaluate similarity of the compared referential utterances from matching other formant frequencies, wherein similarity of the phonograms is determined from evaluation of similarity of all the compared referential utterances.
8 . The system according to claim 7 , wherein the formant frequencies in each of the selected referential utterance are calculated as average values for fixed time intervals in which the formant frequencies are relatively constant.
9 . The system according to claim 7 , wherein the selection and comparison are conducted in relation to referential utterances comprising the same frequency values for the first two formants within the given typical variability limits of formant frequency values for the corresponding type of vowel phonemes in a given language.
10 . The system according to claim 7 , wherein the selection from the phonograms for comparison is conducted in relation to at least two referential utterances of a speech signal related to sounds articulated as differently as possible with maximum and minimum frequency values for the first and the second formants in the given phonogram.
11 . The system according to claim 7 , wherein the computing device additionally subjects a power spectrum of the speech signal for each phonogram to inverse filtering before calculating values of formant frequencies,
wherein the time average for each frequency component of the power spectrum is calculated, at least for particular utterances of the phonogram, and then the original value of the power spectrum of the phonogram for each frequency component of the spectrum is divided by its inverse mean value.
12 . The system according to claim 7 , wherein the computing device additionally subjects a power spectrum of a speech signal for each phonogram to inverse filtering before calculating values of formant frequencies,
wherein the time average for each frequency component of the power spectrum is calculated, at least for individual utterances of the phonogram, and then a logarithm of the spectra is taken, and the average value logarithm of the phonogram signal power spectrum for each frequency component is subtracted from its original value logarithm.
13 . A system for identification of a speaker by phonograms of oral speech, the system comprising:
means for evaluating similarity between a first phonogram of the speaker and a second, or sample, phonogram by matching formant frequencies in referential utterances of a speech signal, wherein the utterances for comparison are selected from the first phonogram and the second phonogram; means for selecting referential utterances of speech signals from the first phonogram and the second phonogram, wherein the referential utterances comprise formant paths of at least three formant frequencies; means for comparing therebetween the selected referential utterances comprising at least two identical formant frequencies; and means for evaluating similarity of the compared referential utterances from matching other formant frequencies, wherein similarity of the phonograms is determined from evaluation of similarity of all the compared referential utterances.
14 . The system according to claim 13 , wherein the formant frequencies in each of the selected referential utterance are calculated as average values for fixed time intervals in which the formant frequencies are relatively constant.
15 . The system according to claim 13 , wherein the selecting and comparing are conducted in relation to referential utterances comprising the same frequency values for the first two formants within the given typical variability limits of formant frequency values for the corresponding type of vowel phonemes in a given language.
16 . The system according to claim 13 , wherein the selecting from the phonograms for comparing is conducted in relation to at least two referential utterances of a speech signal related to sounds articulated as differently as possible with maximum and minimum frequency values for the first and the second formants in the given phonogram.
17 . The system according to claim 13 , additionally comprising
means for subjecting a power spectrum of the speech signal for each phonogram to inverse filtering before calculating values of formant frequencies, wherein the time average for each frequency component of the power spectrum is calculated, at least for particular utterances of the phonogram, and then the original value of the power spectrum of the phonogram for each frequency component of the spectrum is divided by its inverse mean value.
18 . The system according to claim 13 , additionally comprising
means for subjecting a power spectrum of a speech signal for each phonogram to inverse filtering before calculating values of formant frequencies, wherein the time average for each frequency component of the power spectrum is calculated, at least for individual utterances of the phonogram, and then a logarithm of the spectra is taken, and the average value logarithm of the phonogram signal power spectrum for each frequency component is subtracted from its original value logarithm.Join the waitlist — get patent alerts
Track US2013325470A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.