US2018137880A1PendingUtilityA1
Phonation Style Detection
Assignee: GOVERMENT OF THE UNITED STATES AS REPRESENTED BY TE SECRETARY OF THE AIR FORCEPriority: Nov 16, 2016Filed: Feb 16, 2017Published: May 17, 2018
Est. expiryNov 16, 2036(~10.3 yrs left)· nominal 20-yr term from priority
G10L 15/02G10L 25/21G10L 25/84H04L 45/02G06F 21/60G10L 25/24G10L 15/16G10L 25/90G06F 21/32H04L 2209/12H04L 2209/08H04L 63/162H04L 45/54H04L 9/002G06F 21/79G06F 21/72
36
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The invention provides a method for detecting phonation style in dynamic communication environments and making software control decisions based on phonation styles enabling an audio message to be classified based on the phonation style such as, but not limited to: normal phonation, whispered phonation, softly spoken speech phonation, high-level phonation, babble phonation, and non-voice sounds. The purpose of the invention is to introduce the phonation style as a way to control computer software.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for phonation style detection, comprising:
detecting speech activity in a signal; extracting signal features from said detected speech activity; characterizing said extracted signal features; and performing a decision process on said characterized signal features which determines whether said detected speech activity is one of normally spoken speech, loudly spoken speech, softly spoken speech, whisper speech, babble, and non-voice sound.
2 . In the method of claim 1 , characterizing further comprises characterizing said extracted signal features in terms of harmonic measure, signal energy, mixed-excitation, clipping, and voicing.
3 . In the method of claim 2 , performing a decision process further comprises
classifying said speech activity as non-voice sounds in the absence of harmonics and low energy signal features.
4 . In the method of claim 2 , performing a decision process further comprises
classifying said speech activity as softly spoken speech in the absence of harmonics and low energy signal features but in the presence of voicing signal features.
5 . In the method of claim 2 , performing a decision process further comprises
classifying said speech activity as babble in the presence of harmonics and mixed excitation signal features.
6 . In the method of claim 2 , performing a decision process further comprises
classifying said speech activity as loudly spoken speech in the presence of harmonics and clipping but in the absence of mixed excitation signal features.
7 . In the method of claim 2 , performing a decision process further comprises
classifying said speech activity as normally spoken speech in the presence of harmonics but in the absence clipping and mixed excitation signal features.
8 . In the method of claim 2 , performing a decision process further comprises
classifying said speech activity as whisper speech in the absence of harmonics and voicing signal features but in the presence of low energy signal features.
9 . In the method of claim 1 , speech activity detection is performed on substantially 10 to 30 millisecond blocks of said signal.
10 . In the method of claim 1 , speech activity detection further comprises measurement of any one of the following: energy, pitch extraction, autocorrelation, spectral tilt, and cepstral coefficients.
11 . In the method of claim 10 , said speech activity detection further comprises coupling said measurements with classifiers and learning algorithms.
12 . In the method of claim 11 , said classifiers and leaning algorithms are selected from the group comprising state vector machines, neural networks, and Gaussian mixture models.
13 . In the method of claim 10 , said measurement of energy further comprises computing discrete signal energy in the time domain according to:
E ( n )=Σ m=0 N−1 [w ( m ) s ( n−m )] 2
where w(m) is a weighting function; s is signal amplitude; n is a current discrete time sample; m is a current discrete time sample of a window of time; and E(m) is computed signal energy over said time window.
14 . In the method of claim 13 , said weighting function w(m) is selected from the group consisting of rectangular, hamming and triangular window functions.
15 . In the method of claim 10 , said measurement of energy further comprises computing signal energy in the frequency domain according to:
E=Σ f 1 f 2 |X[f]| 2 where X[f] is the Fourier transform of said signal; f 1 and f 2 are the Fourier transform frequency limits; and E is the computed energy of said signal.
16 . In the method of claim 15 , f 1 is 0 and f 2 is one-half the Nyquist rate.Join the waitlist — get patent alerts
Track US2018137880A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.