US2022013120A1PendingUtilityA1

Automatic speech recognition

Assignee: VoicEncode LtdPriority: Jun 14, 2016Filed: Sep 13, 2021Published: Jan 13, 2022
Est. expiryJun 14, 2036(~9.9 yrs left)· nominal 20-yr term from priority
Inventors:Omry Netzer
G06V 10/70G10L 15/22G10L 15/02G10L 25/18G10L 15/30G06V 10/40G06K 9/46G06K 9/3233G06V 10/25
40
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method of speech recognition, sequentially executed by a processor on consecutive speech segments that comprises: obtaining digital information, which is a spectrogram representation, of a speech segment, and extracting from it speech features that characterizes the segment from the spectrogram representation. Then, a consistent structure segment vector based on the speech features is determined onto which machine learning is deployed to determine at least one label of the segment vector. A method of voice recognition and image recognition sequentially executed by a processor, on consecutive voice segments is also described. A system for executing speech, voice, and image recognition is also provided that comprises client devices to obtain and display information, a segment vector generator to determine a consistent structure segment vector based on features, and a machine learning server to determine at least one label of the segment vector.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of speech recognition, sequentially executed by a processor, on a plurality of consecutive speech segments, the method comprising:
 obtaining digital information of a speech segment, the digital information comprising a spectrogram representation; and   assigning at least one label to said speech segment by:
 dividing each of a plurality of time frames of the speech segment to a plurality of frequency bands and each of the plurality of frequency bands to a plurality of frequency bins each having a bin value; 
 calculating a speech feature value of each of a plurality of speech features, for each of a plurality of combinations, each of the combinations is of one of the plurality of time frames and one of the plurality of respective frequency bands, said plurality of speech features includes at least a mean value of the respective bin values, a standard deviation value of the respective bin values, a maximum value of the respective bin values and a voice-unvoiced ratio value of the respective bin values; 
 determining a segment vector based on inner relations between two or more speech features of different combinations from the plurality of combinations, said inner relations represent cross effects between said two or more speech features; and 
 determining the at least one label by classifying said segment vector using machine learning classification algorithm receiving as input at least one labeled segment vector of respective at least one previously analyzed speech segment. 
   
     
     
         2 . The method of  claim 1 , wherein said obtaining digital information further comprising digitizing, by a processor, an analog speech signal originated from a device having a microphone in real time or from a device having an audio recording, wherein, the analog speech signal comprising analog voice portions and non-voice portions; and wherein the digitizing of the analog voice portion produces the digital information of a segment. 
     
     
         3 . The method of  claim 1 , wherein the speech segment represents an element selected from a group comprising of: a syllable, a plurality of syllables, a word, a fraction of a word, a plurality of words and any combination thereof. 
     
     
         4 . The method of  claim 1 , wherein the calculating the speech feature value of each of the plurality of speech features, for each of the plurality of combinations, comprises assembling a plurality of matrixes and an index matrix, having identical number of cells, wherein each matrix of the plurality of matrixes represent a different feature of the plurality of speech features, wherein assembling the index matrix is based on a spectrogram having said plurality of time frames and said plurality of frequency bands, wherein the index matrix dimensions correlates with the plurality of time frames and the plurality of frequency bands of the spectrogram, wherein the plurality of matrixes overlap with the index matrix, and wherein a content of each cell of each matrix of the plurality of matrixes represents a speech feature value of a time frame and a frequency band indicated by the index matrix. 
     
     
         5 . The method of  claim 4 , wherein one or more portions of frequency bands of the index matrix having a time duration which expands along a total number of time frames smaller than a minimal duration defined by a threshold of minimum number of consecutive time frames, are filtered out of the index matrix and the plurality of matrixes. 
     
     
         6 . The method of  claim 4 , wherein contiguous time frames containing similar speech features values are replaced with a time interval in the index matrix and the plurality of matrixes. 
     
     
         7 . The method of  claim 4 , wherein said inner relations between two or more speech features are inner relations between band pairs, wherein the determining a segment vector further comprises compiling a plurality of components each comprising equal number of operands, wherein the first component of the plurality of components is an index component corresponding with the index matrix while the rest of the plurality of components are features components corresponding with the features matrixes, wherein a total number of operands is all possible combinations of frequency bands pairs, and wherein the index component indicates operands having band pairs presence in the segment vector. 
     
     
         8 . The method of  claim 7 , wherein the segment vector further comprises said inner relations. 
     
     
         9 . The method of  claim 7 , wherein properties of operands, having pairs presence, of each feature component are determined by calculating cross effect between sets of aggregated pairs, wherein each set of aggregated pairs is associated with a predetermined time zone of the segment. 
     
     
         10 . The method of  claim 1 , wherein said at least one label comprising at least one alphanumeric character manifestation of a speech segment and wherein said at least one label is a representation of at least one member of a group consisting of: an accent, a pronunciation level, an age of a speaker and a gender of the speaker. 
     
     
         11 . A system for speech recognition, comprising:
 at least one hardware processor adapted to execute code, said code comprising code instructions to sequentially conduct analysis on a plurality of consecutive speech segments, said analysis comprising:   obtaining digital information of a speech segment, the digital information comprising a spectrogram representation; and   assigning at least one label to said speech segment by:
 dividing each of a plurality of time frames of the speech segment to a plurality of frequency bands and each of the plurality of frequency bands to a plurality of frequency bins each having a bin value; 
 calculating a speech feature value of each of a plurality of speech features, for each of a plurality of combinations, each of the combinations is of one of the plurality of time frames and one of the plurality of respective frequency bands, said plurality of speech features includes at least a mean value of the respective bin values, a standard deviation value of the respective bin values, a maximum value of the respective bin values and a voice-unvoiced ratio value of the respective bin values; 
 determining a segment vector based on inner relations between two or more speech features of different combinations from the plurality of combinations, said inner relations represent cross effects between said two or more speech features; and 
 determining the at least one label by classifying said segment vector using machine learning classification algorithm receiving as input at least one labeled segment vector of respective at least one previously analyzed speech segment. 
   
     
     
         12 . The system of  claim 11 , wherein said obtaining said digital information is conducted from devices selected from a group comprising of: image capturing device, video capturing device, images storage, video storage, a real time sound sensor and a sound recording system. 
     
     
         13 . The system of  claim 11 , wherein said obtaining digital information further comprising digitizing an analog speech signal originated from a device having a microphone in real time or from a device having an audio recording, wherein the analog speech signal comprising analog voice portions and non-voice portions, and wherein the digitizing of the analog voice portion produces the digital information of a segment. 
     
     
         14 . The system of  claim 11 , wherein the speech segment represents an element selected from a group comprising of: a syllable, a plurality of syllables, a word, a fraction of a word, a plurality of words and any combination thereof. 
     
     
         15 . The system of  claim 11 , wherein the calculating the speech feature value of each of the plurality of speech features, for each of the plurality of combinations, comprises assembling a plurality of matrixes and an index matrix, having identical number of cells, wherein each matrix of the plurality of matrixes represent a different feature of the plurality of speech features, wherein assembling the index matrix is based on a spectrogram having said plurality of time frames and said plurality of frequency bands, wherein the index matrix dimensions correlates with the plurality of time frames and the plurality of frequency bands of the spectrogram, wherein the plurality of matrixes overlap with the index matrix, and wherein a content of each cell of each matrix of the plurality of matrixes represents a speech feature value of a time frame and a frequency band indicated by the index matrix. 
     
     
         16 . The system of  claim 15 , wherein one or more portions of frequency bands of the index matrix having a time duration which expands along a total number of time frames smaller than a minimal duration defined by a threshold of minimum number of consecutive time frames, are filtered out of the index matrix and the plurality of matrixes; and 
     
     
         17 . The system of  claim 15 , wherein contiguous time frames containing similar speech features values are replaced with a time interval in the index matrix and the plurality of matrixes. 
     
     
         18 . The system of  claim 15 , wherein said inner relations between two or more speech features are inner relations between band pairs, wherein the determining a segment vector further comprises compiling a plurality of components each comprising equal number of operands, wherein the first component of the plurality of components is an index component corresponding with the index matrix while the rest of the plurality of components are features components corresponding with the features matrixes, wherein a total number of operands is all possible combinations of frequency bands pairs, and wherein the index component indicates operands having band pairs presence in the segment vector; and
 wherein the segment vector further comprises said inner relations.   
     
     
         19 . The system of  claim 15 , wherein properties of operands, having pairs presence, of each feature component are determined by calculating cross effect between sets of aggregated pairs, wherein each set of aggregated pairs is associated with a predetermine time zone of the segment. 
     
     
         20 . The system of  claim 19 , wherein properties of operands, having pairs presence, of each feature component are determined by calculating cross effect between sets of aggregated pairs, wherein each set of aggregated pairs is associated with a predetermine time zone of the segment.

Join the waitlist — get patent alerts

Track US2022013120A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.