US2012215528A1PendingUtilityA1

Speech recognition system, speech recognition request device, speech recognition method, speech recognition program, and recording medium

Assignee: NAGATOMO KENTAROPriority: Oct 28, 2009Filed: Oct 12, 2010Published: Aug 23, 2012
Est. expiryOct 28, 2029(~3.3 yrs left)· nominal 20-yr term from priority
G10L 15/30G10L 15/26G10L 17/00G10L 15/02G10L 15/22G10L 2015/025G10L 15/187G06F 21/32
40
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Provided is a speech recognition system, including: a first information processing device including a speech recognition processing unit for receiving data to be used for speech recognition transmitted via a network, carrying out speech recognition processing, and returning resultant data; and a second information processing device connected to the first information processing device via the network. The second information processing device performs conversion of the data into data having a format that disables a content thereof from being perceived and also enables the speech recognition processing unit to perform the speech recognition processing. Thereafter, the second information processing device transmits the data to be used for the speech recognition by the speech recognition processing unit and constructs resultant data returned from the first information processing device into a content of a valid and perceivable recognition result.

Claims

exact text as granted — not AI-modified
1 . A speech recognition system, comprising:
 a first information processing device comprising a speech recognition processing unit for receiving data which are used for speech recognition and which are transmitted via a network, carrying out speech recognition processing, and returning resultant data; and   a second information processing device connected to the first information processing device via the network, for performing conversion of data used for the speech recognition by the speech recognition processing unit into data having a format that disables a content thereof from being perceived and that also enables the speech recognition processing unit to perform the speech recognition processing to obtain converted data in the format and to thereafter transmit the converted data to the speech recognition processing unit, and constructing the resultant data returned from the first information processing device into a content of a valid and perceivable recognition result.   
     
     
         2 . A speech recognition system, comprising:
 a first information processing device comprising a speech recognition processing unit for receiving data to be used for speech recognition transmitted via a network, carrying out speech recognition processing, and returning resultant data; and   a second information processing device which is connected to the first information processing device via the network, which transmits the data to be used for the speech recognition by the speech recognition processing unit after performing mapping thereof by using a mapping function unknown to the first information processing device, and constructing a speech recognition result by modifying the resultant data returned from the first information processing device into the same result as a result of performing the speech recognition without using the mapping function.   
     
     
         3 . A speech recognition system, comprising a plurality of information processing devices that are connected to one another via a network and comprise a speech recognition processing unit in at least one information processing device, wherein:
 the information processing device with the speech recognition processing unit receives at least one data structure of data to be used for speech recognition processing by the speech recognition processing unit;   wherein:   the at least one data structure of the data is converted by using a mapping function and transmitted to the information processing device with the speech recognition processing unit;   the information processing device with the speech recognition processing unit carries out the speech recognition processing based on the converted data structure and transmits a result thereof; and   the result of carrying out the speech recognition processing which is affected by the mapping function is constructed into a result of carrying out the speech recognition processing which is not affected by the mapping function.   
     
     
         4 . A speech recognition system according to  claim 2 , wherein the mapping function that is used comprises a mapping function Φ in which, when the mapping function Φ={φ} maps a data structure X and a data structure Y to φ_x{X} and φ_y{Y}, respectively, with regard to a function F(X,Y) used by the speech recognition processing unit, values of F(X,Y) and F(φ_x{X}, φ_y{Y}) are constantly the same or constantly less than a given threshold value, or a ratio therebetween is constantly fixed. 
     
     
         5 . A speech recognition system according to  claim 2 , wherein a data structure used by the speech recognition processing unit indicates a reference relationship between a given index and a reference destination in relation to an index that refers to specific data included in the data structure. 
     
     
         6 . A speech recognition system according to  claim 2 , wherein the mapping function comprises a function in which:
 with regard to a reference relationship between an index that refers to specific data included in a given data structure and a reference destination, a destination to which a given arbitrary index refers before mapping does not necessarily match a destination to which the same index refers after the mapping; and   data at the reference destination to which any one of indices refers before the mapping is always referred to by any one of the indices after the mapping.   
     
     
         7 . A speech recognition system according to  claim 6 , wherein the mapping function comprises shuffling of indices that refer to the specific data included in the given data structure. 
     
     
         8 . A speech recognition system according to  claim 6 , wherein the mapping function adds an arbitrary number of indices to the specific data included in the given data structure. 
     
     
         9 . A speech recognition system according to  claim 2 , wherein at least one item of data to be used for speech recognition which is subjected to mapping by using the mapping function is retained before the mapping only on an information processing device for inputting a sound to be subjected to the speech recognition. 
     
     
         10 . A speech recognition system according to  claim 2 , wherein the data to be used by the speech recognition processing unit has a structure to which at least one selected from the group consisting of a structure of an acoustic model, a structure of a language model, and a structure of a feature vector is mapped. 
     
     
         11 . A speech recognition system according to  claim 10 , wherein:
 indices indicating respective features included in the feature vector are mapped by using the mapping function given by a device for inputting a sound to be subjected to speech recognition; and   indices to models associated with respective features within the acoustic model are mapped by using the mapping function given by the device for inputting the sound to be subjected to the speech recognition.   
     
     
         12 . A speech recognition system according to  claim 11 , wherein:
 phoneme IDs being indices to phonemes included in the acoustic model are mapped by using the mapping function given by the device for inputting the sound;   phoneme ID strings indicating pronunciations of respective words included in the language model are mapped by using the mapping function given by the device for inputting the sound; and   at least information on representation character strings of the respective words included in the language model is deleted.   
     
     
         13 . A speech recognition system according to  claim 12 , wherein word IDs being indices to the respective words included in the language model are mapped by using the mapping function given by the device for inputting the sound. 
     
     
         14 . A speech recognition system according to  claim 2 , comprising the information processing device which is operable in response to the speech data and which comprises at least an acoustic likelihood computation unit and is configured to:
 map phoneme ID strings indicating pronunciations of respective words included in the language model by using the mapping function given by the information processing device, and delete at least information on representation character strings of the respective words included in the language model;   compute acoustic likelihoods of all known phonemes or necessary phonemes for each frame of the speech data to generate a sequence of a group of the phoneme IDs and acoustic likelihoods that are mapped by using the mapping function given by the information processing device; and   transmit the sequence of the group of the mapped phoneme IDs and acoustic likelihoods and the language model after the mapping to the information processing device comprising a hypothesis search unit.   
     
     
         15 . A speech recognition system according to  claim 2 , comprising the information processing device which is operable in response to speech data and which is configured to:
 divide the speech data into blocks;   map a time sequence among the divided blocks by using the mapping function given by the information processing device for inputting speech data;   transmit the blocks of speech to an information processing device for performing speech recognition based on the time sequence after the mapping;   receive any one of a feature vector or a sequence of a group of phoneme IDs and acoustic likelihoods from the information processing device for performing the speech recognition; and   restore the time sequence by using an inverse function to the mapping function given by the information processing device for inputting speech data.   
     
     
         16 . A speech recognition request device, comprising:
 a communication unit connected via a network to a speech recognition device comprising a speech recognition processing unit for receiving data to be used for speech recognition transmitted via the network, carrying out speech recognition processing, and returning resultant data;   an information conversion unit for converting the data to be used for the speech recognition by the speech recognition processing unit into data having a format that disables a content thereof from being perceived and also enables the speech recognition processing unit to perform the speech recognition processing; and   an recognition result construction unit for reconstructing the resultant data returned from the speech recognition device after performing the speech recognition on the converted data into a speech recognition result that is perceivable as a content of being valid recognition result, based on the converted content.   
     
     
         17 . A speech recognition request device, comprising:
 a communication unit connected via a network to a speech recognition device comprising a speech recognition processing unit for receiving data to be used for speech recognition transmitted via the network, carrying out speech recognition processing, and returning resultant data;   an information conversion unit for mapping the data to be used for the speech recognition by the speech recognition processing unit by using a mapping function unknown to the speech recognition device; and   an recognition result construction unit which is operable on the basis of the mapping function and which constructs the resultant data returned from the speech recognition device to obtain, from the resultant data, the same result as a result of performing the speech recognition without using the mapping function.   
     
     
         18 . A speech recognition request device according to  claim 17 , wherein the information conversion unit maps a data structure of the data to be used for the speech recognition which is transmitted to the speech recognition processing unit so as to indicate a reference relationship between a predetermined index and a reference destination in relation to an index that refers to specific data included in the data structure. 
     
     
         19 . A speech recognition request device according to  claim 17 , wherein the mapping function comprises a function in which:
 with regard to a reference relationship between an index that refers to specific data included in a given data structure and a reference destination, a destination to which a given arbitrary index refers before mapping does not necessarily match a destination to which the same index refers after the mapping; and   data at the reference destination to which any one of indices refers before the mapping are always referred to by any one of the indices after the mapping.   
     
     
         20 . A speech recognition request device according to  claim 17 , wherein:
 indices indicating respective features included in the feature vector are mapped by using the mapping function; and   indices to models associated with respective features within the acoustic model are mapped by using the mapping function.   
     
     
         21 . A speech recognition request device according to  claim 17 , wherein:
 phoneme IDs being indices to phonemes included in the acoustic model are mapped by using the mapping function;   phoneme ID strings indicating pronunciations of respective words included in the language model are mapped by using the mapping function; and   at least information on representation character strings of the respective words included in the language model is deleted.   
     
     
         22 . A speech recognition request device according to  claim 17 , further comprising an acoustic likelihood computation unit and being configured to:
 map phoneme ID strings indicating pronunciations of respective words included in the language model by using the mapping function, and delete at least information on representation character strings of the respective words included in the language model;   compute acoustic likelihoods of all known phonemes or necessary phonemes for each frame of the speech data to generate a sequence of a group of the phoneme IDs and acoustic likelihoods that are mapped by using the mapping function given by the speech recognition device; and   transmit the sequence of the group of the mapped phoneme IDs and acoustic likelihoods and the language model after the mapping to the speech recognition device comprising a hypothesis search unit.   
     
     
         23 . A speech recognition request device according to  claim 17 , further configured to:
 divide speech data of a sound to be subjected to the speech recognition into a plurality of blocks;   map a time sequence among the divided blocks by using the mapping function;   transmit the blocks of speech to the speech recognition device based on the time sequence after the mapping; and   receive result data on the speech recognition transmitted from the speech recognition device, and restore the time sequence by using an inverse function to the mapping function.   
     
     
         24 . An information processing device, comprising:
 means for storing an acoustic model, a language model, and conversion/reconstruction data used for conversion that achieves secrecy and restoration;   a first conversion means for acquiring the acoustic model, the language model, and the conversion/reconstruction data, and converting a data structure of each model used for speech recognition into a data structure having the secrecy;   a second conversion means for converting a sound to be subjected to identification into data, and converting a data structure of the data into a data structure having the secrecy;   means for transmitting the converted data to an acoustic recognition device via a network; and   means for constructing a recognition result equivalent to a result of performing the speech recognition without using the first conversion means and the second conversion means, based on a result of the speech recognition received from the acoustic recognition device via the network, the acoustic model, the language model, and the conversion/reconstruction data.   
     
     
         25 . A speech recognition method, comprising:
 connecting a speech recognition device comprising a speech recognition processing unit and a speech recognition request device for requesting the speech recognition device for speech recognition to each other via a network;   converting, by the speech recognition request device, at least one data structure of data to be used for speech recognition processing by the speech recognition processing unit by using a mapping function, and transmitting the resultant to the speech recognition device;   carrying out, by the speech recognition device, the speech recognition processing based on the converted data structure, and transmitting a result thereof to the speech recognition request device; and   constructing, by the speech recognition request device, the result of carrying out the speech recognition processing which is affected by the mapping function into a result of carrying out the speech recognition processing which is not affected by the mapping function.   
     
     
         26 . A speech recognition method according to  claim 25 , wherein the data to be used by the speech recognition processing unit, which is converted and transmitted from the speech recognition request device to the speech recognition device, has a structure to which at least one selected from the group consisting of a structure of an acoustic model, a structure of a language model, and a structure of a feature vector is mapped. 
     
     
         27 . A speech recognition method according to  claim 25 , wherein the mapping function comprises a function of shuffling indices that refer to specific data included in a given data structure or adding an arbitrary number of indices to the indices that refer to the specific data included in the given data structure. 
     
     
         28 . A speech recognition method according to  claim 25 , wherein the mapping function that is used comprises a mapping function Φ in which, when the mapping function Φ={φ} maps a data structure X and a data structure Y to φ_x {X} and φ_y{Y}, respectively, with regard to a function F(X,Y) used by the speech recognition processing unit, values of F(X,Y) and F(φ_x {X}, φ_y{Y}) are constantly the same or constantly less than a given threshold value, or a ratio therebetween is constantly fixed. 
     
     
         29 . A non-transitory recording medium having recorded thereon a speech recognition program which is used in a control unit of an information processing device coupled through a network to a speech recognition processing device comprising a speech recognition unit which receives data to be used through the network, which carries out speech recognition processing, and which returns resultant data via the network;
 the speech recognition program making the control unit operate as:   a communication unit connected via a network to the speech recognition device;   an information conversion unit which converts the data used for the speech recognition by the speech recognition processing unit into data of a format that disables a content thereof from being perceived and also enables the speech recognition processing unit to perform the speech recognition processing; and   an recognition authentication result construction unit which reconstructs the resultant data returned from the speech recognition device after performing the speech recognition on the converted data into a speech recognition result that enables a content being a valid and perceivable recognition result, based on the converted content.   
     
     
         30 . A non-transitory recording medium having recorded thereon a speech recognition program used in a control unit of an information processing device which is coupled through a network to an acoustic recognition device and which comprises;
 means for storing an acoustic model, a language model, and conversion/reconstruction data used for conversion that achieves secrecy and restoration; and   means for transmitting the converted data to the an acoustic recognition device via a network,   the speech recognition program making the control unit operate as:   a first conversion means for acquiring the acoustic model, the language model, and the conversion/reconstruction data, to convert a data structure of each model used for speech recognition into a data structure having the secrecy;   a second conversion means for converting a sound to be subjected to identification into data, to convert a data structure of the data into a data structure having the secrecy; and   means for constructing a recognition result equivalent to a result of performing the speech recognition without using the first conversion means and the second conversion means, based on a result of the speech recognition received from the acoustic recognition device via the network, the acoustic model, the language model, and the conversion/reconstruction data.   
     
     
         31 . (canceled) 
     
     
         32 . (canceled)

Join the waitlist — get patent alerts

Track US2012215528A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.