US2015340035A1PendingUtilityA1

Automated generation of phonemic lexicon for voice activated cockpit management systems

Assignee: SERBAN DOINITAPriority: Feb 7, 2014Filed: Aug 1, 2015Published: Nov 26, 2015
Est. expiryFeb 7, 2034(~7.5 yrs left)· nominal 20-yr term from priority
G10L 15/187G10L 15/18G10L 17/22G10L 15/14G10L 15/02
25
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system, method and program for acquiring from an input text a character string set and generating the pronunciation thereof which should be recognized as a word is disclosed. The system selects from an input text, plural candidate character strings which are phonemic character candidates or allophones to be recognized as a word; generates plural pronunciation candidates of the selected candidate character string and outputs the optimum pronunciation candidate to be recognized as a word; generates phonemic dictionary by combining data in which the pronunciation candidate with optimal recognition is respectively associated with the character strings; generates recognition data in which character strings respectively indicating plural words contained in the input speech are associated with pronunciations; and outputs a combination contained in the recognition data, out of combinations each consisting of one of the candidate character strings and the one of the pronunciations candidates with the optimum recognition.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . An automation method of acquiring, from an input text and an input speech, a set of allophone character string and a pronunciation thereof which should be recognized as a word, a word in a sentence, and a word in a procedure, the automation method comprising operating one or more processors executing stored program instructions automatically to:
 select, from the input text, at least one allophone candidate character string which is a candidate to be recognized as a word;   generate at least one pronunciation candidate of each of the selected allophone candidate character strings by combining predetermined pronunciations of all allophone characters contained in the selected allophone candidate character string, while one or more pronunciations are predetermined for each of the allophone characters;   generate score data by combining data in which the generated pronunciation candidates are respectively associated with the allophone character strings, with language model data prepared by previously recording numerical values based on scores at which the respective words appear in the text and speech; the score data indicating appearance accuracy of the respective sets each consisting of an allophone character string indicating a word, a word in a sentence, a sentence in a procedure, and the pronunciation thereof;   based on the generated score data, perform speech recognition on the input speech to generate recognition data in which allophone character strings respectively indicating plural words contained in the input speech are associated with pronunciations; and   select and output a combination contained in the recognition data, out of combinations each comprising one of the allophone candidate character strings and one of the pronunciation candidates.   
     
     
         2 . A computer program product embodied in computer readable memory for enabling an information processing apparatus to function as a system for acquiring, from an input text and input speech, a set of allophone character strings and the pronunciation thereof which should be recognized as a word, a word in a sentence, and a sentence in a procedure, the computer program product comprising stored program instructions which, when executed by one or more processors, enable the information processing apparatus to function as:
 a candidate selecting unit for selecting, form the input text, at least one allophone candidate character string which is a candidate to be recognized as a word, a word in a sentence, and a sentence in a procedure;   a pronunciation generating unit for generating at least one pronunciation candidate of each of the selected allophone candidate character strings by combining pronunciations of all allophone characters contained in the selected allophone candidate character strings, while one or more pronunciations are predetermined for each of the allophone characters;   a score generating unit for generating confidence score data by combining data in which the generated pronunciations candidates are respectively associated with the allophone character strings, with language model data prepared by previously recording numerical values based on accuracy with which respective words appear in the text, the accuracy data indicating the appearance accuracy of respective sets each consisting of an allophone character string indicating a word, a word in a sentence, and a sentence in a procedure, and the pronunciation thereof;   a speech recognizing unit for performing, based on the generated confidence data, speech recognition on the input speech to generate recognition data in which allophone character strings respectively indicating plural words contained in the input speech are associated with pronunciations; and   an outputting unit for selecting and outputting a combination contained in the recognition data, out of combinations each consisting of one of the allophone candidate character strings and one of the candidates of a pronunciation thereof.   
     
     
         3 . A method for acquiring, from an input text and an input speech, a set of an allophone character string and a pronunciation thereof which should be recognized as a word, a word in a sentence, and a sentence in a procedure, the method comprising:
 an allophone candidate selecting unit wherein the allophone selecting unit repeats processing of adding other allophone characters to a certain allophone character string containing an input text character by character at the front-end or the tail-end of the certain character string until and optimization score in the input text of an allophone character string obtained by such addition is reached, and selects the allophone character string before the addition as the allophone candidate character string, and;   an allophone candidate selecting unit comprising one or more processors executing stored program instructions for selecting from the input text, at least one allophone candidate character string which is a candidate to be recognized as a word, a word in a sentence, and a sentence in a procedure;   a pronunciation generating unit comprising one or more processors executing stored program instructions for generating at least one pronunciation candidate of each of the selected allophone candidate character strings on the basis of respective allophone characters contained in the selected allophone candidate character strings; and   a word acquiring unit comprising one or more processors executing stored program instructions for selecting and outputting one of the generated allophone candidate character strings and corresponding one of the pronunciation candidates, on conditions that the selected pronunciation candidate is contained in the input text, and that two contexts in the input speech are similar to each other to an extent not less than a predetermined criterion, one of the contexts having the selected pronunciation candidate appear, and the other of the contexts having the selected allophone candidate character string appear.

Join the waitlist — get patent alerts

Track US2015340035A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.