US2002198713A1PendingUtilityA1

Method and apparatus for perfoming spoken language translation

Priority: Jan 29, 1999Filed: Jun 21, 2001Published: Dec 26, 2002
Est. expiryJan 29, 2019(expired)· nominal 20-yr term from priority
G10L 15/1815G10L 15/26
42
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and an apparatus for performing spoken language translation are provided, wherein a speech input is received comprising at least one source language. The speech input comprises words, sentences, and phrases in a natural spoken language. Source expressions are recognized in the source language. Misrecognitions of the source expressions resulting from factors comprising noise and speaker variation are minimized by the generation of intermediate data structures that encode at least one recognition hypothesis. Furthermore, misrecognitions are minimized by the generation of candidate recognized source expressions by processing the intermediate data structures using models comprising a general language model and a domain model. A recognized source expression is selected and confirmed by a user through a user interface. The recognized source expressions are translated from the source language to a target language, and a speech output is synthesized from the translated target language source expressions. Moreover, a meaning of the speech input is detected, and the meaning is rendered in the synthesized translated output.

Claims

exact text as granted — not AI-modified
What is claimed is:  
     
         1 . A method for performing spoken language translation, comprising: 
 receiving at least one speech input comprising at least one source language;    recognizing at least one source expression of the at least one source language;    translating the recognized at least one source expression from the at least one source language to at least one target language;    synthesizing at least one speech output from the translated at least one target language; and    providing the at least one speech output.    
     
     
         2 . The method of  claim 1 , further comprising minimizing misrecognitions of the at least one source expression, wherein the misrecognitions result from factors selected from the group comprising noise and speaker variation.  
     
     
         3 . The method of  claim 2 , wherein minimizing misrecognitions comprises: 
 generating at least one intermediate data structure, wherein the at least one intermediate data structure comprises at least one word graph and at least one n-best list, wherein the at least one intermediate data structure encodes at least one recognition hypothesis; and    generating at least one candidate recognized source expression by processing the at least one intermediate data structure using at least one model, wherein the at least one model is a model selected from the group comprising a general language model and a domain model.    
     
     
         4 . The method of  claim 3 , further comprising selecting one of the at least one candidate recognized source expressions, wherein the selection is performed using an interface selected from the group comprising at least one graphical user interface and at least one voice command interface.  
     
     
         5 . The method of  claim 3 , further comprising confirming one of the at least one candidate recognized source expressions, wherein the confirmation is performed using an interface selected from a group comprising at least one graphical user interface and at least one voice command interface.  
     
     
         6 . The method of  claim 1 , wherein the at least one speech input comprises natural spoken language, wherein the natural spoken language comprises at least one word, at least one phrase, and at least one sentence.  
     
     
         7 . The method of  claim 1 , further comprising: 
 detecting at least one meaning of the at least one speech input, wherein the at least one meaning comprises statements and questions; and    rendering the at least one meaning in the synthesized at least one speech output.    
     
     
         8 . The method of  claim 1 , wherein translating comprises: 
 performing morphological analysis of the recognized at least one source expression using at least one source language dictionary and at least one source language morphological rule;    generating at least one sequence of analyzed morphemes;    performing syntactic source language analysis using grammar rule-based processing and example-based processing;    generating at least one source language syntactic representation based on the source language analysis; and    performing source language to target language transfer using at least one example database and at least one thesaurus, wherein the morphological analysis and the syntactic source language analysis are independent of the transfer and a domain.    
     
     
         9 . The method of  claim 8 , further comprising: 
 generating at least one target language syntactic representation;    performing target language syntactic generation using at least one set of target language syntactic generation rules;    generating at least one sequence of target language morpheme specifications; and    performing target language morphological generation using at least one target language dictionary and at least one set of target language morphological generation rules.    
     
     
         10 . The method of  claim 8 , wherein the grammar rule-based processing comprises: 
 syntactic and morphological analysis in the at least one source language; and    syntactic and morphological generation in the at least one target language.    
     
     
         11 . The method of  claim 8 , wherein the example-based processing comprises performing the transfer from the at least one source language to the at least one target language using an example database, wherein the example database comprises at least one stored pair of corresponding expressions in the at least one source language and the at least one target language.  
     
     
         12 . An apparatus for spoken language translation comprising: 
 at least one processor;    an input coupled to the at least one processor, the input capable of receiving speech signals comprising at least one source language, the at least one processor configured to translate the received speech signals by, 
 recognizing at least one source expression of the at least one source language;  
 translating the recognized at least one source expression from the at least one source language to at least one target language; and  
 synthesizing at least one speech output from the translated at least one target language;  
   an output coupled to the at least one processor, the output capable of providing the synthesized at least one speech output.    
     
     
         13 . The apparatus of  claim 12 , wherein the processor is further configured to translate by minimizing misrecognitions of the at least one source expression, wherein the misrecognitions result from factors selected from the group comprising noise and speaker variation.  
     
     
         14 . The apparatus of  claim 13 , wherein the processor is further configured to minimize misrecognitions by: 
 generating at least one intermediate data structure, wherein the at least one intermediate data structure comprises at least one word graph and at least one n-best list, wherein the at least one intermediate data structure encodes at least one recognition hypothesis; and    generating at least one candidate recognized source expression by processing the at least one intermediate data structure using at least one model, wherein the at least one model is a model selected from the group comprising a general language model and a domain model.    
     
     
         15 . The apparatus of  claim 14 , wherein the processor is further configured to minimize misrecognitions by selecting one of the at least one candidate recognized source expressions, wherein the selection is performed using an interface selected from the group comprising at least one graphical user interface and at least one voice command interface.  
     
     
         16 . The apparatus of  claim 14 , wherein the processor is further configured to minimize misrecognitions by confirming one of the at least one candidate recognized source expressions, wherein the confirmation is performed using an interface selected from the group comprising it least one graphical user interface and at least one voice command interface.  
     
     
         17 . The apparatus of  claim 12 , wherein the at least one speech input comprises natural spoken language, wherein the natural spoken language comprises at least one word, at least one phrase, and at least one sentence.  
     
     
         18 . The apparatus of  claim 12 , wherein the processor is further configured to translate by: 
 detecting at least one meaning of the at least one speech input, wherein the at least one meaning comprises statements and questions; and    rendering the at least one meaning in the synthesized at least one speech output.    
     
     
         19 . The apparatus of  claim 12 , wherein translating comprises: 
 performing morphological analysis of the recognized at least one source expression using at least one source language dictionary and at least one source language morphological rule;    generating at least one sequence of analyzed morphemes;    performing syntactic source language analysis using grammar rule-based processing and example-based processing;    generating at least one source language syntactic representation; and    performing source language to target language transfer using at least one example database and at least one thesaurus, wherein the morphological analysis and the syntactic source language analysis are independent of the transfer and a domain.    
     
     
         20 . The apparatus of  claim 19 , wherein translating further comprises: 
 generating at least one target language syntactic representation;    performing target language syntactic generation using at least one set of target language syntactic generation rules;    generating at least one sequence of target language morpheme specifications; and    performing target language morphological generation using at least one target language dictionary and at least one set of target language morphological generation rules.    
     
     
         21 . The apparatus of  claim 19 , wherein the grammar rule-based processing comprises: 
 syntactic and morphological analysis in the at least one source language; and    syntactic and morphological generation in the at least one target language.    
     
     
         22 . The apparatus of  claim 19 , wherein the example-based processing comprises performing the transfer from the at least one source language to the at least one target language using an example database, wherein the example database comprises at least one stored pair of corresponding expressions in the at least one source language and the at least one target language.  
     
     
         23 . The apparatus of  claim 12 , further comprising at least one input device selected from the group comprising at least one microphone, at least one keyboard, at least one cursor, and at least one touch-sensitive screen.  
     
     
         24 . The apparatus of  claim 12 , further comprising at least one analog-to-digital converter, at least one digital-to-analog converter, at least one amplifier, and at least one output device selected from the group comprising at least one speaker and at least one display device.  
     
     
         25 . A computer readable medium containing executable instructions which, when executed in a processing system, causes the system to perform a method for spoken language translation, the method comprising: 
 receiving at least one speech input comprising at least one source language;    recognizing at least one source expression of the at least one source language;    translating the recognized at least one source expression from the at least one source language to at least one target language;    synthesizing at least one speech output from the translated at least one target language; and    providing the at least one speech output.    
     
     
         26 . The computer readable medium of  claim 25 , wherein the method further comprises minimizing misrecognitions of the at least one source expression, wherein the misrecognitions result from factors comprising noise and speaker variation.  
     
     
         27 . The computer readable medium of  claim 26 , wherein minimizing of misrecognitions comprises: 
 generating at least one intermediate data structure, wherein the at least one intermediate data structure comprises at least one word graph and at least one n-best list, wherein the at least one intermediate data structure encodes at least one recognition hypothesis; and    generating at least one candidate recognized source expression by processing the at least one intermediate data structure using at least one model, wherein the at least one model comprises a general language model and a domain model.    
     
     
         28 . The computer readable medium of  claim 27 , wherein the method further comprises selecting one of the at least one candidate recognized source expressions, wherein the selection is performed using an interface comprising at least one graphical user interface and at least one voice command interface.  
     
     
         29 . The computer readable medium of  claim 27 , wherein the method further comprises confirming one of the at least one candidate recognized source expressions, wherein the confirmation is performed using an interface comprising at least one graphical user interface and at least one voice command interface.  
     
     
         30 . The computer readable medium of  claim 25 , wherein the at least one speech input comprises natural spoken language, wherein the natural spoken language comprises at least one word, at least one phrase, and at least one sentence.  
     
     
         31 . The computer readable medium of  claim 25 , wherein the method further comprises: 
 detecting at least one meaning of the at least one speech input, wherein the at least one meaning comprises statements and questions; and    rendering the at least one meaning in the synthesized at least one speech output.    
     
     
         32 . The computer readable medium of  claim 25 , wherein translating comprises: 
 performing morphological analysis of the recognized at least one source expression using at least one source language dictionary and at least one source language morphological rule;    generating at least one sequence of analyzed morphemes;    performing syntactic source language analysis using grammar rule-based processing and example-based processing;    generating at least one source language syntactic representation; and    performing source language to target language transfer using at least one example database and at least one thesaurus, wherein the morphological analysis and the syntactic source language analysis are independent of the transfer and a domain.    
     
     
         33 . The computer readable medium of  claim 32 , wherein the method further comprises: 
 generating at least one target language syntactic representation;    performing target language syntactic generation using at least one set of target language syntactic generation rules;    generating at least one sequence of target language morpheme specifications; and    performing target language morphological generation using at least one target language dictionary and at least one set of target language morphological generation rules.    
     
     
         34 . The computer readable medium of  claim 32 , wherein the grammar rule-based processing comprises: 
 syntactic and morphological analysis in the at least one source language; and    syntactic and morphological generation in the at least one target language.    
     
     
         35 . The computer readable medium of  claim 32 , wherein the example-based processing comprises performing the transfer from the at least one source language to the at least one target language using an example database, wherein the example database comprises at least one stored pair of corresponding expressions in the at least one source language and the at least one target language.

Join the waitlist — get patent alerts

Track US2002198713A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.