US2008015842A1PendingUtilityA1

Statistical method and apparatus for learning translation relationships among phrases

Assignee: MICROSOFT CORPPriority: Nov 20, 2002Filed: Jul 9, 2007Published: Jan 17, 2008
Est. expiryNov 20, 2022(expired)· nominal 20-yr term from priority
Inventors:Robert C. Moore
G06F 40/44G06F 40/45
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention learns phrase translation relationships by receiving a parallel aligned corpus with phrases to be learned identified in a source language. Candidate phrases in a target language are generated and an inside score is calculated based on word association scores for words inside the source language phrase and candidate phrase. An outside score is calculated based on word association scores for words outside the source language phrase and candidate phrase. The inside and outside scores are combined to obtain a joint score.

Claims

exact text as granted — not AI-modified
1 . A system for identifying phrase translations in multi-word target units for identified source language phrases in multi-word source units, comprising: 
 an individual word association model configured to generate one or more candidate phrases and a score for each candidate phrase based on word associations between words inside the source and target language phrases and word associations between words outside the source and target language phrases.    
   
   
       2 . The system of  claim 1  wherein the source and target units are part of an aligned corpus, and further comprising: 
 a cross-sentence model configured to modify the score based on other candidate phrases generated for the source language phrase across the corpus, to obtain a modified score.    
   
   
       3 . The system of  claim 2  and further comprising: 
 a conversion model configured to convert the modified score into a desired confidence metric indicative of a confidence level associated with the candidate phrase as a translation of the source language phrase.    
   
   
       4 . A method of generating candidate phrases in multi-word target units in a target language as hypothesized translations of an identified phrase in a multi-word source unit in a source language, comprising: 
 identifying first target language words in the target unit that are most strongly associated with a word in the source language phrase;    identifying second target language words in the target unit that have a word in the source language phrase that is most strongly associated it; and    generating the candidate phrases as phrases that begin and end with first or second target language words.    
   
   
       5 . The method of  claim 4  wherein generating candidate phrases further comprises: 
 generating additional candidate phrases as phrases that begin with words starting with a capital letter and that end with first or second target language words.

Join the waitlist — get patent alerts

Track US2008015842A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.