US2005033566A1PendingUtilityA1

Natural language processing method

Assignee: CANON KKPriority: Jul 9, 2003Filed: Jul 8, 2004Published: Feb 10, 2005
Est. expiryJul 9, 2023(expired)· nominal 20-yr term from priority
Inventors:Michio Aizawa
G06F 40/268
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A sentence appended with information associated with a pause setting position is input (S 201 ), and a morphological analysis process is applied to the sentence to divide the sentence into words and to determine parts of speech of respective words (S 203 ). Part-of-speech sequences, each of which includes parts of speech of a total of N (N≧2) words before and after each word boundary, are obtained for respective word boundaries, and the frequencies of occurrence of arrangements of the parts of speech are calculated for respective groups of part-of-speech sequences with the same arrangements of parts of speech (S 206 ). Pause counts, each of which indicates the number of times of setting of a pause setting position indicated by the pause setting position data between parts of speech in the part-of-speech sequence, are calculated for respective groups of the part-of-speech sequences with the same arrangements of parts of speech (S 208 ). Pause insertability values are calculated using the frequencies of occurrence and pause counts for respective groups (S 210 ).

Claims

exact text as granted — not AI-modified
1 . A natural language processing method comprising: 
 a reception step of receiving a sentence appended with information associated with a pause setting position;    a part-of-speech acquisition step of acquiring parts of speech of respective words in the sentence;    an acquisition step of acquiring frequencies of occurrence of arrangements of the parts of speech for respective part-of-speech sequence groups with the same arrangements of parts of speech corresponding to arrangements of the words;    a count step of counting the number of pause setting positions each of which is present between parts of speech in the part-of-speech sequence for respective part-of-speech sequence groups on the basis of the information associated with the pause setting position; and    a calculation step of calculating pause insertability values using the frequencies of occurrence and the number of setting positions for respective part-of-speech sequence groups.    
   
   
       2 . The method according to  claim 1 , wherein the calculation step includes a step of calculating each pause insertability value based on a ratio between the frequency of occurrence and the number of setting positions.  
   
   
       3 . The method according to  claim 1 , wherein the information associated with the pause setting position is a comma included in the sentence, and the count step includes a step of counting the number of commas included between parts of speeches in the part-of-speech sequences as the number of pause setting positions.  
   
   
       4 . The method according to  claim 1 , further comprising, in data of a sentence, which is divided into words, parts of speech of which are determined: 
 a pause setting step of setting a pause in a part-of-speech sequence having a largest pause insertability value calculated in the calculation step of various part-of-speech sequences which are located in a period from a first word boundary to a second word boundary separated a predetermined number of words from the first word boundary, and    in that the pause setting step includes a step of setting a position where the pause is set as a new first word boundary, and repeating pause setting process of the pause setting step for the next period.    
   
   
       5 . A natural language processing apparatus comprising: 
 reception means for receiving a sentence appended with information associated with a pause setting position;    part-of-speech acquisition means for acquiring parts of speech of respective words in the sentence;    acquisition means for acquiring frequencies of occurrence of arrangements of the parts of speech for respective part-of-speech sequence groups with the same arrangements of parts of speech corresponding to arrangements of the words;    count means for counting the number of pause setting positions each of which is present between parts of speech in the part-of-speech sequence for respective part-of-speech sequence groups on the basis of the information associated with the pause setting position; and    calculation means for calculating pause insertability values using the frequencies of occurrence and the number of setting positions for respective part-of-speech sequence groups.    
   
   
       6 . A program for making a computer execute a natural language processing method of  claim 1 .  
   
   
       7 . A computer readable storage medium storing a program of  claim 6.

Join the waitlist — get patent alerts

Track US2005033566A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.