US2010211381A1PendingUtilityA1

System and Method of Creating and Using Compact Linguistic Data

Assignee: RESEARCH IN MOTION LTDPriority: Jul 3, 2002Filed: Apr 27, 2010Published: Aug 19, 2010
Est. expiryJul 3, 2022(expired)· nominal 20-yr term from priority
Y10S707/99937G06F 40/216
39
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method of creating and using compact linguistic data are provided. Frequencies of words appearing in a corpus are calculated. Each unique character in the words is mapped to a character index, and characters in the words are replaced with the character indexes. Sequences of characters are mapped to substitution indexes, and the sequences of characters in the words are replaced with the substitution indexes. The words are grouped by common prefixes, and each prefix is mapped to location information for the group of words which start with the prefix.

Claims

exact text as granted — not AI-modified
1 . A system of creating compact linguistic data, comprising:
 a corpus; and   a linguistic data analyzer, wherein the linguistic data analyzer calculates frequencies of words appearing in the corpus, maps each unique character in the words to a character index, replaces each character in the words with the character index to which the character is mapped, maps sequences of characters that appear in the words to substitution indexes, replaces each sequence of characters in each word with the substitution index to which the sequence of characters are mapped, arranges the words into groups where each group contains words that start with a common prefix, and maps each prefix to location information for the group of words which start with the prefix, and   wherein the compact linguistic data includes the unique characters, the character indexes, the substitution indexes, the location information, the groups of words, and the frequencies of the words.   
     
     
         2 .- 13 . (canceled) 
     
     
         14 . A system of creating compact linguistic data, comprising:
 a corpus; and   a linguistic data analyzer,   wherein the linguistic data analyzer
 calculates frequencies of words appearing as independent words in the corpus, 
 maps each unique character in the words to a character index, 
 replaces each character in the words with the character index to which the character is mapped, 
 maps sequences of characters that appear in the words to substitution indexes, 
 replaces each sequence of characters in each word with the substitution index to which the sequence of characters are mapped, 
 arranges the words into groups where each group contains words that start with a common prefix, 
 maps each prefix to location information for the group of words which start with the prefix to create a prefix index, and 
 removes the prefix from the words in the groups of words; and 
   wherein the compact linguistic data includes the unique characters, the character indexes, the substitution indexes, the prefix index, the groups of words, and the frequencies of the words.   
     
     
         15 . The system of  claim 14 , further comprising:
 a user interface, comprising:   a text input device; and   a text output device; and   a text input logic unit,   wherein the text input logic unit receives a text prefix from the text input device, retrieves a plurality of predicted words from the compact linguistic data that start with the text prefix, selects one of the plurality of predicted words for display based on the frequencies of the plurality of predicted words, and displays the one predicted words using the text output device.   
     
     
         16 . The system of  claim 15 , wherein the text input device is a keyboard. 
     
     
         17 . The system of  claim 16 , wherein the keyboard is a reduced keyboard. 
     
     
         18 . The system of  claim 15 , wherein the user interface and the text input logic unit are implemented on a mobile communication device. 
     
     
         19 . The system of  claim 15 , wherein the text input logic unit selects one of the groups of words as the plurality of predicted words. 
     
     
         20 . The system of  claim 19 , wherein the selected group of words is selected based on the frequency of the group of words. 
     
     
         21 . The system of  claim 15 , wherein the text input logic unit is configured to update the frequencies of the words based on whether the predicted word is input by a device user. 
     
     
         22 . The system of  claim 14 , wherein for each group, only the maximum frequency, which is the highest frequency value in the group, is retained with full precision, and the frequencies of words with less than the maximum frequency are retained as a percentage of the maximum frequency. 
     
     
         23 . A computer-implemented method of creating compact linguistic data, comprising:
 performing, by a processor, the operations of:
 calculating frequencies of words appearing as independent words in the corpus, 
 mapping each unique character in the words to a character index, 
 replacing each character in the words with the character index to which the character is mapped, 
 mapping sequences of characters that appear in the words to substitution indexes, 
 replacing each sequence of characters in each word with the substitution index to which the sequence of characters are mapped, 
 arranging the words into groups where each group contains words that start with a common prefix, 
 mapping each prefix to location information for the group of words which start with the prefix to create a prefix index, and 
 removing the prefix from the words in the groups of words; and 
   storing, in electronic format, the unique characters, the character indexes, the substitution indexes, the prefix index, the groups of words, the frequencies of the words and the frequencies of the groups of words as compact linguistic data.   
     
     
         24 . The method of  claim 23 , further comprising:
 receiving a text prefix from a text input logic unit of a text input device,   retrieving a plurality of predicted words from the compact linguistic data that start with the text prefix,   selecting one of the plurality of predicted words for display based on the frequencies of the plurality of predicted words, and   displaying the one predicted words using a text output device.   
     
     
         25 . The method of  claim 24 , wherein the text input device is a keyboard. 
     
     
         26 . The method of  claim 25 , wherein the keyboard is a reduced keyboard. 
     
     
         27 . The method of  claim 24 , wherein the text input logic unit is implemented on a mobile communication device. 
     
     
         28 . The method of  claim 24 , wherein the text input logic unit selects one of the groups of words as the plurality of predicted words. 
     
     
         29 . The method of  claim 28 , wherein the selected group of words is selected based on the frequency of the group of words. 
     
     
         30 . The method of  claim 24 , wherein the text input logic unit is configured to update the frequencies of the words based on whether the predicted word is input by a device user. 
     
     
         31 . The method of  claim 23 , wherein for each group, only the maximum frequency, which is the highest frequency value in the group, is retained with full precision, and the frequencies of words with less than the maximum frequency are retained as a percentage of the maximum frequency.

Join the waitlist — get patent alerts

Track US2010211381A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.