Method and apparatus for data processing and word processing in Chinese using a phonetic Chinese language
Abstract
A method and apparatus for data processing and word processing in the Chinese language. A Phonetic Chinese Language (PCL) is defined in which any ideogram can be unambiguously represented by a Phonetic Chinese Word (PCW) no more than four characters in length, each word being composed of letters selected from a defined set of letters that can each be uniquely represented by a 7-bit digital code. Each PCW represents one and only one ideogram and provides the full sound and tone information required to pronounce it. Ambiguities caused by homonyms and homotones are avoided. PCL words are translated into their corresponding ideograms and vice versa by means of a stored monosyllabic dictionary. A method for unambiguously separating a polysyllabic PCL character string into separate words is also provided, which makes it unnecessary to employ a polysyllabic dictionary. Also disclosed is a method of forming an alphagrammic listing from PCL character strings by separating the strings into separate characters and listing them in alphabetical order, provided that homotones and identical ideograms are grouped together even if strict alphabetical ordering of the string would have separated them. The disclosure also includes a keyboard adapted for efficiently entering PCL characters for processing.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A method of digitally encoding and storing the ideographic Chinese language in a computer, comprising the steps of: a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being pronounced as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting one and only one digital representation for each selected ideogram which is usable in said computer for outputting said ideograms; c) selecting a set of letters for a phonetic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at least one such PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms and distinguish between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms; d) selecting one and only one digital representation for each PCA letter which is usable in said computer for outputting said PCA letter; and e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representations of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideogram and its corresponding PCW including distinguishing between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms.
2. A method as in claim 1, wherein said PCA letters represent the following language elements: a) a plurality of vowels; b) a plurality of tones with which said vowels are pronounced; and c) a plurality of consonants.
3. A method of digitally encoding and storing the ideographic Chinese language in a computer, comprising the steps of: 1) a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being pronounced as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting one and only one digital representation for each selected ideogram which is usable in said computer for outputting said ideogram; c) selecting a set of letters for a phonetic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at least one such PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms; d) selecting one and only one digital representation for each PCA letter which is usable in said computer for outputting said PCA letter; and e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representations of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideograms and its corresponding PCW; 2) wherein said PCA letters represent the following language elements; a) a plurality of vowels; b) a plurality of tones with which said vowels are pronounced; and c) a plurality of consonants; and 3) wherein said vowels include a) a plurality of voweltones, each of which represents a given vowel sound pronounced with a given tone, and b) a plurality of semi-consonants, each of which represents a given vowel sound irrespective of tone.
4. A method as in claim 3, wherein said plurality of tones includes four tones.
5. A method as in claim 4, wherein each of said voweltones comprises a base character and an indicia incorporated therein which indicates the tone.
6. A method of digitally encoding and storing the ideographic Chinese language in a computer, comprising the steps of: 1) a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being pronounced as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting one and only one digital representation for each selected ideogram which is usable in said computer for outputting said ideogram; c) selecting a set of letters for a phonetic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at least one PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms; d) selecting one and only one digital representation for each PCA letter which is usable in said computer for outputting said PCA letter; and e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representations of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideogram and its corresponding PCW; 2) wherein said PCA letters represent the following language elements; a) a plurality of vowels; b) a plurality of tones with which said vowels are pronounced; and c) a plurality of consonants; and 3) wherein said consonants include a) a plurality of short consonants, each of which represents a respective consonant sound; b) a plurality of long consonants, each of which represents a respective consonant sound pronounced with a respective vowel sound; and c) a silent zero consonant.
7. A method of digitally encoding and storing the ideographic Chinese language in a computer, comprising the steps of: 1) a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being pronounced as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting one and only one digital representation for each selected ideogram which is usable in said computer for outputting said ideogram; c) selecting a set of letters for a phonetic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at least one PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms; d) selecting one and only one digital representation for each PCA letter which is usable in said computer for outputting said PCA letter; and e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representation of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideogram and its corresponding PCW; 2) wherein said PCA letters represent the following language elements; a) a plurality of vowels; b) a plurality of tones with which said vowels are pronounced; and c) a plurality of consonants; 3) wherein each such PCW has the form TS+Q, wherein a) TS is a tone-syllable having one of the forms CV, CSV, SV, and V; C being a consonant, S being a semi-consonant, and V being a voweltone; and b) Q is a generalized tone-syllable modifier which indicates meaning for distinguishing between homotones.
8. A method as in claim 7, wherein Q has one of the forms φ and G, wherein a) φ is the null set; and b) G is a generalized semantic classifier comprising a PCA letter added to the tone-syllable TS to the extend necessary for distinguishing between homotones.
9. A method as in claim 8, wherein G has one of the forms C, V, S and Z, wherein Z is the zero consonant.
10. A method as in claim 4, wherein a vowel sound "i" is represented by three groups of distinct PCA letters.
11. A method as in claim 10, wherein a vowel sound "u" and a vowel sound "u" are each represented by two groups of distinct PCA letters.
12. A method as in claim 11, wherein the PCA can distinguish between 255 homotones for PCWs wherein the only vowel sound is "i", "u", or "u"; 170 homotones for PCWs ending in the vowel sound "i"; and 85 homotones for all other PCWs.
13. A method as in claim 6, wherein said plurality of tones includes four tones; said vowels including a plurality of voweltones, each of which represents a given vowel sound pronounced with a given tone, and a plurality of semi-consonants, each of which represents a given vowel sound irrespective of tone; four of said voweltones respectively representing the four tones; and further representing the vowel sound "i" when they follow on of said short consonants.
14. A method as in claim 13, wherein four of said voweltones respectively represent the vowel sound "e" pronounced with said four tones; but represent the vowel sound "o" when they follow the sounds "b", "p", "m", and "f" and the semi-consonants.
15. A method as in claim 14, wherein four of said voweltones respectively represent the four tones; and further represent the vowel sound "er" when they are written alone or following the zero consonant; and further represent the vowel sound "i" when they follow the short consonants.
16. A method as in claim 9, comprising selecting a primary set of at least about 8000 ideograms which are those most frequently used in the Chinese language.
17. A method as in claim 16, wherein at least about 3900 ideograms of said primary set, which account for at least about 97 percent of usage, are uniquely identified by PCWs having one of the forms TS+φ, TS+V*, and TS+Z, V* being the same voweltone as that in the tone-syllable TS.
18. A method as in claim 17, wherein all of the remaining ideograms of the Chinese language are uniquely identified by PCWs having the form TS+G, where G is a PCA letter other than V* or Z.
19. A method as in claim 18, wherein at least about 80 percent of the remaining approximately 4100 ideograms of the primary set are each uniquely identified by employing a semantic classifier G which is a PCA letter similar to an ideographic radical having a meaning similar to that of the ideogram to be identified.
20. A method as in claim 1, wherein each PCW comprises no more than 4 PCA letters.
21. A method as in claim 20, wherein each PCW comprises a frequency-weighted average of 2.4 PCA letters.
22. A text processing method which includes digitally encoding and storing the ideographic Chinese language in a computer, comprising the steps of: a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being pronounced as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting one and only one digital representation for each selected ideogram which is usable in said computer for outputting said ideogram; c) selecting a set of letters for a phonetic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at least one such PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms and distinguish between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms; d) selecting one and only one digital representation for each PCA letter which is usable in said computer for outputting said PCA letter; e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representations of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideograms and its corresponding PCW including distinguishing between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms; entering a continuous string of phonetic Chinese language characters into said computer memory, said string of characters including at least two groups of characters, each group of characters defining a phonetic Chinese word of variable character length; and processing said continuous string in said computer memory so as to accurately determine the beginning and end of each phonetic Chinese word in said string.
23. A method as in claim 22, further comprising the step of referring to the stored monosyllabic dictionary to unambiquously determine the one and only one ideogram corresponding to each such phonetic Chinese word.
24. A method of creating an alphagrammic listing of a set of word strings, which includes digitally encoding and storing the ideographic Chinese language in a computer, the method comprising the steps of: a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being pronounced as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting one and only one digital representation for each selected ideogram which is usable in said computer for outputting said ideogram; c) selecting a set of letters for a phonetic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at last one such PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms and distinguish between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms; d) selecting one and only one digital representation for each PCA letter which is usable in said computer for outputting said PCA letter; e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representations of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideogram and its corresponding PCW including distinguishing between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms; each word string including a plurality of phonetic Chinese words, each phonetic Chinese word (PCW) representing one and only one Chinese ideogram and providing the sound and tone information required to pronounce that ideogram, and distinguishing between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms, said PCA letters having a predetermined alphabetical order, said method of creating an alphagrammic listing comprising the steps of: 1) storing said set of word strings in the computer memory; and 2) sorting said set of word strings in alphagrammic order, wherein 3) said word strings are listed in the alphabetical order of the characters in that word string; 4) said alphabetical order being overridded to the extend that; (a) all strings whose corresponding first Chinese ideograms are identical are listed together for purposes of ordering said strings; and (b) all words in said word strings pronounced with the same sound and tone are listed together for purposes of ordering said strings; 5) all strings listed together in said steps (a) and (b) being listed in alphabetical order with respect to one another.
25. A method of processing character strings, comprising a) entering a string of letters of a phonetic Chinese alphabet (PCA) in a computer memory; wherein 1) said PCA includes respective pluralities of voweltones (V), semi-consonants (S), and consonants (C), and including a zero consonant (Z); 2) said string of letters includes at least two separate phonetic Chinese words (PCWs), each said PCW having the form TS+Q, wherein TS is a tone-syllable having one of the forms CV, CSV, SV and V, and Q is a generalized meaning-indicating modifier having one of two forms, namely a PCA letter and the omission of any PCA letter; provided that Q cannot take the form of one voweltone (RV) which is employed to indicate the retroflex ideogram when it occurs at the end of a character string; 3) each of said PCWs represents one and only one Chinese ideogram and provides the sound and tone information required to pronounce that ideogram; and 4) each non-initial PCW that has the form V+Q is preceded in such string by the zero consonant, and each noninitial PCW that has the form SV+Q is preceded in such string by the zero consonant whenever such last-mentioned PCW follows a PCW having one of the forms CVC and CSVC; and b) separating said string in said computer memory unambiguously into said separate phonetic Chinese words included therein.
26. A method as in claim 25, further comprising a step of referring to a stored monosyllabic dictionary to unambiguously determine the ideogram corresponding to each such phonetic Chinese word.
27. A method as in claim 25, further comprising a) defining a predetermined alphabetical order for said PCA letters; b) entering at least two of said strings of PCA letters in a computer memory; and c) sorting said strings in said computer memory in alphagrammic order wherein 1) said strings are listed in the alphabetical order of the letters in that string, 2) said alphabetical order being overridden to the extent that i) all strings whose corresponding first Chinese ideograms are identical are listed together for purposes of alphabetization, and ii) all PCWs in said strings pronounced with the same sound and tone are listed together for purposes of alphabetization of said strings; all strings listed together in said steps (2) (i) and (2) (ii) being listed in alphabetical order with respect to one another.
28. A method of encoding and storing Chinese ideograms in a computer, comprising the steps of: a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being pronounced as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting a set of letters for a phonetic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at least one such PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms; c) selecting one and only one 7-bit digital representation for each selected PCA letter and each selected ideogram which are usable in said computer for outputting said ideograms and said PCA letters; d) selecting one and only one phonetic Chinese word (PCW) composed of PCA letters for uniquely identifying each selected ideogram; and e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representations of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideograms and its corresponding PCW, including distinguishing between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms.
29. A method as in claim 28, wherein said 7-bit digital representation for each PCA letter is within the range 80H-FFH.
30. A method as in claim 29, wherein said 7-bit digital representation is within the range 80H-DFH.
31. A method as in claim 28, wherein said 7-bit digital representation is within the range 81H-DEH.
32. A method as in claim 24, further comprising the step of referring to the stored monosyllabic dictionary to unambiguously determine the one and only one ideogram corresponding to each such phonetic Chinese word.
33. A method as in claim 25, further comprising digitally encoding and storing the ideographic Chinese language in said computer by the steps of: a) selecting a set of Chinese ideograms to be encoded and stored, each of said Chinese ideograms being processed as a monosyllable having a predetermined consonant sound, vowel sound, and vowel tone; b) selecting one and only one digital representation for each selected ideogram which is usable in said computer for outputting said ideogram; c) selecting a set of letters for a phenolic Chinese alphabet (PCA) which can be formed into phonetic Chinese words (PCWs) each comprising at least one such PCA letter, which fully identify the sound and tone pronunciation of such selected ideograms and distinguish between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms; d) selecting one and only one digital representation for each PCA letter letter which is usable in said computer for outputting said PCA letter; and e) storing a monosyllabic dictionary in a computer memory in said computer which associates the digital representations of said ideograms and PCA letters so as to identify a one-to-one relationship between the respective digital representations of each selected ideogram and its corresponding PCW including distinguishing between all homotone ideograms having identical sound and tone pronunciation in said selected set of Chinese ideograms.Join the waitlist — get patent alerts
Track US5175803A — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.