US2016062979A1PendingUtilityA1

Word classification based on phonetic features

Assignee: GOOGLE INCPriority: Aug 27, 2014Filed: Sep 5, 2014Published: Mar 3, 2016
Est. expiryAug 27, 2034(~8.1 yrs left)· nominal 20-yr term from priority
G06F 40/279G06F 40/242G06F 16/90344G06F 40/30G06F 17/2785G06F 17/27G06F 17/2765G06F 17/2735G06F 40/237
52
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining a textual term; determining, by one or more computers, a vector representing a phonetic feature of the textual term; comparing the vector representing the phonetic feature of the textual term with a reference vector representing a phonetic feature of a reference textual term; and classifying the textual term based on the comparing the vector with the reference vector.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method, comprising:
 obtaining a textual term;   determining, by one or more computers, a vector representing a phonetic feature of the textual term;   comparing the vector representing the phonetic feature of the textual term with a reference vector representing a phonetic feature of a reference textual term; and   classifying the textual term based on the comparing the vector with the reference vector.   
     
     
         2 . The method of  claim 1 , wherein obtaining the textual term comprises obtaining the textual term from a resource stored at a remote computer. 
     
     
         3 . The method of  claim 1 , wherein obtaining the textual term comprises obtaining the textual term from a search query. 
     
     
         4 . The method of  claim 1 , wherein determining the vector representing the phonetic feature of the textual term comprises determining a pronunciation of the textual term. 
     
     
         5 . The method of  claim 1 ,
 wherein the textual term includes a plurality of characters,   wherein determining the vector representing the phonetic feature of the textual term comprises determining a first vector representing a phonetic feature of a subset of the plurality of characters of the textual term, and   wherein classifying the textual term comprises:
 determining that the subset of the plurality of characters of the textual term is similar to the reference textual term based on the comparing the vector with the reference vector; and 
 in response to determining that the subset of the plurality of characters of the textual term is similar to the reference textual term, associating a definition of the reference textual term to the subset of the plurality of characters of the textual term. 
   
     
     
         6 . The method of  claim 5 , comprising:
 determining a second vector representing a phonetic feature of a second subset of the plurality of characters of the textual term;   comparing the second vector with a second reference vector representing a phonetic feature of a second reference textual term;   determining that the second subset of the plurality of characters is similar to the second reference textual term based on comparing the second vector with the second reference vector; and   in response to determining that the second subset of the plurality of characters is similar to the second reference textual term, associating a definition of the second reference textual term to the second subset of the plurality of characters of the textual term.   
     
     
         7 . The method of  claim 1 , wherein comparing the vector with the reference vector comprises determining a cosine distance between the vector and the reference vector. 
     
     
         8 . The method of  claim 7 , wherein classifying the textual term comprises:
 determining that the cosine distance is within a specific distance; and   in response to determining that the cosine distance is within the specific distance, classifying the textual term as being similar to the reference textual term.   
     
     
         9 . The method of  claim 1 , wherein classifying the textual term comprises
 determining a likelihood that the textual term is similar to the reference textual term; and.   classifying the textual term based on the likelihood at the textual term is similar to the reference textual term.   
     
     
         10 . The method of  claim 1 , wherein classifying the textual term comprises:
 determining that the textual term is similar to the reference textual term; and   in response to determining that the textual term is similar to the reference textual term, associating a definition of the reference textual term to the textual term.   
     
     
         11 . The method of  claim 1 ,
 wherein obtaining the textual term comprises obtaining one or more textual terms that are surrounding the textual term, and   wherein determining the vector representing the phonetic feature of the textual term comprises determining the vector using (i) the phonetic feature of the textual term and (ii) the one or more textual terms that are surrounding the textual term.   
     
     
         12 . A computer-readable medium storing software having stored thereon instructions, which, when executed by one or more computers, cause the one or more computers to perform operations of:
 obtaining a textual term;   determining, by one or more computers, a vector representing a phonetic feature of the textual term;   comparing the vector representing the phonetic feature of the textual term with a reference vector representing a phonetic feature of a reference textual term; and   classifying the textual term based on the comparing the vector with the reference vector.   
     
     
         13 . The computer-readable medium of  claim 12 ,
 wherein the textual term includes a plurality of characters,   wherein determining the vector representing the phonetic feature of the textual term comprises determining a first vector representing a phonetic feature of a subset of the plurality of characters of the textual term, and   wherein classifying the textual term comprises:
 determining that the subset of the plurality of characters of the textual term is similar to the reference textual term based on the comparing the vector with the reference vector; and 
 in response to determining that the subset of the plurality of characters of the textual term is similar to the reference textual term, associating a definition of the reference textual term to the subset of the plurality of characters of the textual term. 
   
     
     
         14 . The computer-readable medium of  claim 13 , wherein the operations comprise:
 determining a second vector representing a phonetic feature of a second subset of the plurality of characters of the textual term;   comparing the second vector with a second reference vector representing a phonetic feature of a second reference textual term;   determining that the second subset of the plurality of characters is similar to the second reference textual term based on comparing the second vector with the second reference vector; and   in response to determining that the second subset of the plurality of characters is similar to the second reference textual term, associating a definition of the second reference textual term to the second subset of the plurality of characters of the textual term.   
     
     
         15 . The computer-readable medium of  claim 12 , wherein comparing the vector with the reference vector comprises determining a cosine distance between the vector and the reference vector. 
     
     
         16 . The computer-readable medium of  claim 12 ,
 wherein obtaining the textual term comprises obtaining one or more textual terms that are surrounding the textual term, and   wherein determining the vector representing the phonetic feature of the textual term comprises determining the vector using (i) the phonetic feature of the textual term and (ii) the one or more textual terms that are surrounding the textual term.   
     
     
         17 . A system comprising:
 one or more processors and one or more computer storage media storing instructions that are operable, when executed by the one or more processors, to cause the one or more processors to perform operations comprising:   obtaining a textual term;   determining, by one or more computers, a vector representing a phonetic feature of the textual term;   comparing the vector representing the phonetic feature of the textual term with a reference vector representing a phonetic feature of a reference textual term; and   classifying the textual term based on the comparing the vector with the reference vector.   
     
     
         18 . The system of  claim 17 ,
 wherein the textual term includes a plurality of characters,   wherein determining the vector representing the phonetic feature of the textual term comprises determining a first vector representing a phonetic feature of a subset of the plurality of characters of the textual term, and   wherein classifying the textual term comprises:
 determining that the subset of the plurality of characters of the textual term is similar to the reference textual term based on the comparing the vector with the reference vector; and 
 in response to determining that the subset of the plurality of characters of the textual term is similar to the reference textual term, associating a definition of the reference textual term to the subset of the plurality of characters of the textual term. 
   
     
     
         19 . The system of  claim 18 , wherein the operations comprise:
 determining a second vector representing a phonetic feature of a second subset of the plurality of characters of the textual term;   comparing the second vector with a second reference vector representing a phonetic feature of a second reference textual term;   determining that the second subset of the plurality of characters is similar to the second reference textual term based on comparing the second vector with the second reference vector; and   in response to determining that the second subset of the plurality of characters is similar to the second reference textual term, associating a definition of the second reference textual term to the second subset of the plurality of characters of the textual term.   
     
     
         20 . The system of  claim 17 ,
 wherein obtaining the textual term comprises obtaining one or more textual terms that are surrounding the textual term, and   wherein determining the vector representing the phonetic feature of the textual term comprises determining the vector using (i) the phonetic feature of the textual term and (ii) the one or more textual terms that are surrounding the textual term.

Join the waitlist — get patent alerts

Track US2016062979A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.