US2016307000A1PendingUtilityA1

Index-side diacritical canonicalization

Individually held — no corporate assignee on recordPriority: Nov 9, 2010Filed: Nov 9, 2010Published: Oct 20, 2016
Est. expiryNov 9, 2030(~4.3 yrs left)· nominal 20-yr term from priority
G06F 16/3337G06F 40/247G06F 21/64G06F 40/232
35
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for index-side synonym expansion. One method includes obtaining a token sequence for a resource and indexing a particular token in the token sequence. The indexing includes obtaining a diacritically canonicalized form of the particular token; determining that the diacritically canonicalized form of the particular token is different from the particular token; and storing data associating the resource with both the particular token and the different diacritically canonicalized form of the particular token as index terms for the resource in a search engine.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 - 30 . (canceled) 
     
     
         31 . A system comprising:
 one or more computers programmed to perform operations comprising:   obtaining a sequence of tokens that occurs in a resource;   for each of one or more tokens in the sequence of tokens:
 determining that a different, diacritically canonicalized, representative token has been pre-selected for the token, and 
 storing, in an index of index terms associated with the resource, (i) the token, and (ii) the different, diacritically canonicalized, representative token that has been pre-selected for the token, annotated with a prefix identifying the representative token as being derived from the token and not actually being the token; and 
 using, by a search engine, the index to identify the resource as responsive to a search query that includes the diacritically canonicalized, representative token. 
   
     
     
         32 - 50 . (canceled) 
     
     
         51 . The system of  claim 31 , wherein the token sequence comprises tokens extracted from the resource or from metadata for the resource. 
     
     
         52 . The system of  claim 31 , wherein the operations further comprise storing data in the index of index terms that indicates that the different, diacritically canonicalized representative token is a diacritically canonicalized token. 
     
     
         53 . The system of  claim 52 , wherein the data indicates that the representative token is not included in the sequence of tokens in the resource. 
     
     
         54 . The system of  claim 31 , wherein the operations further comprise:
 receiving a search query comprising one or more tokens;   identifying a first token in the search query, wherein the first token comprises one or more characters with diacritical marks;   determining whether a diacritically canonicalized representative token has been pre-selected for first token; and   augmenting the search query to include the diacritically canonicalized representative token for the first token.   
     
     
         55 . The system of  claim 54 , wherein the operations further comprise augmenting the search query to include both (i) the diacritically canonicalized representative token for the first token and (ii) the diacritically canonicalized representative token for the first token with information identifying the diacritically canonicalized representative token as a diacritically canonicalized token. 
     
     
         56 . The system of  claim 54 , wherein the operations further comprise assigning a weight to each token in the augmented search query, including assigning a weight to the diacritically canonicalized representative token so that resources matching the first token in the search query are weighted more highly than resources matching only the diacritically canonicalized representative token. 
     
     
         57 . The system of  claim 31 , wherein the diacritically canonicalized representative token has no characters with any diacritical marks. 
     
     
         58 . The system of  claim 31 , wherein an index term is a token associated with an indexed resource, and wherein the index term is used to identify indexed resources that include a match between the index term and a token of a query. 
     
     
         59 . A computer-implemented method, comprising:
 obtaining a sequence of tokens that occurs in a resource;   for each of one or more tokens in the sequence of tokens:
 determining that a different, diacritically canonicalized, representative token has been pre-selected for the token, and 
 storing, in an index of index terms associated with the resource, (i) the token, and (ii) the different, diacritically canonicalized, representative token that has been pre-selected for the token, annotated with a prefix identifying the representative token as being derived from the token and not actually being the token; and 
   using, by a search engine, the index to identify the resource as responsive to a search query that includes the diacritically canonicalized, representative token.   
     
     
         60 . The method of  claim 59 , wherein the token sequence comprises tokens extracted from the resource or from metadata for the resource. 
     
     
         61 . The method of  claim 59 , further comprising storing data in the index of index terms that indicates that the representative token that is not included in the sequence of tokens is a diacritically canonicalized form. 
     
     
         62 . The method of  claim 61 , wherein the data indicates that the representative token is not included in the sequence of tokens in the resource. 
     
     
         63 . The method of  claim 59 , further comprising:
 receiving a search query comprising one or more tokens;   identifying a first token in the search query, wherein the first token comprises one or more characters with diacritical marks;   determining whether a diacritically canonicalized representative token has been pre-selected for first token; and   augmenting the search query to include the diacritically canonicalized representative token for the first token.   
     
     
         64 . The method of  claim 63 , further comprising augmenting the search query to include both (i) the diacritically canonicalized representative token for the first token and (ii) the diacritically canonicalized representative token for the first token with information identifying the diacritically canonicalized representative token as a diacritically canonicalized token. 
     
     
         65 . The method of  claim 63 , wherein the operations further comprise assigning a weight to each token in the augmented search query, including assigning a weight to the diacritically canonicalized representative token so that resources matching the first token in the search query are weighted more highly than resources matching only the diacritically canonicalized representative token. 
     
     
         66 . The method of  claim 59 , wherein the diacritically canonicalized representative token has no characters with any diacritical marks. 
     
     
         67 . The method of  claim 59 , wherein an index term is a token associated with an indexed resource, and wherein the index term is used to identify indexed resources that include a match between the index term and a token of a query. 
     
     
         68 . A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
 obtaining a sequence of tokens that occurs in a resource;   for each of one or more tokens in the sequence of tokens:
 determining that a different, diacritically canonicalized, representative token has been pre-selected for the token, and 
 storing, in an index of index terms associated with the resource, (i) the token, and (ii) the different, diacritically canonicalized, representative token that has been pre-selected for the token, annotated with a prefix identifying the representative token as being derived from the token and not actually being the token; and 
   using, by a search engine, the index to identify the resource as responsive to a search query that includes the diacritically canonicalized, representative token.   
     
     
         69 . The computer readable medium of  claim 68 , wherein an index term is a token associated with an indexed resource, and wherein the index term is used to identify indexed resources that include a match between the index term and a token of a query.

Join the waitlist — get patent alerts

Track US2016307000A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.