US2023075113A1PendingUtilityA1

System and method for unsupervised text normalization using distributed representation of words

Assignee: AT & T IP I LPPriority: Oct 3, 2014Filed: Nov 14, 2022Published: Mar 9, 2023
Est. expiryOct 3, 2034(~8.2 yrs left)· nominal 20-yr term from priority
G06Q 10/40G06F 40/58G06F 40/232G06Q 50/01
75
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system, method and computer-readable storage devices for providing unsupervised normalization of noisy text using distributed representation of words. The system receives, from a social media forum, a word having a non-canonical spelling in a first language. The system determines a context of the word in the social media forum, identifies the word in a vector space model, and selects an “n-best” vector paths in the vector space model, where the n-best vector paths are neighbors to the vector space path based on the context and the non-canonical spelling. The system can then select, based on a similarity cost, a best path from the n-best vector paths and identify a word associated with the best path as the canonical version.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A method comprising:
 composing a correctly-spelled word finite state machine with a finite state transducer, wherein the finite state transducer comprises a vector space model trained from a corpus of noisy text, and wherein words within the finite state transducer are clustered based on context, to yield a modified finite state machine;   receiving a non-canonical spelling for a word, wherein the non-canonical spelling comprises a variant spelling of a canonical spelling of the word;   processing the non-canonical spelling via the modified finite state machine to yield a proposed word; and   outputting the proposed word as a canonical form of the non-canonical spelling, the proposed word determined according to a best path through the modified finite state machine.   
     
     
         2 . The method of  claim 1 , further comprising:
 performing a best path function on the modified finite state machine, wherein the best path function comprises:
 selecting n-best vector paths in a vector space model which are neighbors to the non-canonical spelling; and 
 selecting, based on a similarity cost, the best path from the n-best vector paths. 
   
     
     
         3 . The method of  claim 1 , wherein the non-canonical spelling is classified in the finite state transducer based on a word context. 
     
     
         4 . The method of  claim 1 , wherein the non-canonical spelling comprises a compound word. 
     
     
         5 . The method of  claim 1 , wherein the outputting the proposed word is performed as part of a translation from a first language to a second language. 
     
     
         6 . The method of  claim 2 , wherein the similarity cost is based on a type of the non-canonical spelling. 
     
     
         7 . The method of  claim 6 , wherein the type of the non-canonical spelling is an abbreviation. 
     
     
         8 . A system comprising:
 a processor; and   a computer-readable storage device storing instructions which, when executed by the processor, cause the processor to perform operations, the operations comprising:
 composing a correctly-spelled word finite state machine with a finite state transducer, wherein the finite state transducer comprises a vector space model trained from a corpus of noisy text, and wherein words within the finite state transducer are clustered based on context, to yield a modified finite state machine; 
 receiving a non-canonical spelling for a word, wherein the non-canonical spelling comprises a variant spelling of a canonical spelling of the word; 
 processing the non-canonical spelling via the modified finite state machine to yield a proposed word; and 
 outputting the proposed word as a canonical form of the non-canonical spelling, the proposed word determined according to a best path through the modified finite state machine. 
   
     
     
         9 . The system of  claim 8 , wherein the computer-readable storage device stores additional instructions which, when executed by the processor, cause the processor to perform operations further comprising:
 performing a best path function on the modified finite state machine, wherein the best path function comprises:
 selecting n-best vector paths in a vector space model which are neighbors to the non-canonical spelling; and 
 selecting, based on a similarity cost, the best path from the n-best vector paths. 
   
     
     
         10 . The system of  claim 8 , wherein the non-canonical spelling is classified in the finite state transducer based on a word context. 
     
     
         11 . The system of  claim 8 , wherein the non-canonical spelling comprises a compound word. 
     
     
         12 . The system of  claim 8 , wherein the outputting the proposed word is performed as part of a translation from a first language to a second language. 
     
     
         13 . The system of  claim 9 , wherein the similarity cost is based on a type of the non-canonical spelling. 
     
     
         14 . The system of  claim 13 , wherein the type of the non-canonical spelling is an abbreviation. 
     
     
         15 . A method comprising:
 receiving a non-canonical spelling for a word, wherein the non-canonical spelling comprises a variant spelling of a canonical spelling of the word;   processing the non-canonical spelling via a modified finite state machine to yield a proposed word, wherein the modified finite state machine is generated by composing a correctly-spelled word finite state machine with a finite state transducer, the finite state transducer comprising a vector space model trained from a corpus of noisy text, and wherein words within the finite state transducer are clustered based on context; and   outputting the proposed word as a canonical form of the non-canonical spelling, the proposed word determined according to a best path through the modified finite state machine.   
     
     
         16 . The method of  claim 15 , further comprising:
 performing a best path function on the modified finite state machine, wherein the best path function comprises:
 selecting n-best vector paths in a vector space model which are neighbors to the non-canonical spelling; and 
 selecting, based on a similarity cost, the best path from the n-best vector paths. 
   
     
     
         17 . The method of  claim 15 , wherein the non-canonical spelling is classified in the finite state transducer based on a word context. 
     
     
         18 . The method of  claim 15 , wherein the non-canonical spelling comprises a compound word. 
     
     
         19 . The method of  claim 15 , wherein the outputting the proposed word is performed as part of a translation from a first language to a second language. 
     
     
         20 . The method of  claim 16 , wherein the similarity cost is based on a type of the non-canonical spelling.

Join the waitlist — get patent alerts

Track US2023075113A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.