US2017308607A1PendingUtilityA1

Method and System for a Semantic Search Engine

Assignee: INBENTAPriority: Nov 21, 2014Filed: Jul 7, 2017Published: Oct 26, 2017
Est. expiryNov 21, 2034(~8.3 yrs left)· nominal 20-yr term from priority
Inventors:Jordi Torras
G06F 16/3344G06F 17/30684
24
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Semantic Search Engine using Lexical Functions and Meaning-Text Criteria, that outputs a response (R) as the result of a semantic matching process consisting in comparing a natural language query (Q) with a plurality of contents (C), formed of phrases or expressions obtained from a contents' database ( 6 ), and selecting the response (R) as being the contents corresponding to the comparison having a best semantic matching degree. It involves the transformation of the contents (C) and the query in individual words or groups of tokenized words (W 1, W 2 ), which are transformed in its turn into semantic representations (LSC 1, LSC 2 ) thereof, by applying the rules of Meaning Text Theory and through Lexical Functions, the said semantic representations (LSC 1, LSC 2 ) consisting each of a couple formed of a lemma (L) plus a semantic category (SC).

Claims

exact text as granted — not AI-modified
1 . Semantic search engine, that outputs a responses (R) as the result of a semantic matching process comprising in detecting the meanings of a query (Q) and comparing it with detected meanings of contents (C), formed of phrases or expressions obtained from a contents' database ( 6 ), and selecting responses (R) as being contents corresponding to the comparison having semantic matches, comprising the following steps:
 for the query (Q):
 detecting and formalizing all the meanings of the query (Q) into a global semantic representation (LSCS 1 ) that gives the full meaning of the query (Q), by transforming individual or groups of words (W 1 ) of the query (Q) into semantic representations consisting of pairs of lemma (L) plus a semantic category SC (LSC 1 ), retrieved from the lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ), 
 weighting semantic representations LSC 1  in the basis of their category index and their frequency (LSC 1 +FSW 1 ) generating global weighted semantic representations (LSCS 1 +FSWS 1 ) of the query Q, 
   and for every contents (C):
 detecting and formalizing all the meanings of the contents (C) into a global semantic representation (LSCS 2 ) that gives the full meaning of the content (C), by transforming individual or groups of words (W 2 ) of the contents C into semantic representations consisting of pairs of lemma (L) plus a semantic category SC (LSC 2 ), retrieved from the lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ). 
 weighting semantic representations LSC 2  in the basis of their category index and their frequency (LSC 2 +FSW 2 ) generating global weighted semantic representations (LSCS 2 +FSWS 2 ) of the contents C, 
 calculating a semantic matching degree in a matching process, between a global weighted semantic representation (LSCS 1 +FSWS 1 ) of the query (Q) and a global weighted semantic representation (LSCS 2 +FSWS 2 ) of the indexed contents (C), assigning a score 
   and retrieving the contents (C) which have the best matches (score) between their global weighted semantic representation (LSCS 2 +FSWS 2 ) and the query (Q) global weighted semantic representation (LSCS 1 +FSWS 1 ) from the database ( 6 ), and allocate them to respective responses (R).   
     
     
         2 . The semantic search engine of  claim 1 , wherein the step of detecting and formalizing all the meanings of the query (Q), includes applying lexical functions rules (LFR) to change or contract the pairs of lemma (L) plus a semantic category SC (LSC 1 ), retrieved from the lexicon in lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ), generating different versions of the global semantic representation (LSCS 1 ), of the query (Q). 
     
     
         3 . The semantic search engine of  claim 1 , wherein the step of detecting and formalizing all the meanings of the contents (C) applying lexical functions rules (LFR) to the sequence of (LSC 2 ) representing the contents' global meaning (LSCS 2 ) to transform, contract or expand the pairs of lemma (L) plus a semantic category SC (LSC 2 ), retrieved from the lexicon in lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ). 
     
     
         4 . The semantic search engine of  claim 1 , further comprising, prior to the step of calculating a semantic matching degree, indexing, for each global weighted semantic representation (LSCS 2 +FSWS 2 ), each single semantic representation (LSC 2 ), its frequency balanced semantic weight (FSW 2 ), its semantic approximation factor (SAF) and all its expansions (LSC 2 ′) allowed by the lexical functions rules (LFR) and, for each expansion (LSC 2 ′), its semantic approximation factor (SAF′), in the contents (C) into contents database ( 6 ), 
     
     
         5 . The semantic search engine of  claim 1 , wherein transformations of individual or groups of tokenized words (W 1 , W 2 ), both of the query (Q) and of the contents (C), into semantic representations (LSC 1 , LSC 2 ) thereof, are performed by applying the rules of Meaning Text Theory and through Lexical Functions rules and assignments, the semantic representations (LSC 1 , LSC 2 ) consisting each of a couple formed of a lemma (L) plus a semantic category (SC). 
     
     
         6 . The semantic search engine of  claim 1 , further comprising a lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ) consisting of a database with multiple registers ( 100 ), each composed of several fields: an entry word (W); a semantic category (SC) and a lemma (L) of the entry word which are combined to formally represent the meaning of the word (LSC); several other meanings (LSC′) associated to the meaning of the word (LSC) through lexical functions (LF 1 -LF 6 ) comprising at least a synonyms (syn 0 ; syn 1 ; syn 2 , . . . ); contraries; superlatives; adjectives associated to a noun; and verbs associated a noun, and a set of expansion, contraction and transformation rules based on lexical functions associations (LFR). 
     
     
         7 . The semantic search engine of  claim 1 , wherein the lexicon and Lexical Functions assignments and rules (LSCLF) ( 5 ) is implemented in a database regularly updatable on a time basis and on a project-basis. 
     
     
         8 . The semantic search engine of  claim 1 , wherein the matching process for calculating the semantic matching degree between the query (Q) and a contents (C) comprises,
 For each semantic representation (LSC 1 ) of the global semantic representation of the query Q (LSCS 1 ) retrieved by lexical server  4 ,
 assigning a category index (ISC) that is proportional to semantic category (SC) importance, 
 normalizing the category index assignation to get a semantic weight based on category (SWC 1 ) to make sure that all category indexes (ISC) for a global semantic representation of the query Q (LSCS 1 ) add exactly one, by dividing its category index (ISC) by the sum of category indexes of all semantic representation (LSC 1 ) of the global semantic representation of the query Q (LSCS 1 ) 
 assigning a frequency index (FREQ) to each semantic representation (LSC 1 ) through a precalculated meaning-frequency table. 
 calculating and normalizing a frequency balanced semantic weight (FSW 1 ) by dividing the SWC 1  by 1+log2 of the meaning-frequency value (FREQ) for each meaning LSC 1  of the global semantic representation (LSCS 1 ) in query Q and normalizing them in order that all FSW 1  of the global semantic representation of the query Q (LSCS 1 ) add 1. 
   for each semantic representations (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 ),
 assigning a category index (ISC) that is proportional to semantic category (SC) importance, 
 normalizing the category index assignation to get a semantic weight based on category (SWC) to make sure that all category indexes (ISC) for a global semantic representation of the contents C (LSCS 2 ) add exactly one, by dividing its category index (ISC) by the sum of category indexes of all semantic representation (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 ) 
 assigning a frequency index (FREQ) to each semantic representation (LSC 2 ) through a precalculated meaning-frequency table, where each meaning (LSC 2 ) present in the indexed contents C has a computed frequency (FREQ) that takes into account the number of times that it appears in different contents C and the Semantic Approximation Factor (SAF) that defines the quality of each appearance (Considering the SAF of a LSC 2  appearing as itself=1, maximum quality of the appearance). The frequency index application is based on the theory that Information decreases as probability of a meaning increases, based on a logarithmic proportion. 
 calculating and normalizing a frequency balanced semantic weight (FSW) by dividing the SWC by 1+log2 of the meaning-frequency value (FREQ) for each meaning LSC 2  of the global semantic representation (LSCS 2 ) in contents C and normalizing them in order that all FSW of the global semantic representation of the contents C (LSCS 2 ) add 1. 
   
     
     
         9 . The semantic search engine of  claim 1 , wherein, for each semantic representation (LSC 1 ) of the global semantic representation of the query Q (LSCS 1 +FSWS 1 ),
 if its lemma (L) and semantic category (SC) combination (LSC 1 ) matches a semantic representation (LSC 2 ) of the global weighted to semantic representation of the contents C (LSCS 2 +FSWS 2 ), or a semantic representation (LSC 2 ′) assigned to the semantic representation (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 +FSWS 2 ), through a lexical function (LF 1 , LF 2 , LF 3 , . . . ) in a register  100  of the lexicon and Lexical Functions assignments and rules (LSCLF) ( 5 ), then calculate in block  32 , a partial positive similarity as PPS=FSW 1 ×SAF, being SAF a Semantic Approximation Factor varying between 0 and 1, accounting for the semantic distance between LSC 1  and the LSC 2  or the LSC 2 's Lexical functions assignment (LSC 2 ′) matched. SAF allows to point the difference between matching the same meaning (LSC 1 =LSC 2  where SAF=1) or matching a meaning related to the original LSC 2  present in contents C through a lexical function assignment (LSC 1 =LSC 2 's Lfn assignment where SAF=factor attached to the used Lexical Function rule (LFR) that expands the original meaning). In  FIG. 3 , two PPS outputs from block  32  are shown (PPS 1  and PPS 2 ), and   if the semantic representation (LSC 1 ) doesn't match any semantic representation (LSC 2 ) of the global weighted semantic representation of the contents C (LSCS 2 +FSWS 2 ), or a semantic representation (LSC 2 ′) assigned to the semantic representation (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 +FSWS 2 ), through a lexical function (LF 1 , LF 2 , LF 3 , . . . ) then calculate a partial positive similarity as PPS=0.   
     
     
         10 . The semantic search engine of  claim 1 , wherein a Total Positive Similarity (POS_SIM) is calculated, in block  33 , as the sum of all the aforesaid partial positives similarities (PPS) of the global weighted semantic representation (LSCS 1 +FSWS 1 ) of the query (Q). 
     
     
         11 . The semantic search engine of  claim 1 , wherein, for every semantic representation (LSC 2 ) of the global weighted semantic representation (LSCS 2 +FSWS 2 ) of the contents (C) that did not contribute to the total Positive Similarity (POS_SIM), then calculate, at block  32 , a partial negative similarity as PNS=frequency balanced semantic weight (FSW 2 ) of the semantic representation (LSC 2 ) of the global weighted semantic representation (LSCS 2 +FSW 2 ) of the contents C with no correspondence with any semantic representation (LSC 1 ) in the global weighted semantic representations of the query Q (LSCS 1 +FSWS 1 ). 
     
     
         12 . The semantic search engine of  claim 11 , wherein a Total Negative Similarity (NEG_SIM) is calculated, at block  34 , as the sum of all the aforesaid partial negative similarities (PNS) of the global weighted semantic representation (LSCS 2 +FSWS 2 )×Negative weight factor of the contents (C). 
     
     
         13 . The semantic search engine of  claim 12 , wherein, for each content (C) a coincidence score (COINC 1 ; COINC 2 ) is calculated, at block  35 , as the difference between the Total Positive Similarity (POS_SIM) and Total Negative Similarity (NEG_SIM) when the linguistic type for the content C is phrase, and calculated taking Total Positive Similarity (POS_SIM) value as the coincidence score (COINC 1 ; COINC 2 ) when the linguistic type for the content (C) is free text. 
     
     
         14 . The semantic search engine of  claim 13 , wherein, in block  36 , a semantic matching degree between the query (Q) and a content (C) is calculated for each coincidence (COINC 1 ; COINC 2 ) between the global weighted semantic representation of the query Q (LSCS 1 +FSWS 1 ) and the global semantic representation (LSCS 2 +FSWS 2 ) of the content (C), as the coincidence (COINC 1 ; COINC 2 ) for the REL factor (reliability of the matching) of content (C). 
     
     
         15 . The semantic search engine of  claim 14 , wherein that a response (R) to the query (Q) is selected as the content (C) having the higher semantic matching degree.

Join the waitlist — get patent alerts

Track US2017308607A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.