Method and System for a Semantic Search Engine
Abstract
Semantic Search Engine using Lexical Functions and Meaning-Text Criteria, that outputs a response (R) as the result of a semantic matching process consisting in comparing a natural language query (Q) with a plurality of contents (C), formed of phrases or expressions obtained from a contents' database ( 6 ), and selecting the response (R) as being the contents corresponding to the comparison having a best semantic matching degree. It involves the transformation of the contents (C) and the query in individual words or groups of tokenized words (W 1, W 2 ), which are transformed in its turn into semantic representations (LSC 1, LSC 2 ) thereof, by applying the rules of Meaning Text Theory and through Lexical Functions, the said semantic representations (LSC 1, LSC 2 ) consisting each of a couple formed of a lemma (L) plus a semantic category (SC).
Claims
exact text as granted — not AI-modified1 . Semantic search engine, that outputs a responses (R) as the result of a semantic matching process comprising in detecting the meanings of a query (Q) and comparing it with detected meanings of contents (C), formed of phrases or expressions obtained from a contents' database ( 6 ), and selecting responses (R) as being contents corresponding to the comparison having semantic matches, comprising the following steps:
for the query (Q):
detecting and formalizing all the meanings of the query (Q) into a global semantic representation (LSCS 1 ) that gives the full meaning of the query (Q), by transforming individual or groups of words (W 1 ) of the query (Q) into semantic representations consisting of pairs of lemma (L) plus a semantic category SC (LSC 1 ), retrieved from the lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ),
weighting semantic representations LSC 1 in the basis of their category index and their frequency (LSC 1 +FSW 1 ) generating global weighted semantic representations (LSCS 1 +FSWS 1 ) of the query Q,
and for every contents (C):
detecting and formalizing all the meanings of the contents (C) into a global semantic representation (LSCS 2 ) that gives the full meaning of the content (C), by transforming individual or groups of words (W 2 ) of the contents C into semantic representations consisting of pairs of lemma (L) plus a semantic category SC (LSC 2 ), retrieved from the lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ).
weighting semantic representations LSC 2 in the basis of their category index and their frequency (LSC 2 +FSW 2 ) generating global weighted semantic representations (LSCS 2 +FSWS 2 ) of the contents C,
calculating a semantic matching degree in a matching process, between a global weighted semantic representation (LSCS 1 +FSWS 1 ) of the query (Q) and a global weighted semantic representation (LSCS 2 +FSWS 2 ) of the indexed contents (C), assigning a score
and retrieving the contents (C) which have the best matches (score) between their global weighted semantic representation (LSCS 2 +FSWS 2 ) and the query (Q) global weighted semantic representation (LSCS 1 +FSWS 1 ) from the database ( 6 ), and allocate them to respective responses (R).
2 . The semantic search engine of claim 1 , wherein the step of detecting and formalizing all the meanings of the query (Q), includes applying lexical functions rules (LFR) to change or contract the pairs of lemma (L) plus a semantic category SC (LSC 1 ), retrieved from the lexicon in lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ), generating different versions of the global semantic representation (LSCS 1 ), of the query (Q).
3 . The semantic search engine of claim 1 , wherein the step of detecting and formalizing all the meanings of the contents (C) applying lexical functions rules (LFR) to the sequence of (LSC 2 ) representing the contents' global meaning (LSCS 2 ) to transform, contract or expand the pairs of lemma (L) plus a semantic category SC (LSC 2 ), retrieved from the lexicon in lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ).
4 . The semantic search engine of claim 1 , further comprising, prior to the step of calculating a semantic matching degree, indexing, for each global weighted semantic representation (LSCS 2 +FSWS 2 ), each single semantic representation (LSC 2 ), its frequency balanced semantic weight (FSW 2 ), its semantic approximation factor (SAF) and all its expansions (LSC 2 ′) allowed by the lexical functions rules (LFR) and, for each expansion (LSC 2 ′), its semantic approximation factor (SAF′), in the contents (C) into contents database ( 6 ),
5 . The semantic search engine of claim 1 , wherein transformations of individual or groups of tokenized words (W 1 , W 2 ), both of the query (Q) and of the contents (C), into semantic representations (LSC 1 , LSC 2 ) thereof, are performed by applying the rules of Meaning Text Theory and through Lexical Functions rules and assignments, the semantic representations (LSC 1 , LSC 2 ) consisting each of a couple formed of a lemma (L) plus a semantic category (SC).
6 . The semantic search engine of claim 1 , further comprising a lexicon and Lexical Functions assignments and rules (LSCLF) database ( 5 ) consisting of a database with multiple registers ( 100 ), each composed of several fields: an entry word (W); a semantic category (SC) and a lemma (L) of the entry word which are combined to formally represent the meaning of the word (LSC); several other meanings (LSC′) associated to the meaning of the word (LSC) through lexical functions (LF 1 -LF 6 ) comprising at least a synonyms (syn 0 ; syn 1 ; syn 2 , . . . ); contraries; superlatives; adjectives associated to a noun; and verbs associated a noun, and a set of expansion, contraction and transformation rules based on lexical functions associations (LFR).
7 . The semantic search engine of claim 1 , wherein the lexicon and Lexical Functions assignments and rules (LSCLF) ( 5 ) is implemented in a database regularly updatable on a time basis and on a project-basis.
8 . The semantic search engine of claim 1 , wherein the matching process for calculating the semantic matching degree between the query (Q) and a contents (C) comprises,
For each semantic representation (LSC 1 ) of the global semantic representation of the query Q (LSCS 1 ) retrieved by lexical server 4 ,
assigning a category index (ISC) that is proportional to semantic category (SC) importance,
normalizing the category index assignation to get a semantic weight based on category (SWC 1 ) to make sure that all category indexes (ISC) for a global semantic representation of the query Q (LSCS 1 ) add exactly one, by dividing its category index (ISC) by the sum of category indexes of all semantic representation (LSC 1 ) of the global semantic representation of the query Q (LSCS 1 )
assigning a frequency index (FREQ) to each semantic representation (LSC 1 ) through a precalculated meaning-frequency table.
calculating and normalizing a frequency balanced semantic weight (FSW 1 ) by dividing the SWC 1 by 1+log2 of the meaning-frequency value (FREQ) for each meaning LSC 1 of the global semantic representation (LSCS 1 ) in query Q and normalizing them in order that all FSW 1 of the global semantic representation of the query Q (LSCS 1 ) add 1.
for each semantic representations (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 ),
assigning a category index (ISC) that is proportional to semantic category (SC) importance,
normalizing the category index assignation to get a semantic weight based on category (SWC) to make sure that all category indexes (ISC) for a global semantic representation of the contents C (LSCS 2 ) add exactly one, by dividing its category index (ISC) by the sum of category indexes of all semantic representation (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 )
assigning a frequency index (FREQ) to each semantic representation (LSC 2 ) through a precalculated meaning-frequency table, where each meaning (LSC 2 ) present in the indexed contents C has a computed frequency (FREQ) that takes into account the number of times that it appears in different contents C and the Semantic Approximation Factor (SAF) that defines the quality of each appearance (Considering the SAF of a LSC 2 appearing as itself=1, maximum quality of the appearance). The frequency index application is based on the theory that Information decreases as probability of a meaning increases, based on a logarithmic proportion.
calculating and normalizing a frequency balanced semantic weight (FSW) by dividing the SWC by 1+log2 of the meaning-frequency value (FREQ) for each meaning LSC 2 of the global semantic representation (LSCS 2 ) in contents C and normalizing them in order that all FSW of the global semantic representation of the contents C (LSCS 2 ) add 1.
9 . The semantic search engine of claim 1 , wherein, for each semantic representation (LSC 1 ) of the global semantic representation of the query Q (LSCS 1 +FSWS 1 ),
if its lemma (L) and semantic category (SC) combination (LSC 1 ) matches a semantic representation (LSC 2 ) of the global weighted to semantic representation of the contents C (LSCS 2 +FSWS 2 ), or a semantic representation (LSC 2 ′) assigned to the semantic representation (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 +FSWS 2 ), through a lexical function (LF 1 , LF 2 , LF 3 , . . . ) in a register 100 of the lexicon and Lexical Functions assignments and rules (LSCLF) ( 5 ), then calculate in block 32 , a partial positive similarity as PPS=FSW 1 ×SAF, being SAF a Semantic Approximation Factor varying between 0 and 1, accounting for the semantic distance between LSC 1 and the LSC 2 or the LSC 2 's Lexical functions assignment (LSC 2 ′) matched. SAF allows to point the difference between matching the same meaning (LSC 1 =LSC 2 where SAF=1) or matching a meaning related to the original LSC 2 present in contents C through a lexical function assignment (LSC 1 =LSC 2 's Lfn assignment where SAF=factor attached to the used Lexical Function rule (LFR) that expands the original meaning). In FIG. 3 , two PPS outputs from block 32 are shown (PPS 1 and PPS 2 ), and if the semantic representation (LSC 1 ) doesn't match any semantic representation (LSC 2 ) of the global weighted semantic representation of the contents C (LSCS 2 +FSWS 2 ), or a semantic representation (LSC 2 ′) assigned to the semantic representation (LSC 2 ) of the global semantic representation of the contents C (LSCS 2 +FSWS 2 ), through a lexical function (LF 1 , LF 2 , LF 3 , . . . ) then calculate a partial positive similarity as PPS=0.
10 . The semantic search engine of claim 1 , wherein a Total Positive Similarity (POS_SIM) is calculated, in block 33 , as the sum of all the aforesaid partial positives similarities (PPS) of the global weighted semantic representation (LSCS 1 +FSWS 1 ) of the query (Q).
11 . The semantic search engine of claim 1 , wherein, for every semantic representation (LSC 2 ) of the global weighted semantic representation (LSCS 2 +FSWS 2 ) of the contents (C) that did not contribute to the total Positive Similarity (POS_SIM), then calculate, at block 32 , a partial negative similarity as PNS=frequency balanced semantic weight (FSW 2 ) of the semantic representation (LSC 2 ) of the global weighted semantic representation (LSCS 2 +FSW 2 ) of the contents C with no correspondence with any semantic representation (LSC 1 ) in the global weighted semantic representations of the query Q (LSCS 1 +FSWS 1 ).
12 . The semantic search engine of claim 11 , wherein a Total Negative Similarity (NEG_SIM) is calculated, at block 34 , as the sum of all the aforesaid partial negative similarities (PNS) of the global weighted semantic representation (LSCS 2 +FSWS 2 )×Negative weight factor of the contents (C).
13 . The semantic search engine of claim 12 , wherein, for each content (C) a coincidence score (COINC 1 ; COINC 2 ) is calculated, at block 35 , as the difference between the Total Positive Similarity (POS_SIM) and Total Negative Similarity (NEG_SIM) when the linguistic type for the content C is phrase, and calculated taking Total Positive Similarity (POS_SIM) value as the coincidence score (COINC 1 ; COINC 2 ) when the linguistic type for the content (C) is free text.
14 . The semantic search engine of claim 13 , wherein, in block 36 , a semantic matching degree between the query (Q) and a content (C) is calculated for each coincidence (COINC 1 ; COINC 2 ) between the global weighted semantic representation of the query Q (LSCS 1 +FSWS 1 ) and the global semantic representation (LSCS 2 +FSWS 2 ) of the content (C), as the coincidence (COINC 1 ; COINC 2 ) for the REL factor (reliability of the matching) of content (C).
15 . The semantic search engine of claim 14 , wherein that a response (R) to the query (Q) is selected as the content (C) having the higher semantic matching degree.Join the waitlist — get patent alerts
Track US2017308607A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.