US12299023B2ActiveUtilityA1

Document retrieval system

Assignee: SEMICONDUCTOR ENERGY LABPriority: Oct 25, 2019Filed: Oct 14, 2020Granted: May 13, 2025
Est. expiryOct 25, 2039(~13.3 yrs left)· nominal 20-yr term from priority
G06F 40/268G06F 40/284G06F 16/3347G06F 40/30G06F 16/332
64
PatentIndex Score
0
Cited by
21
References
8
Claims

Abstract

A document retrieval system that retrieves documents, with concepts of the documents taken into account, is provided. The document retrieval system ( 100 ) includes an input unit ( 101 ), a first processing unit ( 102 ), a storage unit ( 105 ), a second processing unit ( 103 ), and an output unit ( 104 ). The input unit ( 101 ) has a function of inputting a first document ( 20 ), the first processing unit ( 102 ) has a function of creating a first graph structure ( 21 ) from the first document ( 20 ), the storage unit ( 105 ) has a function of storing a second graph structure ( 11 ), the second processing unit ( 103 ) has a function of calculating a similarity between the first graph structure ( 21 ) and the second graph structure ( 11 ), the output unit ( 104 ) has a function of supplying information, the first processing unit ( 102 ) has a function of dividing the first document ( 20 ) into a plurality of tokens, a node and an edge of the first graph structure ( 21 ) have a label, and the label includes the plurality of tokens.

Claims

exact text as granted — not AI-modified
The invention claimed is: 
     
       1. A non-transitory computer readable storage medium having instructions stored thereon which, when executed by one or more processers, cause the one or more processers to perform operations for document retrieval, the operations comprising:
 inputting a first document; 
 creating a first graph structure from the first document; 
 storing a second graph structure; 
 vectorizing the first graph structure and the second graph structure; 
 comparing the vectorized first graph structure and the vectorized second graph structure to perform document retrieval; 
 supplying information; and 
 dividing the first document into a plurality of tokens; 
 wherein an edge of the first graph structure comprises a label, 
 wherein the label comprises the plurality of tokens, and 
 wherein, in a case where the label has an antonym, generating a new graph structure by reversing a direction of the edge of the first graph structure and replacing the label of the edge by the antonym. 
 
     
     
       2. The non-transitory computer readable storage medium according to  claim 1 , the operations further comprising giving a part of speech to a token. 
     
     
       3. The non-transitory computer readable storage medium according to  claim 1 , the operations further comprising:
 performing a modification analysis, and 
 wherein the processing unit is configured to combine some of the tokens in accordance with a result of the modification analysis. 
 
     
     
       4. The non-transitory computer readable storage medium according to  claim 1 , the operations further comprising replacing a token having a representative word or a superordinate by the representative word or the superordinate. 
     
     
       5. The non-transitory computer readable storage medium according to  claim 1 , wherein the second graph structure is created in the processing unit, from a second document. 
     
     
       6. The non-transitory computer readable storage medium according to  claim 1 , the operations further comprising vectorizing the first graph structure and the second graph structure using Weisfeiler-Lehman Kernels. 
     
     
       7. The non-transitory computer readable storage medium according to  claim 2 , the operations further comprising, in a case where a part of speech given to a first token is a noun and a part of speech given to a second token that is placed right before the first token is an adjective, combining the second token to the first token. 
     
     
       8. The non-transitory computer readable storage medium according to  claim 2 , the operations further comprising, in a case where a part of speech given to a third token and a part of speech given to a fourth token that is placed right after the third token are each a noun, combining the third token to the fourth token.

Join the waitlist — get patent alerts

Track US12299023B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.