US2024004930A1PendingUtilityA1

System and method for pre-indexing filtering and correction of documents in search systems

Assignee: OPEN TEXT HOLDINGS INCPriority: Sep 25, 2019Filed: Jul 10, 2023Published: Jan 4, 2024
Est. expirySep 25, 2039(~13.1 yrs left)· nominal 20-yr term from priority
G06F 16/901G06F 16/93G06F 40/10G06F 16/31G06F 16/144G06F 40/284G06F 16/2465G06F 16/316
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Embodiments as disclosed herein provide a search system with an pre-indexing filter that provides both a sophisticated and contextually tailored approach to filtering documents and a corrector that is adapted to alter a document that has been designated to be filtered out from the indexing process and determine if the altered document should be indexed. The alteration of the document may be tied to the attributes, rules or thresholds used to initially filter the document from the indexing process. The filtering criteria can thus be tailored to a specific context such that both the initial filtering and the alteration process may be better suited for application in that context.

Claims

exact text as granted — not AI-modified
1 - 21 . (canceled) 
     
     
         22 . A search system, comprising:
 a processor;   a non-transitory computer readable medium, having instructions executable on the processor for:
 receiving a first set of tokens for a document from a text extractor; 
 obtaining a detector score indicating that each of the received first set of tokens is of an associated type of token; 
 obtaining a filter score based on the detector score, the filter score based on a scoring rule associated with the associated type; 
 determining whether the document should be indexed based on an application of a verdict rule, the verdict rule comprising an expression for evaluating the filter score; and 
 in response to determining that the document should be indexed, providing the document to an indexer adapted to index the first set of tokens for the document. 
   
     
     
         23 . The search system of  claim 22 , wherein in response to determining that the document should not be indexed, discarding the document based on a suitability criteria. 
     
     
         24 . The search system of  claim 22 , wherein in response to determining that the document should not be indexed, initiating forwarding of the document to a corrector. 
     
     
         25 . The search system of  claim 24 , wherein the corrector is adapted to:
 create an altered document comprising a second set of tokens, and   determine whether the altered document should be indexed based on a suitability criteria.   
     
     
         26 . The search system of  claim 25 , wherein the suitability criteria is document size or a number of the second set of tokens. 
     
     
         27 . The search system of  claim 22 , wherein indexing the document comprises indexing the first set of tokens. 
     
     
         28 . The search system of  claim 22 , wherein obtaining the detector score comprises initiating evaluation of each of the received first set of tokens by one or more detectors that are adapted to produce the detector score. 
     
     
         29 . The search system of  claim 22 , wherein the associated type comprises an attribute or type that is associated with the token or the document. 
     
     
         30 . A non-transitory computer readable medium, comprising instructions for:
 receiving a first set of tokens for a document from a text extractor;   obtaining a detector score indicating that each of the received first set of tokens is of an associated type of token;   obtaining a filter score based on the detector score, the filter score based on a scoring rule associated with the associated type;   determining whether the document should be indexed based on an application of a verdict rule, the verdict rule comprising an expression for evaluating the filter score; and   in response to determining that the document should be indexed, providing the document to an indexer adapted to index the first set of tokens for the document.   
     
     
         31 . The non-transitory computer readable medium of  claim 30 , wherein in response to determining that the document should not be indexed, discarding the document based on a suitability criteria. 
     
     
         32 . The non-transitory computer readable medium of  claim 30 , wherein in response to determining that the document should not be indexed, initiating forwarding of the document to a corrector. 
     
     
         33 . The non-transitory computer readable medium of  claim 32 , wherein the corrector is adapted to:
 create an altered document comprising a second set of tokens, and   
       determine whether the altered document should be indexed based on a suitability criteria. 
     
     
         34 . The non-transitory computer readable medium of  claim 33 , wherein the suitability criteria is document size or a number of the second set of tokens. 
     
     
         35 . The non-transitory computer readable medium of  claim 30 , wherein indexing the document comprises indexing the first set of tokens. 
     
     
         36 . The non-transitory computer readable medium of  claim 30 , wherein obtaining the detector score comprises initiating evaluation of each of the received first set of tokens by one or more detectors that are adapted to produce the detector score. 
     
     
         37 . A method, comprising:
 receiving a first set of tokens for a document from a text extractor;   obtaining a detector score indicating that each of the received first set of tokens is of an associated type of token;   obtaining a filter score based on the detector score, the filter score based on a scoring rule associated with the associated type;   determining whether the document should be indexed based on an application of a verdict rule, the verdict rule comprising an expression for evaluating the filter score; and   in response to determining that the document should be indexed, providing the document to an indexer adapted to index the first set of tokens for the document to an indexer adapted to index the first set of tokens for the document.   
     
     
         38 . The method of  claim 37 , wherein in response to determining that the document should not be indexed, discarding the document based on a suitability criteria. 
     
     
         39 . The method of  claim 37 , wherein in response to determining that the document should not be indexed, initiating forwarding of the document to a corrector. 
     
     
         40 . The method of  claim 39 , wherein the corrector is adapted to:
 create an altered document comprising a second set of tokens, and   
       determine whether the altered document should be indexed based on a suitability criteria. 
     
     
         41 . The method of  claim 40 , wherein the suitability criteria is document size or a number of the second set of tokens. 
     
     
         42 . The method of  claim 37 , wherein obtaining the detector score comprises initiating evaluation of each of the received first set of tokens by one or more detectors that are adapted to produce the detector score.

Join the waitlist — get patent alerts

Track US2024004930A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.