US2025173044A1PendingUtilityA1

Natural language processing system and method for documents

Assignee: THOMSON REUTERS ENTPR CENTRE GMBHPriority: Feb 3, 2017Filed: Jan 30, 2025Published: May 29, 2025
Est. expiryFeb 3, 2037(~10.5 yrs left)· nominal 20-yr term from priority
G06N 3/0464G06N 3/09G06N 3/0442G06F 40/106G06F 40/30G06F 16/93G06N 20/00G06F 3/0484G06N 3/045G06N 3/044G06N 7/01G06Q 50/18G06Q 50/06G06Q 10/10G06N 20/10G06F 3/0482
74
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In various embodiments, the disclosed systems and methods may receive documents, analyze the documents, categorize portions of the analyzed documents, and present the images of the documents and at least a portion of the categories. The analysis may include identification of categories and the presentation may include indicia of the portion of the image of the document related to the category. The systems and methods disclosed may allow querying and/or reporting of a plurality of documents to facilitate processing.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for categorizing electronic documents, the method comprising:
 receiving, by a processor, a plurality of electronic documents;   associating, by a plurality of trained machine learning models comprising a paragraph model trained to identify one or more categories associated with paragraphs of text and a sentence model trained to identify one or more subcategories of the one or more categories associated with sentences of text, a category and a subcategory for each of the plurality of electronic documents;   ordering, based at least on the category and the subcategory, the plurality of electronic documents chronologically by:
 identifying, by the processor, a similarity between a category and a subcategory associated with a first document of the plurality of electronic documents and a category and a subcategory associated with a second document of the plurality of electronic documents; and 
 identifying, based on the identified similarity, a chronological order of the first document and the second document; and 
 sorting the first document and the second document based on the identified chronological order; and 
   generating a graphical user interface comprising a navigable document image of the first document and the second document in the identified chronological order.   
     
     
         2 . The method of  claim 1 , wherein the category and the subcategory associated with the first document of the plurality of electronic documents indicates the first document is chronologically before the second document of the plurality of electronic documents. 
     
     
         3 . The method of  claim 2 , wherein the category and the subcategory associated with the second document of the plurality of electronic documents indicate the second document is an addendum to the first document of the plurality of electronic documents. 
     
     
         4 . The method of  claim 3 , wherein the indication of the second document as an addendum to the first document of the plurality of electronic documents comprises a text reference to the category and the subcategory of the first document. 
     
     
         5 . The method of  claim 1 , wherein the category of the first document comprises a temporal component of the first document and the subcategory of the first document comprises a date as the temporal component. 
     
     
         6 . The method of  claim 5 , wherein the subcategory of the first document comprises a text portion of the first document including a semantic identifier of a date as the temporal component. 
     
     
         7 . The method of  claim 1 , wherein the second document of the plurality of electronic documents is more recent than the first document of the plurality of electronic documents, the method further comprising:
 removing an association of the category and the subcategory from the first document of the plurality of electronic documents based on the second document being more recent than the first document.   
     
     
         8 . The method of  claim 1 , wherein the plurality of electronic documents are received as an image file and converted to a text format using optical character recognition software. 
     
     
         9 . A system for categorizing electronic documents, the system comprising:
 a processor; and   a memory storing instructions that, when executed, cause the processor to perform operations comprising:
 receiving a plurality of electronic documents; 
 associating, by a plurality of trained machine learning models comprising a paragraph model trained to identify one or more categories associated with paragraphs of text and a sentence model trained to identify one or more subcategories of the one or more categories associated with sentences of text, a category and a subcategory for each of the plurality of electronic documents; 
 ordering, based at least on the category and the subcategory, the plurality of electronic documents chronologically by:
 identifying, by the processor, a similarity between a category and a subcategory associated with a first document of the plurality of electronic documents and a category and a subcategory associated with a second document of the plurality of electronic documents; and 
 identifying, based on the identified similarity, a chronological order of the first document and the second document; and 
 sorting the first document and the second document based on the identified chronological order; and 
 
   generating a graphical user interface comprising a navigable document image of the first document and the second document in the identified chronological order.   
     
     
         10 . The system of  claim 9 , wherein the category and the subcategory associated with the first document of the plurality of electronic documents indicates the first document is chronologically before the second document of the plurality of electronic documents. 
     
     
         11 . The system of  claim 10 , wherein the category and the subcategory associated with the second document of the plurality of electronic documents indicate the second document is an addendum to the first document of the plurality of electronic documents. 
     
     
         12 . The system of  claim 11 , wherein the indication of the second document as an addendum to the first document of the plurality of electronic documents comprises a text reference to the category and the subcategory of the first document. 
     
     
         13 . The system of  claim 9 , wherein the category of the first document comprises a temporal component of the first document and the subcategory of the first document comprises a date as the temporal component. 
     
     
         14 . The system of  claim 13 , wherein the subcategory of the first document comprises a text portion of the first document including a semantic identifier of a date as the temporal component. 
     
     
         15 . The system of  claim 9 , wherein the second document of the plurality of electronic documents is more recent than the first document of the plurality of electronic documents, the processor further performing the operations of:
 removing an association of the category and the subcategory from the first document of the plurality of electronic documents based on the second document being more recent than the first document.   
     
     
         16 . The system of  claim 9 , wherein the plurality of electronic documents are received as an image file and converted to a text format using optical character recognition software. 
     
     
         17 . A non-transitory computer readable medium containing instructions which, when executed by a computer, cause the computer to perform the operations of:
 receiving a plurality of electronic documents;   associating, by a plurality of trained machine learning models comprising a paragraph model trained to identify one or more categories associated with paragraphs of text and a sentence model trained to identify one or more subcategories of the one or more categories associated with sentences of text, a category and a subcategory for each of the plurality of electronic documents;   ordering, based at least on the category and the subcategory, the plurality of electronic documents chronologically by:
 identifying a similarity between a category and a subcategory associated with a first document of the plurality of electronic documents and a category and a subcategory associated with a second document of the plurality of electronic documents; and 
 identifying, based on the identified similarity, a chronological order of the first document and the second document; and 
 sorting the first document and the second document based on the identified chronological order; and 
   generating a graphical user interface comprising a navigable document image of the first document and the second document in the identified chronological order.   
     
     
         18 . The non-transitory computer readable medium of  claim 17 , wherein the category and the subcategory associated with the first document of the plurality of electronic documents indicates the first document is chronologically before the second document of the plurality of electronic documents. 
     
     
         19 . The non-transitory computer readable medium of  claim 18 , wherein the category and the subcategory associated with the second document of the plurality of electronic documents indicate the second document is an addendum to the first document of the plurality of electronic documents. 
     
     
         20 . The non-transitory computer readable medium of  claim 19 , wherein the indication of the second document as an addendum to the first document of the plurality of electronic documents comprises a text reference to the category and the subcategory of the first document.

Join the waitlist — get patent alerts

Track US2025173044A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.