US11170179B2ActiveUtilityA1

Systems and methods for natural language processing of structured documents

Assignee: JPMORGAN CHASE BANK NAPriority: Jun 30, 2017Filed: Jun 28, 2018Granted: Nov 9, 2021
Est. expiryJun 30, 2037(~10.9 yrs left)· nominal 20-yr term from priority
G06F 40/205G06V 30/416G06F 40/106G06F 40/40G06F 40/30G06K 9/00469
64
PatentIndex Score
2
Cited by
6
References
20
Claims

Abstract

Systems and methods for natural language processing of structured documents. In another embodiment, in an information processing apparatus comprising at least one computer processor, a method for processing a structured document may include: (1) receiving a document; (2) parsing the document into a plurality of components using a statistical parser; (3) extracting a plurality of entities from each component; (4) identifying a potential relationship between two of the plurality of entities; (5) generating a numeric representation for the potential relationship; (6) confirming the potential relationship with a logical regression model; and (7) generating and storing a unified structured file for the document.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A method for processing a structured document, comprising:
 in an information processing apparatus comprising at least one computer processor: 
 receiving a structured legal document; 
 parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network; 
 extracting a plurality of entities from each document section; 
 identifying a potential relationship between two of the plurality of entities; 
 generating a numeric representation for the potential relationship; 
 confirming the potential relationship with a logical regression model; 
 generating and storing a unified structured file for the structured legal document; 
 receiving feedback on an accuracy of at least one of the plurality of document sections identified using the statistical parser; and 
 updating the statistical parser based on the feedback. 
 
     
     
       2. The method of  claim 1 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection. 
     
     
       3. The method of  claim 1 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document. 
     
     
       4. The method of  claim 1 , wherein the step of parsing the structured legal document into a plurality of document sections comprises identifying a relationship among the plurality of document sections. 
     
     
       5. The method of  claim 1 , further comprising:
 generating a score for each sentence or paragraph of the structured legal document. 
 
     
     
       6. The method of  claim 5 , wherein the score is generated using a latent semantic indexing model. 
     
     
       7. The method of  claim 5 , wherein the score is generated using a continuous bag-of-words model. 
     
     
       8. The method of  claim 1 , wherein the potential relationship is based on an ontology. 
     
     
       9. The method of  claim 1 , wherein the potential relationship is based on a hierarchical correspondence rule. 
     
     
       10. The method of  claim 1 , wherein the numeric representation for the potential relationship is based on functional features, tail features, and head features of the potential relationship. 
     
     
       11. The method of  claim 1 , further comprising:
 identifying a plurality of defined terms in the structured legal document. 
 
     
     
       12. The method of  claim 1 , wherein the logical regression model confirms each potential relationship as being true or false. 
     
     
       13. The method of  claim 1 , further comprising:
 generating a graphical representation of the structured legal document. 
 
     
     
       14. A method for processing a structured document, comprising:
 in an information processing apparatus comprising at least one computer processor: 
 receiving a structured legal document; 
 parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network; 
 extracting a plurality of entities from each document section using a Conditional Random Field model; 
 identifying a potential relationship between two of the plurality of entities; 
 generating a numeric representation for the potential relationship; 
 confirming the potential relationship with a logical regression model; 
 generating and storing a unified structured file for the structured legal document; 
 receiving feedback on an accuracy of at least one of the entities extracted using the Conditional Random Field model; and 
 updating the Conditional Random Field model based on the feedback. 
 
     
     
       15. The method of  claim 14 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection. 
     
     
       16. The method of  claim 14 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document. 
     
     
       17. A method for processing a structured document, comprising:
 in an information processing apparatus comprising at least one computer processor: 
 receiving a structured legal document; 
 parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network; 
 extracting a plurality of entities from each document; 
 identifying a potential relationship between two of the plurality of entities; 
 generating a numeric representation for the potential relationship; 
 confirming the potential relationship with a logical regression model; 
 generating and storing a unified structured file for the structured legal document; 
 receiving feedback on an accuracy of at least one of the potential relationships confirmed using the logical regression model; and 
 updating the logical regression model based on the feedback. 
 
     
     
       18. The method of  claim 17 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection. 
     
     
       19. The method of  claim 17 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document. 
     
     
       20. The method of  claim 17 , wherein further comprising generating a score for each sentence or paragraph of the structured legal document, wherein the score is generated using a latent semantic indexing model or a continuous bag-of-words model.

Join the waitlist — get patent alerts

Track US11170179B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.