Systems and methods for natural language processing of structured documents
Abstract
Systems and methods for natural language processing of structured documents. In another embodiment, in an information processing apparatus comprising at least one computer processor, a method for processing a structured document may include: (1) receiving a document; (2) parsing the document into a plurality of components using a statistical parser; (3) extracting a plurality of entities from each component; (4) identifying a potential relationship between two of the plurality of entities; (5) generating a numeric representation for the potential relationship; (6) confirming the potential relationship with a logical regression model; and (7) generating and storing a unified structured file for the document.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A method for processing a structured document, comprising:
in an information processing apparatus comprising at least one computer processor:
receiving a structured legal document;
parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network;
extracting a plurality of entities from each document section;
identifying a potential relationship between two of the plurality of entities;
generating a numeric representation for the potential relationship;
confirming the potential relationship with a logical regression model;
generating and storing a unified structured file for the structured legal document;
receiving feedback on an accuracy of at least one of the plurality of document sections identified using the statistical parser; and
updating the statistical parser based on the feedback.
2. The method of claim 1 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection.
3. The method of claim 1 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document.
4. The method of claim 1 , wherein the step of parsing the structured legal document into a plurality of document sections comprises identifying a relationship among the plurality of document sections.
5. The method of claim 1 , further comprising:
generating a score for each sentence or paragraph of the structured legal document.
6. The method of claim 5 , wherein the score is generated using a latent semantic indexing model.
7. The method of claim 5 , wherein the score is generated using a continuous bag-of-words model.
8. The method of claim 1 , wherein the potential relationship is based on an ontology.
9. The method of claim 1 , wherein the potential relationship is based on a hierarchical correspondence rule.
10. The method of claim 1 , wherein the numeric representation for the potential relationship is based on functional features, tail features, and head features of the potential relationship.
11. The method of claim 1 , further comprising:
identifying a plurality of defined terms in the structured legal document.
12. The method of claim 1 , wherein the logical regression model confirms each potential relationship as being true or false.
13. The method of claim 1 , further comprising:
generating a graphical representation of the structured legal document.
14. A method for processing a structured document, comprising:
in an information processing apparatus comprising at least one computer processor:
receiving a structured legal document;
parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network;
extracting a plurality of entities from each document section using a Conditional Random Field model;
identifying a potential relationship between two of the plurality of entities;
generating a numeric representation for the potential relationship;
confirming the potential relationship with a logical regression model;
generating and storing a unified structured file for the structured legal document;
receiving feedback on an accuracy of at least one of the entities extracted using the Conditional Random Field model; and
updating the Conditional Random Field model based on the feedback.
15. The method of claim 14 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection.
16. The method of claim 14 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document.
17. A method for processing a structured document, comprising:
in an information processing apparatus comprising at least one computer processor:
receiving a structured legal document;
parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network;
extracting a plurality of entities from each document;
identifying a potential relationship between two of the plurality of entities;
generating a numeric representation for the potential relationship;
confirming the potential relationship with a logical regression model;
generating and storing a unified structured file for the structured legal document;
receiving feedback on an accuracy of at least one of the potential relationships confirmed using the logical regression model; and
updating the logical regression model based on the feedback.
18. The method of claim 17 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection.
19. The method of claim 17 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document.
20. The method of claim 17 , wherein further comprising generating a score for each sentence or paragraph of the structured legal document, wherein the score is generated using a latent semantic indexing model or a continuous bag-of-words model.Join the waitlist — get patent alerts
Track US11170179B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.