Regulatory document analysis with natural language processing
Abstract
Technologies are provided for automatically comparing versions of a regulatory document and highlighting meaningful changes to each version of the regulatory document. An analysis engine accepts two inputs of a regulatory document in HTML format. One input is an original version of the regulatory document and one input is a revised version of the regulatory document. The documents are processed by the analysis engine to highlight added content as compared to the original version of the HTML content and the second document being processed to highlight removed content as compared to the revised version of the HTML content. These highlighted documents are then presented to the user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
obtaining a first document and a second document, the second document being a revised version of the first document, wherein the first document is in a first format and the second document is in a second format; processing the first document to generate a first processed document by converting the first document from the first format to a third format, removing content from the first document in the third format, and converting the first document with the content removed back to the first format; processing the second document to generate a second processed document by converting the second document from a fourth format to the third format, removing content from the second document in the third format, and converting the second document with the content removed back to the second format; identifying one or more first content items of the first processed document that are not included in the second processed document; identifying one or more second content items of the second processed document that are not included in the first processed document; and displaying, in a user interface of a display device, the first document and highlighting that indicates the one or more first content items; and displaying, in the user interface of the display device, the second document and highlighting that indicates the one or more second content items.
2 . The method of claim 1 , wherein the third format is a tree data structure comprising a plurality of first nodes, and wherein processing the first document to generate the first processed document comprises removing at least one first node from the plurality of first nodes.
3 . The method of claim 1 , wherein the fourth format is a tree data structure comprising a plurality of second nodes, and wherein processing the second document to generate the second processed document comprises removing at least one second node from the plurality of second nodes.
4 . The method of claim 1 , wherein processing the first document to generate the first processed document further comprises lemmatizing tokens of the first document.
5 . The method of claim 4 , wherein identifying one or more first content items of the first processed document that are not included in the second processed document comprising comparing a set of lemmatized tokens for the first processed document to a set of lemmatized tokens for the second processed document.
6 . The method of claim 1 , wherein processing the second document to generate the second processed document further comprises lemmatizing tokens of the second document.
7 . The method of claim 6 , wherein identifying one or more second content items of the second processed document that are not included in the first processed document comprises comparing a set of lemmatized tokens for the second processed document to a set of lemmatized tokens for the first processed document.
8 . A system comprising:
a processing system; and one or more computer readable storage media storing instructions which, when executed by the processing system, cause the processing system to perform a method comprising:
obtaining a first document and a second document from a database, the first document comprising a first set of content items and the second document comprising a second set of content items;
processing the first document to generate a first processed document, wherein the first processed document comprises a first set of processed content items, wherein a number of content items in the first set of processed content items is less than a number of content items in the first set of content items;
processing the second document to generate a second processed document, wherein the second processed document comprises a second set of processed content items, wherein a number of content items in the second set of processed content items is less than a number of content items in the second set of content items;
identifying a first portion of the first processed document that is different from the second processed document;
identifying a second portion of the second processed document that is different from the first processed document; and
displaying, in a user interface of a display device, the first processed document, an indicator that indicates the first portion; and
displaying, in the user interface of the display device, the second processed document and an indicator that indicates the second portion.
9 . The system of claim 8 , wherein processing the first document to generate a first processed document comprises transforming data representing the first document into a first data structure comprising a plurality of first nodes and removing at least one first node from the plurality of first nodes.
10 . The system of claim 8 , wherein processing the second document to generate a second processed document comprises transforming data representing the second document into a second data structure comprising a plurality of second nodes and removing at least one second node from the plurality of second nodes.
11 . The system of claim 8 , wherein processing the first document to generate a first processed document further comprises lemmatizing words of the first document.
12 . The system of claim 8 , wherein processing the second document to generate a second processed document further comprises lemmatizing words of the second document.
13 . The system of claim 8 , wherein processing the first document to generate a first processed document further comprises removing at least one stop word from words of the first document.
14 . The system of claim 8 , wherein processing the second document to generate a second processed document further comprises removing at least one stop word from words of the second document.
15 . One or more non-transitory computer-readable media storing computer-readable instructions that, when executed by a processing system, cause an electronic device to perform a method comprising:
obtaining a first document and a second document from a database, the first document comprising a first set of content items and the second document comprising a second set of content items; processing the first document to generate a first processed document, wherein the first processed document comprises a first set of processed content items, wherein a number of content items in the first set of processed content items is less than a number of content items in the first set of content items; processing the second document to generate a second processed document, wherein the second processed document comprises a second set of processed content items, wherein a number of content items in the second set of processed content items is less than a number of content items in the second set of content items; identifying a first portion of the first processed document that is different from the second processed document; identifying a second portion of the second processed document that is different from the first processed document; and displaying, in a user interface of a display device, the first processed document, an indicator that indicates the first portion; and displaying, in the user interface of the display device, the second processed document and an indicator that indicates the second portion.
16 . The one or more non-transitory computer-readable media of claim 15 , wherein processing the first document to generate a first processed document comprises transforming data representing the first document into a first data structure comprising a plurality of first nodes and removing at least one first node from the plurality of first nodes.
17 . The one or more non-transitory computer-readable media of claim 15 , wherein processing the second document to generate a second processed document comprises transforming data representing the second document into a second data structure comprising a plurality of second nodes and removing at least one second node from the plurality of second nodes.
18 . The one or more non-transitory computer-readable media of claim 15 , wherein processing the first document to generate a first processed document further comprises lemmatizing words of the first document.
19 . The one or more non-transitory computer-readable media of claim 15 , wherein processing the second document to generate a second processed document further comprises lemmatizing words of the second document.
20 . The one or more non-transitory computer-readable media of claim 15 , wherein the first document is in (Hypertext Markup Language) HTML format and the second document is HTML format, the second document being a revised version of the first document.Join the waitlist — get patent alerts
Track US2023334230A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.