US9619448B2ActiveUtilityA1

Automated document revision markup and change control

Assignee: IBMPriority: Nov 7, 2011Filed: Sep 3, 2015Granted: Apr 11, 2017
Est. expiryNov 7, 2031(~5.3 yrs left)· nominal 20-yr term from priority
G06F 16/93G06F 40/117G06F 40/177G06F 40/197G06F 17/2247G06F 17/245G06F 17/30011G06F 17/2288G06F 17/218G06F 40/143
78
PatentIndex Score
2
Cited by
32
References
20
Claims

Abstract

Automated comparison of Darwin Information Typing Architecture (DITA) documents for revision mark-up includes reading document data from first and second DITA documents into respective document object model trees of nodes, and identifying and collapsing emphasis subtree nodes in the trees into their parent nodes, the collapsing caching emphasis data from the identified subtree nodes. A traversal transforms the model trees into respective node lists and captures adjacent sibling emphasis subtree nodes as single text nodes. The node lists are merged into a merged node list that recognizes matches node pairs having primary sort key information and document structure metadata meeting a match threshold, with differences between matching tokens of the node pairs saved. A merged document object model built from the refined merged node list is transformed into a hypertext mark-up language document.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A computer-implemented method for automated comparison of Darwin Information Typing Architecture (DITA) documents, the method comprising executing on a processor the steps of:
 reading document data from a first DITA table document into a first document object model tree comprising a plurality of nodes, and from a second DITA table document into a second document object model tree comprising a plurality of nodes; 
 normalizing table attributes of the first document object model tree and the second document object model tree; 
 transforming via preorder traversal the first document object model tree into a first pre-order node list output, and the second document object model tree into a second pre-order node list output; 
 constructing unique table header labels for nodes in the first pre-order node list output, and for nodes in the second pre-order node list output; 
 comparing, via a table-specific fuzzy Longest Common Subsequence (LCS) process, the unique table header labels of the first constructed pre-order node list output to the unique table header labels of the second constructed pre-order node list output, to thereby generate a merged header node list; 
 analyzing column header name text to distinguish between new and modified columns in the first constructed pre-order node list output and in the second constructed pre-order node list output; and 
 generating a first column name map from old column names for the first constructed pre-order node list output to column names in the merged header node list, and a second column name map from old column names for the second constructed pre-order node list output to column names in the merged header node list. 
 
     
     
       2. The method of  claim 1 , further comprising:
 integrating computer-readable program code into a computer system comprising the processor, a computer readable memory in circuit communication with the processor, and a computer readable storage medium in circuit communication with the processor; and 
 wherein the processor executes program code instructions stored on the computer-readable storage medium via the computer readable memory and thereby performs the steps of reading the document data from the first DITA table document into the first document object model tree and from the second DITA table document into the second document object model tree, normalizing the table attributes of the first document object model tree and the second document object model tree, transforming the first document object model tree into the first pre-order node list output and the second document object model tree into the second pre-order node list output, constructing the unique table header labels for the nodes in the first pre-order node list output and for the nodes in the first pre-order node list output, comparing the unique table header labels of the first constructed pre-order node list output to the unique table header labels of the second constructed pre-order node list output to generate the merged header node list, analyzing the column header name text to distinguish between new and modified columns in the first constructed pre-order node list output and in the second constructed pre-order node list output, and generating the first column name map and the second column name map. 
 
     
     
       3. The method of  claim 1 , further comprising:
 mapping metadata of the document data read from the first DITA table document to the merged header node list as a function of the first column name map, and metadata of the document data read from the second DITA table document to the merged header node list as a function of the second column name map; 
 reading the mapped metadata of the first DITA table document into first Document Object Model (DOM) tree data, and the mapped metadata of the second DITA table document into second DOM tree data; 
 transforming, via preorder traversal, the first DOM tree data into a first DOM pre-order node list output, and the second DOM tree data into a second DOM pre-order node list output; 
 constructing unique column table header labels for nodes in the first DOM pre-order node list output, and for nodes in the second DOM pre-order node list output; and 
 comparing, via a table-specific fuzzy LCS process, the unique column header labels of the first DOM pre-order node list output to the unique column header labels of the second DOM pre-order node list output to thereby generate the merged header node list. 
 
     
     
       4. The method of  claim 3 , further comprising:
 normalizing the merged header node list as a function of the mapped metadata to correct structural table issues. 
 
     
     
       5. The method of  claim 3 , wherein the step of transforming the first document object model tree into the first pre-order node list output and the second document object model tree into the second pre-order node list output captures adjacent sibling emphasis subtree nodes as single text nodes; and
 the method further comprising: 
 identifying and collapsing emphasis subtree nodes in the first document object model tree into parent nodes in the first document object model tree, and emphasis subtree nodes in the second document object model tree into parent nodes in the second document object model tree, the collapsing comprising caching emphasis data from the identified subtree nodes; 
 recognizing matches of node pairs of the first constructed pre-order node list output and the second constructed pre-order node list output that have primary sort key information and document structure metadata meeting a threshold percentage of match; 
 saving differences between matching tokens of the node pairs; 
 separating out table segments from the merged header node list into a table node list and a non-table node list; 
 recovering the cached emphasis data for the table segments in the table node list; 
 building a merged document object model from the merged node list and the non-table node list and the recovered cached emphasis data for the table segments in the table node list; and 
 transforming the built merged document object model into a hypertext mark-up language document that displays the saved differences between the matching tokens as word-level highlighting mark-ups within the refined tables. 
 
     
     
       6. The method of  claim 5 , further comprising:
 tagging and aligning matching columns as a function of comparing table headers. 
 
     
     
       7. The method of  claim 5 , further comprising:
 re-merging table data with tagged columns so that the tagged columns stay aligned. 
 
     
     
       8. The method of  claim 7 , further comprising:
 in response to a threshold number of words of compared phrases matching, determining that word phrases match in a node text content; and 
 indicating word level differences as word-level highlighting mark-ups within matching compared phrases. 
 
     
     
       9. A system, comprising:
 a processing unit; 
 a computer readable memory coupled to the processing unit; and 
 a computer-readable storage medium coupled to the processing unit; 
 wherein the processing unit executes computer instructions stored on the computer-readable storage medium via the computer readable memory and is thereby caused to: 
 read document data from a first DITA table document into a first document object model tree comprising a plurality of nodes, and from a second DITA table document into a second document object model tree comprising a plurality of nodes; 
 normalize table attributes of the first document object model tree and the second document object model tree; 
 transform via preorder traversal the first document object model tree into a first pre-order node list output, and the second document object model tree into a second pre-order node list output; 
 construct unique table header labels for nodes in the first pre-order node list output, and for nodes in the second pre-order node list output; 
 compare, via a table-specific fuzzy Longest Common Subsequence (LCS) process, the unique table header labels of the first constructed pre-order node list output to the unique table header labels of the second constructed pre-order node list output, to thereby generate a merged header node list; 
 analyze column header name text to distinguish between new and modified columns in the first constructed pre-order node list output and in the second constructed pre-order node list output; and 
 generate a first column name map from old column names for the first constructed pre-order node list output to column names in the merged header node list, and a second column name map from old column names for the second constructed pre-order node list output to column names in the merged header node list. 
 
     
     
       10. The system of  claim 9 , wherein the processing unit executes the program instructions stored on the computer-readable storage medium via the computer readable memory, and thereby further:
 maps metadata of the document data read from the first DITA table document to the merged header node list as a function of the first column name map, and metadata of the document data read from the second DITA table document to the merged header node list as a function of the second column name map; 
 reads the mapped metadata of the first DITA table document into first Document Object Model (DOM) tree data, and the mapped metadata of the second DITA table document into second DOM tree data; 
 transforms, via preorder traversal, the first DOM tree data into a first DOM pre-order node list output, and the second DOM tree data into a second DOM pre-order node list output; 
 constructs unique column table header labels for nodes in the first DOM pre-order node list output, and for nodes in the second DOM pre-order node list output; and 
 compares, via a table-specific fuzzy LCS process, the unique column header labels of the first DOM pre-order node list output to the unique column header labels of the second DOM pre-order node list output to thereby generate the merged header node list. 
 
     
     
       11. The system of  claim 10 , wherein the processing unit executes the program instructions stored on the computer-readable storage medium via the computer readable memory, and thereby normalizes the merged header node list as a function of the mapped metadata to correct structural table issues. 
     
     
       12. The system of  claim 10 , wherein the processing unit executes the program instructions stored on the computer-readable storage medium via the computer readable memory, and thereby:
 transforms the first document object model tree into the first pre-order node list output and the second document object model tree into the second pre-order node list output by capturing adjacent sibling emphasis subtree nodes as single text nodes; 
 identifies and collapses emphasis subtree nodes in the first document object model tree into parent nodes in the first document object model tree, and emphasis subtree nodes in the second document object model tree into parent nodes in the second document object model tree, by caching emphasis data from the identified subtree nodes; 
 recognizes matches of node pairs of the first constructed pre-order node list output and the second constructed pre-order node list output that have primary sort key information and document structure metadata meeting a threshold percentage of match; 
 saves differences between matching tokens of the node pairs; 
 separates out table segments from the merged header node list into a table node list and a non-table node list; 
 recovers the cached emphasis data for the table segments in the table node list; 
 builds a merged document object model from the merged node list and the non-table node list and the recovered cached emphasis data for the table segments in the table node list; and 
 transforms the built merged document object model into a hypertext mark-up language document that displays the saved differences between the matching tokens as word-level highlighting mark-ups within the refined tables. 
 
     
     
       13. The system of  claim 12 , wherein the processing unit executes the program instructions stored on the computer-readable storage medium via the computer readable memory, and thereby tags and aligns matching columns as a function of comparing table headers. 
     
     
       14. The system of  claim 12 , wherein the processing unit executes the program instructions stored on the computer-readable storage medium via the computer readable memory, and thereby:
 re-merges table data with tagged columns so that the tagged columns stay aligned; 
 in response to a threshold number of words of compared phrases matching, determines that word phrases match in a node text content; and 
 indicates word level differences as word-level highlighting mark-ups within matching compared phrases. 
 
     
     
       15. An article of manufacture for automated comparison of Darwin Information Typing Architecture (DITA) documents, comprising:
 a computer readable hardware storage device having computer readable program code embodied therewith, wherein the computer readable hardware storage device is not a transitory signal per se, the computer readable program code comprising instructions for execution by a computer processor that cause the computer processor to: 
 read document data from a first DITA table document into a first document object model tree comprising a plurality of nodes, and from a second DITA table document into a second document object model tree comprising a plurality of nodes; 
 normalize table attributes of the first document object model tree and the second document object model tree; 
 transform via preorder traversal the first document object model tree into a first pre-order node list output, and the second document object model tree into a second pre-order node list output; 
 construct unique table header labels for nodes in the first pre-order node list output, and for nodes in the second pre-order node list output; 
 compare, via a table-specific fuzzy Longest Common Subsequence (LCS) process, the unique table header labels of the first constructed pre-order node list output to the unique table header labels of the second constructed pre-order node list output, to thereby generate a merged header node list; 
 analyze column header name text to distinguish between new and modified columns in the first constructed pre-order node list output and in the second constructed pre-order node list output; and 
 generate a first column name map from old column names for the first constructed pre-order node list output to column names in the merged header node list, and a second column name map from old column names for the second constructed pre-order node list output to column names in the merged header node list. 
 
     
     
       16. The article of manufacture of  claim 15 , wherein the computer readable program code instructions for execution by the computer processor, further cause the computer processor to: map metadata of the document data read from the first DITA table document to the merged header node list as a function of the first column name map, and metadata of the document data read from the second DITA table document to the merged header node list as a function of the second column name map;
 read the mapped metadata of the first DITA table document into first Document Object Model (DOM) tree data, and the mapped metadata of the second DITA table document into second DOM tree data; 
 transform, via preorder traversal, the first DOM tree data into a first DOM pre-order node list output, and the second DOM tree data into a second DOM pre-order node list output; 
 construct unique column table header labels for nodes in the first DOM pre-order node list output, and for nodes in the second DOM pre-order node list output; and 
 compare, via a table-specific fuzzy LCS process, the unique column header labels of the first DOM pre-order node list output to the unique column header labels of the second DOM pre-order node list output to thereby generate the merged header node list. 
 
     
     
       17. The article of manufacture of  claim 16 , wherein the computer readable program code instructions for execution by the computer processor, further cause the computer processor to normalize the merged header node list as a function of the mapped metadata to correct structural table issues. 
     
     
       18. The article of manufacture of  claim 16 , wherein the computer readable program code instructions for execution by the computer processor, further cause the computer processor to:
 transform the first document object model tree into the first pre-order node list output and the second document object model tree into the second pre-order node list output by capturing adjacent sibling emphasis subtree nodes as single text nodes; 
 identify and collapse emphasis subtree nodes in the first document object model tree into parent nodes in the first document object model tree, and emphasis subtree nodes in the second document object model tree into parent nodes in the second document object model tree, by caching emphasis data from the identified subtree nodes; 
 recognize matches of node pairs of the first constructed pre-order node list output and the second constructed pre-order node list output that have primary sort key information and document structure metadata meeting a threshold percentage of match; 
 save differences between matching tokens of the node pairs; 
 separate out table segments from the merged header node list into a table node list and a non-table node list; 
 recover the cached emphasis data for the table segments in the table node list; 
 build a merged document object model from the merged node list and the non-table node list and the recovered cached emphasis data for the table segments in the table node list; and 
 transform the built merged document object model into a hypertext mark-up language document that displays the saved differences between the matching tokens as word-level highlighting mark-ups within the refined tables. 
 
     
     
       19. The article of manufacture of  claim 18 , wherein the computer readable program code instructions for execution by the computer processor, further cause the computer processor to tag and align matching columns as a function of comparing table headers. 
     
     
       20. The article of manufacture of  claim 18 , wherein the computer readable program code instructions for execution by the computer processor, further cause the computer processor to:
 re-merge table data with tagged columns so that the tagged columns stay aligned; 
 in response to a threshold number of words of compared phrases matching, determine that word phrases match in a node text content; and 
 indicate word level differences as word-level highlighting mark-ups within matching compared phrases.

Join the waitlist — get patent alerts

Track US9619448B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.