US2016328374A1PendingUtilityA1

Methods and Data Structures for Improved Searchable Formatted Documents including Citation and Corpus Generation

Individually held — no corporate assignee on recordPriority: Sep 16, 2008Filed: Jul 22, 2016Published: Nov 10, 2016
Est. expirySep 16, 2028(~2.1 yrs left)· nominal 20-yr term from priority
Inventors:Kendyl A. Roman
G06F 40/289G06F 40/205G06F 16/93G06F 40/169G06F 40/103G06F 40/30G06F 17/2785G06F 17/30011G06F 17/241G06F 17/211G06K 9/00463G06F 17/2705G06K 9/18G06F 17/2775G06V 30/224G06V 30/414
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Computer searchable annotated formatted documents are produced by correlating documents stored as a photographic or scanned graphic representations of an actual document (evidence, report, court order, etc.) with textual version of the same documents. A produced document will provide additional details in a computer data structure that supports citation annotation as well as other types of analysis of a document. The computer data structure also supports generation of citation reports and corpus reports. A computer method of creating searchable annotated formatted documents including citation and corpus reports by correlating and correcting text files with photographic or scanned graphic of the original documents. Data structures for correlating and correcting text files with graphic images. Generation of citation reports, concordance reports, and corpus reports. Data structures for citation reports, concordance reports, and corpus reports generation.

Claims

exact text as granted — not AI-modified
1 . A non-transitory computer readable medium encoded with program instructions which are executed by a computer to provide a method of generating internal citations for a formatted document, the instructions comprising the steps of:
 a) obtaining graphic representations of each page of the formatted document, wherein the formatted document is a patent document,   b) optically recognizing characters from the graphic representations, and determining the position of the characters on each page,   c) obtaining a separate and distinct text version of the textual content of the formatted document from a source separate and distinct from the formatted document and the graphic representations obtained therefrom,   d) parsing text from the text version, the parsed text being separate and distinct from the recognized characters,   e) correlating the recognized characters with the parsed text to determine an internal citation for each sentence, wherein the internal citation identifies the document and a citation location inside the document where the corresponding sentence is found,   f) creating a data structure storing data determined in the correlating step.   
     
     
         2 . The computer readable medium of  claim 1  wherein the citation location comprises:
 i) an internal citation column number; and 
 ii) an internal citation line number; 
 
     
     
         3 . The computer readable medium of  claim 1  further comprising a step of:
 using the parsed text to correct the recognized characters, yielding a corrected formatted document. 
 
     
     
         4 . The computer readable medium of  claim 2  further comprising a step of:
 attaching the data structure to the formatted document, 
 wherein, when a portion of text is copied from the formatted document, a corresponding internal citation is included with the copied portion. 
 
     
     
         5 . The computer readable medium of  claim 4  wherein the attaching step yields a searchable annotated formatted document. 
     
     
         6 . The computer readable medium of  claim 2  wherein the data structure comprises citation start data, comprising a start column number and a start line number. 
     
     
         7 . The computer readable medium of  claim 6  wherein the data structure comprises citation end data, comprising an end column number and an end line number. 
     
     
         8 . The computer readable medium of  claim 1  wherein the data structure comprises:
 a) internal citation page number data or internal citation paragraph number data and 
 b) internal citation sentence number data. 
 
     
     
         9 . The computer readable medium of  claim 1  wherein the parsing the text step further includes at least one of the group of:
 a) determining new paragraphs, and 
 b) determining paragraph types. 
 
     
     
         10 . The computer readable medium of  claim 1  wherein the parsing the text step further includes determining document parts, said document parts each comprising a distinct set of pages, wherein the document parts includes at least one of the group of:
 a) abstract, 
 b) drawing, 
 c) specification, and 
 d) claims. 
 
     
     
         11 . The computer readable medium of  claim 1  wherein the parsing the text step further includes determining document sections, said document sections each comprising a distinct group of paragraphs, under one or more headings, wherein the document sections includes at least one of the group of:
 a) field of invention, 
 b) background of invention, 
 c) summary of invention, 
 d) description of drawings, and 
 e) description of invention. 
 
     
     
         12 . The computer readable medium of  claim 1  wherein the determining the position of the characters substep further includes at least one of the group of:
 a) assembling lines, 
 b) allocating lines to columns, and 
 c) calculating line numbers. 
 
     
     
         13 . The computer readable medium of  claim 1  wherein the correlating step further includes at least one of the group of:
 a) determining column numbers, and 
 b) determining line numbers. 
 
     
     
         14 . The computer readable medium of  claim 1  wherein the formatted document contains patent drawing figures, and
 wherein the correlating step further includes at least one of the group of: 
 a) determining figure numbers in the drawing figures, and 
 b) determining figure item numbers in drawings figures. 
 
     
     
         15 . The computer readable medium of  claim 1  wherein the correlating step further includes at least one of the group of:
 a) determining patent claim numbers, and 
 b) parsing patent clauses and determining patent clause numbers. 
 
     
     
         16 . The computer readable medium of  claim 1  further comprising a step of:
 generating a citation document using the correlation data structure. 
 
     
     
         17 . The computer readable medium of  claim 1  further comprising a step of:
 generating a concordance report using the correlation data structure, the concordance report comprising rows comprising: 
 a) a word or phrase, and 
 b) one or more internal citations, indicating where the word or phrase occurs in the formatted document. 
 
     
     
         18 . The computer readable medium of  claim 1  further comprising a step of:
 generating a patent corpus report using the correlation data structure, the patent corpus report comprising rows comprising: 
 a) prior context comprising the entire prior portion of the parsed sentence, 
 b) a word or phrase, 
 c) subsequent context comprising the entire subsequent portion of the parsed sentence, and 
 d) an internal citation. 
 
     
     
         19 . The computer readable medium of  claim 18  wherein the patent corpus report is based on a single word root. 
     
     
         20 . The computer readable medium of  claim 18  wherein the patent corpus report is based on one of the group of:
 a) a phrase, and 
 b) a set of words having similar meaning or usage.

Join the waitlist — get patent alerts

Track US2016328374A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.