US2016342578A1PendingUtilityA1

Systems, Methods, and Media for Generating Structured Documents

Assignee: METRODIGI INCPriority: Jul 26, 2013Filed: Aug 3, 2016Published: Nov 24, 2016
Est. expiryJul 26, 2033(~7 yrs left)· nominal 20-yr term from priority
G06F 40/14G06F 16/951G06F 16/93G06F 40/117G06F 40/10G06F 16/313G06F 40/258G06F 40/205G06F 40/154G06F 40/151G06F 16/958G06F 40/103G06F 17/2705G06F 17/2247G06F 17/2264
38
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems, methods, and media for generating structured documents are provided herein. Methods may include receiving digital source content, the digital source content having source data elements where each source data element includes one or more attributes, determining the one or more attributes for each of the source data elements, tagging each of the source data elements with an identifier based upon their one or more attributes, the identifier defining a function for a particular source data element with a structured document, and generating a structured document from the tagged source data elements.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for converting digital source content into a structured document, the method comprising:
 using an element ranking method comprising:   receiving digital source content, the digital source content comprising source data elements;   determining one or more attributes for each of the source data elements;   comparing an attribute and value for the source data elements to expected values for an identifier;   tagging each of the source data elements with the identifier based upon the attribute, wherein the identifier defines a function for a particular source data element with a structured document; and   generating a structured document from the tagged source data elements.   
     
     
         2 . The method according to  claim 1 , wherein the digital source content is extracted from an at least partially unstructured document. 
     
     
         3 . The method according to  claim 1 , wherein the digital source content comprises one or more images. 
     
     
         4 . The method according to  claim 1 , further comprising:
 comparing the one or more attributes for a source data element to a corpus of exemplary data elements, wherein each of the exemplary data elements includes one or more attributes that define the identifier;   locating an exemplary data element that matches the one or more attributes for the source data element; and   tagging the data element with the identifier of the exemplary data element.   
     
     
         5 . The method according to  claim 4 , further comprising:
 determining if the source data element does not match any of the exemplary data elements; and   assigning a default identifier to the source data element that does not match any of the exemplary data elements.   
     
     
         6 . The method according to  claim 5 , further comprising any of:
 highlighting the source data element having the default identifier in the structured document; and   adding the source data element to a list of source data elements having default identifiers.   
     
     
         7 . The method according to  claim 6 , further comprising:
 receiving feedback from an end-user relative to the default identifier; and   updating the identifier for the source data element based upon the feedback.   
     
     
         8 . The method according to  claim 7 , further comprising updating the corpus of exemplary data elements with the source data element having the updated identifier. 
     
     
         9 . The method according to  claim 4 , further comprising including the tagged source data elements into the corpus of exemplary data elements. 
     
     
         10 . The method according to  claim 4 , wherein the corpus of exemplary data elements is selected based upon a domain which is determined for the digital source content. 
     
     
         11 . The method according to  claim 1 , further comprising assigning a priority level to at least one of the one or more attributes of the source data element. 
     
     
         12 . The method according to  claim 1 , further comprising:
 examining the digital source content for code content, the code content defining executable code;   determining one or more code attributes for the code content;   tagging the code content according to the one or more code attributes; and   modifying the code content within the structured document based upon the tagging of the code content.   
     
     
         13 . The method according to  claim 1 , further comprising:
 determining if the source data element does not match any of the classified data elements of the previously generated structured documents; and   assigning a default identifier to the source data element that does not match any of the classified data elements of previously analyzed structured documents.   
     
     
         14 . The method according to  claim 1 , wherein at least a portion of the one or more attributes for each of the data elements is pre-defined in the source document that comprises the data elements. 
     
     
         15 . The method according to  claim 1 , further comprising generating an electronic publication from the structured document, the electronic publication being configured for use with an electronic document reader. 
     
     
         16 . The method according to  claim 1 , wherein any of the digital source content or the structured document comprises any format selected from ePub, HTML, JavaScript, Cascading Style Sheets, Extensible Markup Language, plain text, Comma Separated Value data, and images. 
     
     
         17 . A system for converting digital source content into a structured document, the system comprising:
 an interface for receiving digital source content, the digital source content comprising source data elements where each source data element includes one or more attributes;   a processor; and   a memory for storing executable instructions that comprise:
 a parsing module that determines the one or more attributes for each of the source data elements; 
   an assignment module that:
 compares an attribute and value for the source data elements to expected values for an identifier and tags each of the source data elements with the identifier based upon their one or more attributes; 
 determines, for data elements of the digital source content, a semantic naming convention, a structure, a granularity, an order, and a styling by comparing the data elements of the digital source content to classified data elements of previously generated structured documents; 
 identifies data elements of the digital source content that match classified data elements of previously analyzed structured documents; and 
 assigns any of the semantic naming convention, the structure, the granularity, the order, and the styling of the classified data elements of previously generated structured documents to the matching data elements; and 
 a document generator module that generates the structured document. 
   
     
     
         18 . The system according to  claim 17 , wherein the digital source content comprises one or more images. 
     
     
         19 . The system according to  claim 17 , wherein the parsing module is further configured to:
 compare the one or more attributes for the source data element to a corpus of exemplary data elements, wherein each of the exemplary data elements includes one or more attributes that define the identifier;   locate an exemplary data element that matches the one or more attributes for the source data element; and   wherein the assignment module is configured to tag the data element with the identifier of the exemplary data element.   
     
     
         20 . The system according to  claim 19 , wherein the assignment module is further configured to:
 determine if the source data element does not match any of the exemplary data elements; and   assign a default identifier to the source data element that does not match any of the exemplary data elements.

Join the waitlist — get patent alerts

Track US2016342578A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.