US2015324338A1PendingUtilityA1

Identification of Layout and Content Flow of an Unstructured Document

Assignee: APPLE INCPriority: Jan 2, 2009Filed: May 12, 2015Published: Nov 12, 2015
Est. expiryJan 2, 2029(~2.4 yrs left)· nominal 20-yr term from priority
G06F 16/93G06F 40/10G06F 40/117G06F 40/174G06F 40/205G06F 40/186G06F 40/103G06F 40/106G06F 40/126G06F 40/40G06F 17/2294G06F 17/28G06F 18/00G06V 30/413G06V 30/414G06F 40/143G06F 40/163
61
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Some embodiments provide a method for analyzing an unstructured document that includes a number of glyphs, each of which has a position in the unstructured document. Based on positions of the glyphs in the unstructured document, the method creates associations between different sets of glyphs in order to identify different sets of glyphs as different words. The method creates associations between different sets of words in order to identify different sets of words as different paragraphs. The method defines associations between paragraphs that are not contiguous in order to define a reading order for the paragraphs.

Claims

exact text as granted — not AI-modified
1 - 25 . (canceled) 
     
     
         26 . A method for analyzing a document comprising a plurality of glyphs, each glyph having a position in the document, the method comprising:
 based on positions of the glyphs in the document, creating associations between glyphs as a text line;   identifying a split in the text line based on identified gaps between glyphs in the text line, the identified split dividing the text line into two separate portions;   based on the identified split, defining the portion of the text line on one side of the split as a first text line and the portion of the text line on the other side of the split as a second text line; and   assigning the first and second text lines to different paragraphs.   
     
     
         27 . The method of  claim 26 , wherein creating associations between glyphs comprises performing cluster analysis on the positions of the glyphs to identify horizontal spacing between glyphs. 
     
     
         28 . The method of  claim 27 , wherein the cluster analysis identifies clusters of horizontal spacing sizes to identify spacing between words and spacing within words. 
     
     
         29 . The method of  claim 26  further comprising assigning the different paragraphs with the first and second text lines to different columns. 
     
     
         30 . The method of  claim 29 , wherein the different columns are on a same page of the document. 
     
     
         31 . The method of  claim 26 , wherein identifying a split comprises determining whether a gap between glyphs in the text line that represents the split spans across a certain number of text lines. 
     
     
         32 . The method of  claim 26 , wherein the document is an unstructured document. 
     
     
         33 . A non-transitory machine readable medium storing a program which when executed by at least on processing unit analyzes a document comprising a plurality of glyphs, each glyph having a position in the document, the program comprising sets of instructions for:
 based on positions of the glyphs in the document, creating associations between glyphs as a text line;   identifying a split in the text line based on identified gaps between glyphs in the text line, the identified split dividing the text line into two separate portions;   based on the identified split, defining the portion of the text line on one side of the split as a first text line and the portion of the text line on the other side of the split as a second text line; and   assigning the first and second text lines to different paragraphs.   
     
     
         34 . The non-transitory machine readable medium of  claim 33 , wherein the set of instructions for creating associations between glyphs comprises a set of instructions for performing cluster analysis on the positions of the glyphs to identify horizontal spacing between glyphs. 
     
     
         35 . The non-transitory machine readable medium of  claim 34 , wherein the cluster analysis identifies clusters of horizontal spacing sizes to identify spacing between words and spacing within words. 
     
     
         36 . The non-transitory machine readable medium of  claim 33 , wherein the program further comprises a set of instructions for assigning the different paragraphs with the first and second text lines to different columns. 
     
     
         37 . The non-transitory machine readable medium of  claim 36 , wherein the different columns are on a same page of the document. 
     
     
         38 . The non-transitory machine readable medium of  claim 33 , wherein the set of instructions for identifying a split comprises a set of instructions for determining whether a gap between glyphs in the text line that represents the split spans across a certain number of text lines. 
     
     
         39 . The non-transitory machine readable medium of  claim 33 , wherein the document is an unstructured document. 
     
     
         40 . A system comprising:
 a set of processing units; and   a non-transitory machine readable medium storing a program which when executed by at least one of the processing units analyzes a document comprising a plurality of glyphs, each glyph having a position in the document, the program comprising sets of instructions for:
 based on positions of the glyphs in the document, creating associations between glyphs as a text line; 
 identifying a split in the text line based on identified gaps between glyphs in the text line, the identified split dividing the text line into two separate portions; 
 based on the identified split, defining the portion of the text line on one side of the split as a first text line and the portion of the text line on the other side of the split as a second text line; and 
 assigning the first and second text lines to different paragraphs. 
   
     
     
         41 . The system of  claim 40 , wherein the set of instructions for creating associations between glyphs comprises a set of instructions for performing cluster analysis on the positions of the glyphs to identify horizontal spacing between glyphs. 
     
     
         42 . The system of  claim 41 , wherein the cluster analysis identifies clusters of horizontal spacing sizes to identify spacing between words and spacing within words. 
     
     
         43 . The system of  claim 40 , wherein the program further comprises a set of instructions for assigning the different paragraphs with the first and second text lines to different columns. 
     
     
         44 . The system of  claim 43 , wherein the different columns are on a same page of the document. 
     
     
         45 . The system of  claim 40 , wherein the set of instructions for identifying a split comprises a set of instructions for determining whether a gap between glyphs in the text line that represents the split spans across a certain number of text lines.

Join the waitlist — get patent alerts

Track US2015324338A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.