Identification of Layout and Content Flow of an Unstructured Document
Abstract
Some embodiments provide a method for analyzing an unstructured document that includes a number of glyphs, each of which has a position in the unstructured document. Based on positions of the glyphs in the unstructured document, the method creates associations between different sets of glyphs in order to identify different sets of glyphs as different words. The method creates associations between different sets of words in order to identify different sets of words as different paragraphs. The method defines associations between paragraphs that are not contiguous in order to define a reading order for the paragraphs.
Claims
exact text as granted — not AI-modified1 - 25 . (canceled)
26 . A method for analyzing a document comprising a plurality of glyphs, each glyph having a position in the document, the method comprising:
based on positions of the glyphs in the document, creating associations between glyphs as a text line; identifying a split in the text line based on identified gaps between glyphs in the text line, the identified split dividing the text line into two separate portions; based on the identified split, defining the portion of the text line on one side of the split as a first text line and the portion of the text line on the other side of the split as a second text line; and assigning the first and second text lines to different paragraphs.
27 . The method of claim 26 , wherein creating associations between glyphs comprises performing cluster analysis on the positions of the glyphs to identify horizontal spacing between glyphs.
28 . The method of claim 27 , wherein the cluster analysis identifies clusters of horizontal spacing sizes to identify spacing between words and spacing within words.
29 . The method of claim 26 further comprising assigning the different paragraphs with the first and second text lines to different columns.
30 . The method of claim 29 , wherein the different columns are on a same page of the document.
31 . The method of claim 26 , wherein identifying a split comprises determining whether a gap between glyphs in the text line that represents the split spans across a certain number of text lines.
32 . The method of claim 26 , wherein the document is an unstructured document.
33 . A non-transitory machine readable medium storing a program which when executed by at least on processing unit analyzes a document comprising a plurality of glyphs, each glyph having a position in the document, the program comprising sets of instructions for:
based on positions of the glyphs in the document, creating associations between glyphs as a text line; identifying a split in the text line based on identified gaps between glyphs in the text line, the identified split dividing the text line into two separate portions; based on the identified split, defining the portion of the text line on one side of the split as a first text line and the portion of the text line on the other side of the split as a second text line; and assigning the first and second text lines to different paragraphs.
34 . The non-transitory machine readable medium of claim 33 , wherein the set of instructions for creating associations between glyphs comprises a set of instructions for performing cluster analysis on the positions of the glyphs to identify horizontal spacing between glyphs.
35 . The non-transitory machine readable medium of claim 34 , wherein the cluster analysis identifies clusters of horizontal spacing sizes to identify spacing between words and spacing within words.
36 . The non-transitory machine readable medium of claim 33 , wherein the program further comprises a set of instructions for assigning the different paragraphs with the first and second text lines to different columns.
37 . The non-transitory machine readable medium of claim 36 , wherein the different columns are on a same page of the document.
38 . The non-transitory machine readable medium of claim 33 , wherein the set of instructions for identifying a split comprises a set of instructions for determining whether a gap between glyphs in the text line that represents the split spans across a certain number of text lines.
39 . The non-transitory machine readable medium of claim 33 , wherein the document is an unstructured document.
40 . A system comprising:
a set of processing units; and a non-transitory machine readable medium storing a program which when executed by at least one of the processing units analyzes a document comprising a plurality of glyphs, each glyph having a position in the document, the program comprising sets of instructions for:
based on positions of the glyphs in the document, creating associations between glyphs as a text line;
identifying a split in the text line based on identified gaps between glyphs in the text line, the identified split dividing the text line into two separate portions;
based on the identified split, defining the portion of the text line on one side of the split as a first text line and the portion of the text line on the other side of the split as a second text line; and
assigning the first and second text lines to different paragraphs.
41 . The system of claim 40 , wherein the set of instructions for creating associations between glyphs comprises a set of instructions for performing cluster analysis on the positions of the glyphs to identify horizontal spacing between glyphs.
42 . The system of claim 41 , wherein the cluster analysis identifies clusters of horizontal spacing sizes to identify spacing between words and spacing within words.
43 . The system of claim 40 , wherein the program further comprises a set of instructions for assigning the different paragraphs with the first and second text lines to different columns.
44 . The system of claim 43 , wherein the different columns are on a same page of the document.
45 . The system of claim 40 , wherein the set of instructions for identifying a split comprises a set of instructions for determining whether a gap between glyphs in the text line that represents the split spans across a certain number of text lines.Join the waitlist — get patent alerts
Track US2015324338A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.