Document processing system and method
Abstract
A computer-implemented method comprises storing a plurality of construction project specification documents in a data storage system. The method further comprises, for each of the plurality of construction project specification documents, extracting a plurality of blocks of text from the document, extracting formatting information for the document based on the extracted text blocks, generating a location descriptor for each of the text blocks based on the formatting information, the location descriptor indicating the location of the text block within the document, determining the type of text contained in each of the text blocks, and storing the location descriptor and the type of text contained in the text block for each of the text blocks.
Claims
exact text as granted — not AI-modified1 .- 18 . (canceled)
19 . A system for efficiently managing a database of construction project specification documents, the system comprising:
the database having a plurality of construction project specification documents stored therein, wherein the construction project specification documents are formatted to have a predefined uniform organizational structure, the predefined uniform organizational structure defining different parts of the construction project specification documents according to a specific standard for organizing construction specification documents; memory; a processor; a machine-readable storage medium having a program of instructions thereon, wherein, for each of the plurality of construction project specification documents, the program of instructions, when executed by the processor, causes the system to:
extract a plurality of blocks of text from each construction project specification document,
extract formatting information for each construction project specification document based on the extracted text blocks,
generate a location descriptor for each of the text blocks based on the formatting information, the location descriptor indicating the location of the text block within each construction project specification document, the location descriptor including an indication of a specific part of the different parts of a given construction project specification document of the plurality of construction project specification documents, wherein to generate the location descriptor, the program of instructions utilizes the predefined uniform organizational structure of the plurality of construction project specification documents according to the specific standard for organizing construction specification documents and a sequential offset,
determine a type of text contained in each of the text blocks,
recognize a plurality of entities contained in the plurality of blocks of text, and
store, for each of the text blocks, the location descriptor and the type of text contained in the text block in the database; and
generate a plurality of relatedness scores by comparing each of the plurality of entities against each of the remaining plurality of entities, each of the plurality of relatedness scores corresponding to each pair of compared entities, each of the plurality of relatedness scores based on a likelihood of the compared entities appearing in a common block of text, and
store each relatedness score in the database; and,
the system further including a named entity recognition system configured to perform entity identification and entity extraction, the named entity recognition system identifies entities as one of a plurality of different types of entities, the plurality of different types of entities including named entities, structure indicating entities, and relationship indicating entities.
20 . A system as defined in claim 19 , wherein, to extract the plurality of blocks of text from the document, the program of instructions dissects pages of the document into the text blocks based on headings, subheadings, and indenting within the document.
21 . A system as defined in claim 19 , wherein the specific standard for organizing construction specification documents is a MasterFormat standard.
22 . A system as defined in claim 21 , wherein the location descriptor includes a hierarchical descriptor generated based on the predefined uniform organizational structure of the plurality of construction project specification documents, and wherein the hierarchical descriptor reflects the location of the text block by indicating a division of the predefined uniform organizational structure and a code of the predefined uniform organizational structure.
23 . A system as defined in claim 19 ,
wherein the program of instructions is further configured to identify construction project specification documents that satisfy search criteria received from a user; and wherein the system further comprises user interface logic configured to generate a user interface, the user interface logic configured to generate a plurality of charts for display to the user, wherein the user can interact with the charts to specify modified search criteria, and wherein the user interface logic is configured to receive modified search criteria from the user via one of the charts and update the remaining charts to reflect the modified search criteria.
24 . A system as defined in claim 23 , wherein the plurality of charts include a timeline reflecting the number of construction project specification documents that satisfy the search criteria as a function of time.
25 . A system as defined in claim 23 , wherein the plurality of charts include a map reflecting the number of construction project specification documents that satisfy the search criteria as a function of geographic region.
26 . A system as defined in claim 23 , wherein the plurality of charts include a chart reflecting the number of construction project specification documents associated with projects that are at a specified stage of completion.
27 . A computer-implemented method for efficiently managing a database of construction project specification documents, the method comprising:
storing a plurality of construction project specification documents in the database, wherein the construction project specification documents are formatted to have a predefined uniform organizational structure, the predefined uniform organizational structure defining different parts of the construction project specification documents according to a specific standard for organizing construction project specification documents; for each of the plurality of construction project specification documents:
extracting a plurality of blocks of text from the documents,
extracting formatting information for the documents based on the extracted text blocks,
generating a location descriptor for each of the text blocks based on the formatting information, the location descriptor indicating the location of the text block within the documents, the location descriptor including an indication of a specific part of the different parts of a given construction project specification document of the plurality of construction project specification documents, wherein the predefined uniform organizational structure of the plurality of construction project specification documents according to the specific standard for organizing construction specification documents and a sequential offset are utilized to generate the location descriptor,
determining a type of text contained in each of the text blocks, recognizing a plurality of entities contained in the plurality of blocks of text, via a named entity recognition system configured to perform entity identification and entity extraction, the named entity recognition system identifies entities as one of a plurality of different types of entities, the plurality of different types of entities including named entities, structure indicating entities, and relationship indicating entities, and
storing, for each of the text blocks, the location descriptor and the type of text contained in the text block in the database;
generate a plurality of relatedness scores by comparing each of the plurality of entities against each of the remaining plurality of entities, each of the plurality of relatedness scores corresponding to each pair of compared entities, each of the plurality of relatedness scores based on a likelihood of the compared entities appearing in a common block of text; and
store each relatedness score in the database.
28 . A method as defined in claim 27 , wherein extracting the plurality of blocks of text comprises dissecting pages of the document into the text blocks based on headings, subheadings, and indenting within the document.
29 . A method as defined in claim 27 , wherein the location descriptor includes a hierarchical descriptor generated based on the predefined uniform organizational structure of the construction project specification documents, and wherein the hierarchical descriptor reflects the location of the text block by indicating a division of the predefined uniform organizational structure and a code of the predefined uniform organizational structure.
30 . A method as defined in claim 27 ,
wherein the system is further configured to identify documents that satisfy search criteria received from a user; and wherein the system further comprises user interface logic configured to generate a user interface, the user interface logic configured to generate a plurality of charts for display to the user, wherein the user can interact with the charts to specify modified search criteria, and wherein the user interface logic is configured to receive modified search criteria from the user via one of the charts and update the remaining charts to reflect the modified search criteria.
31 . A computer-implemented method for efficiently managing a database of construction project specification documents, the method comprising:
storing a plurality of construction project specification documents in the database, wherein the construction project specification documents are formatted to have a predefined uniform organizational structure, the predefined uniform organizational structure defining different parts of the construction project specification documents according to a specific standard for organizing construction project specification documents; for each of the plurality of construction project specification documents:
extracting a plurality of blocks of text from the documents,
extracting formatting information for the documents based on the extracted text blocks,
generating a location descriptor for each of the text blocks based on the formatting information, the location descriptor indicating the location of the text block within the documents, the location descriptor including an indication of a specific part of the different parts of a given construction project specification document of the plurality of construction project specification documents, wherein the predefined uniform organizational structure of the plurality of construction project specification documents according to the specific standard for organizing construction specification documents and a sequential offset are utilized to generate the location descriptor,
determining a type of text contained in each of the text blocks, and storing, for each of the text blocks, the location descriptor and the type of text contained in the text block,
recognizing a plurality of entities contained in the plurality of blocks of text, via a named entity recognition system configured to perform entity identification and entity extraction, the named entity recognition system identifies entities as one of a plurality of different types of entities, the plurality of different types of entities including named entities, structure indicating entities, and relationship indicating entities;
generate a plurality of relatedness scores by comparing each of the plurality of entities against each of the remaining plurality of entities, each of the plurality of relatedness scores corresponding to each pair of compared entities, each of the plurality of relatedness scores based on a likelihood of the compared entities appearing in a common block of text;
receiving a search query comprising search criteria from a user electronically via a graphical user interface;
analyzing the construction project specification documents to determine a number of documents that satisfy the search criteria; and
responsive to the search query, generating a display reflecting data regarding the number of documents that satisfy the search criteria.
32 . The system as defined in claim 19 , wherein the named entities includes product names, company names, cities, states, standards or personal names.
33 . The system as defined in claim 19 , wherein the structure indicating entities includes words or phrases indicating document structure.Join the waitlist — get patent alerts
Track US2020159985A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.