Schema generation using natural language processing
Abstract
In a method for generating a schema for a corpus of data, a first corpus of data is received, wherein the first corpus of data includes unstructured text. A processor identifies a set of one or more entity relationships within the first corpus of data, wherein an entity relationship comprises a first entity, a second entity, and a specified relationship between the entities. A processor compares the set of one or more entity relationships to a second corpus of data, wherein the second corpus of data includes text of a subject matter different than the corpus of data. A processor determines a score for each entity relationship based on the comparison to the second corpus of data. A processor generates a schema for the first corpus of data based on the score for each entity relationship of the set of one or more entity relationships.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 - 7 . (canceled)
8 . A computer program product for generating a schema for a corpus of data, the computer program product comprising:
one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the program instructions comprising: program instructions to receive a first corpus of data, wherein the first corpus of data includes, at least, unstructured text; program instructions to identify a set of one or more entity relationships within the first corpus of data, wherein an entity relationship comprises a first entity, a second entity, and a specified relationship between the first entity and the second entity; program instructions to compare the set of one or more entity relationships to a second corpus of data, wherein the second corpus of data includes, at least, text of a subject matter different than the corpus of data; program instructions to determine a score for each entity relationship of the set of one or more entity relationships based on, at least, the comparison to the second corpus of data; and program instructions to generate a schema for the first corpus of data based on, at least, the score for each entity relationship of the set of one or more entity relationships.
9 . The computer program product of claim 8 , further comprising:
program instructions, stored on the one or more computer readable storage media, to generate a graph, wherein the graph illustrates the set of one or more entity relationships and merges duplicate entities into a single entity.
10 . The computer program product of claim 9 , wherein program instructions to generate the schema for the first corpus of data comprise:
program instructions to generate the schema for the first corpus of data based on a hierarchy indicated within the graph and by the score for each entity relationship of the set of one or more entity relationships.
11 . The computer program product of claim 8 , wherein program instructions to determine a score for each entity relationship of the set of one or more entity relationships are further based on a frequency of occurrence of each respective entity relationship within the first corpus of data.
12 . The computer program product of claim 8 , wherein program instructions to compare the set of one or more entity relationships to the second corpus of data include natural language processing.
13 . The computer program product of claim 8 , wherein the second corpus of data comprises an encyclopedia.
14 . The computer program product of claim 8 , wherein the second entity of an entity relationship of the set of entity relationships is an attribute of the first entity.
15 . A computer system for a corpus of data, the computer system comprising:
one or more computer processors, one or more computer readable storage media, and program instructions stored on the computer readable storage media for execution by at least one of the one or more processors, the program instructions comprising: program instructions to receive a first corpus of data, wherein the first corpus of data includes, at least, unstructured text; program instructions to identify a set of one or more entity relationships within the first corpus of data, wherein an entity relationship comprises a first entity, a second entity, and a specified relationship between the first entity and the second entity; program instructions to compare the set of one or more entity relationships to a second corpus of data, wherein the second corpus of data includes, at least, text of a subject matter different than the corpus of data; program instructions to determine a score for each entity relationship of the set of one or more entity relationships based on, at least, the comparison to the second corpus of data; and program instructions to generate a schema for the first corpus of data based on, at least, the score for each entity relationship of the set of one or more entity relationships.
16 . The computer system of claim 15 , further comprising:
program instructions, stored on the computer readable storage media for execution by at least one of the one or more processors, to generate a graph, wherein the graph illustrates the set of one or more entity relationships and merges duplicate entities into a single entity.
17 . The computer system of claim 16 , wherein program instructions to generate the schema for the first corpus of data comprise:
program instructions to generate the schema for the first corpus of data based on a hierarchy indicated within the graph and by the score for each entity relationship of the set of one or more entity relationships.
18 . The computer system of claim 15 , wherein program instructions to determine a score for each entity relationship of the set of one or more entity relationships are further based on a frequency of occurrence of each respective entity relationship within the first corpus of data.
19 . The computer system of claim 15 , wherein program instructions to compare the set of one or more entity relationships to the second corpus of data include natural language processing.
20 . The computer system of claim 15 , wherein the second corpus of data comprises an encyclopedia.Join the waitlist — get patent alerts
Track US2016283523A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.