Analyzing data from structured and unstructured sources
Abstract
Aspects of the present invention disclose a method for analyzing data from a plurality of data sources. The method includes extracting features of data received from a first source and from a second source by analyzing the data received from the first source of data and from the second source. The method includes processors determining a topic modeling framework, wherein the topic modeling framework detects a semantic structure of the features of the data received from the first data source and the second source. The method includes processors applying the topic modeling framework to the data received from the first source of data the second source of data. The method includes generating a final entity output, wherein the final entity output includes a cluster of entity mentions that the applied topic modeling framework extracts from the first source of data and the second source of data are combined.
Claims
exact text as granted — not AI-modified1 . A method for analyzing data from a plurality of data sources, the method comprising: extracting, by one or more processors, features of data received from a first source and features of the data received from a second source by analyzing the data received from the first source of data and the data received from the second source of data, wherein extracting features of data received from a first source and features of the data received from a second source, further comprises: determining, by one or more processors, that text included in the data received from the first source does not include concepts that relate to a knowledge resource based on an analysis of the text, wherein the data received from the first source includes unstructured data, and wherein the data received from the second source includes structured data that is formatted and contained in a relational database; analyzing, by one or more processors, unstructured data and structured data utilizing semi-supervised learning and unsupervised learning; determining, by one or more processors, a topic modeling framework, wherein the topic modeling framework detects a semantic structure of the features of the data received from the first data source and the data received from the second data source, wherein determining the topic modeling framework further comprises: querying, by one of more processors, a knowledge resource to identify concepts that are associated with the features extracted from the data received from the first source and the features extracted from the data received from the second source; and determining, by one or more processors, the topic modeling framework based on the features extracted from the data received from the first source and the features extracted from the data received from the second source; applying, by one or more processors, the topic modeling framework to the data received from the first source of data and to the data received from the second source of data, wherein applying the topic modeling framework further comprises: activating, by one or more processors, a tokenization process, wherein a tokenization process subdivides a plurality of text during application of the topic modeling framework; and constructing, by one or more processors, a plurality of identical entity chains from the first source and the second source; and generating, by one or more processors, a final entity output, wherein the final entity output includes a cluster of entity mentions that the applied topic modeling framework extracts from the first source of data and the second source of data are combined, wherein generating a final entity output, further comprises: integrating, by one or more processors, a data ranking model, wherein the data ranking is based on a measure of the similarity of the data to the identified topic model, to identify data that refers to the same entity; generating, by one or more processors, an identical entity from the data received from a first source and of the data received from a second source, wherein an identical entity is constructed from a mention that refers to similar entities; generating, by one more processors, a chain of individual entities from the data received from a first source and of the data received from the second source by extracting the generated identical entities and; eliminating, by one or more processors, manual annotations of co-referring relations from the data received from a first source and of the data received from the second source.
Join the waitlist — get patent alerts
Track US2018365592A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.