Data management platform using metadata repository
Abstract
An analytical computing environment for large data sets comprises a software platform for data management. The platform provides various automation and self-service features to enable those users to rapidly provision and manage an agile analytics environment. The platform leverages a metadata repository, which tracks and manages all aspects of the data lifecycle. The repository maintains various types of platform metadata including, for example, status information (load dates, quality exceptions, access rights, etc.), definitions (business meaning, technical formats, etc.), lineage (data sources and processes creating a data set, etc.), and user data (user rights, access history, user comments, etc.). Within the platform, the metadata is integrated with all platform services, such as load processing, quality controls and system use. As the system is used, the metadata gets richer and more valuable, supporting additional automation and quality controls.
Claims
exact text as granted — not AI-modifiedWhat is claimed is as follows:
1 . A method comprising:
receiving, by a management server, data from a plurality of data sources, wherein at least one data source of the plurality of data sources comprises base metadata; storing, by the management server, the data in a distributed file system cluster; generating, by the management server, platform metadata for use in managing the data across a set of data management platform services, wherein the set of data management platform services comprises a data shopping component and a data preparation component; receiving, by the data shopping component, a plurality inputs associated with data fields stored in a shopping cart of the data shopping component; generating, by the data shopping component based on the received plurality of inputs, a view of data associated with the data fields, wherein the data fields identify at least a first data source and a second data source of the plurality of data sources having distinct source system formats; processing, by the data preparation component based on the view of the data, one or more commands comprising a join command, a filter command, or a transform command; and generating, by the data preparation component based on the one or more commands, a custom dataset comprising at least a portion of the data.
2 . The method of claim 1 , wherein the set of data management platform services further comprises a load processing service, and the method further comprises:
receiving, by the load processing service, a portion of the data from at least one data source of the plurality of data sources, wherein the load processing service performs a quality control on the portion of the data.
3 . The method of claim 2 , wherein the set of data management platform services further comprises a source processing service, and the method further comprises:
receiving, by the source processing service from the load processing service, the portion of the data, wherein the source processing service formats and profiles the portion of the data received from the load processing service.
4 . The method of claim 3 , wherein the set of data management platform services further comprises a subject processing service, and the method further comprises:
receiving, by the subject processing service from the source processing service, the portion of the data, wherein the subject processing service standardizes the portion of the data received from the source processing service.
5 . The method of claim 1 , wherein the view of the data integrates a subset of the data stored in the first data source and a subset of the data stored in the second data source without modifying the distinct source system formats.
6 . The method of claim 5 , further comprising:
receiving, at a graphical user interface (GUI) in communication with the management server, a selection of a shop-for-data function; and generating, based on the selection of the shop-for-data function, the view of the data.
7 . The method of claim 6 , wherein the view of the data comprises a tracking of data lineage comprising one or more of the plurality of data sources and one or more calculations used to generate the view of the data.
8 . A system comprising:
a management server configured to:
receive data from a plurality of data sources, wherein at least one data source of the plurality of data sources comprises base metadata;
store the data in a distributed file system cluster; and
generate platform metadata for use in managing the data across a set of data management platform services, wherein the set of data management platform services comprises a data shopping component and a data preparation component;
the data shopping component configured to:
receive a plurality inputs associated with data fields stored in a shopping cart of the data shopping component; and
generate, based on the received plurality of inputs, a view of data associated with the data fields, wherein the data fields identify at least a first data source and a second data source of the plurality of data sources having distinct source system formats; and
the data preparation component configured to:
process, based on the view of the data, one or more commands comprising a join command, a filter command, or a transform command; and
generate, based on the one or more commands, a custom dataset comprising at least a portion of the data.
9 . The system of claim 8 , wherein the set of data management platform services further comprises a load processing service, and the system further comprises:
the load processing service configured to:
receive a portion of the data from at least one data source of the plurality of data sources, wherein the load processing service performs a quality control on the portion of the data.
10 . The system of claim 9 , wherein the set of data management platform services further comprises a source processing service, and the system further comprises:
the source processing service configured to:
receive, from the load processing service, the portion of the data, wherein the source processing service formats and profiles the portion of the data received from the load processing service.
11 . The system of claim 10 , wherein the set of data management platform services further comprises a subject processing service, and the system further comprises:
the subject processing service configured to:
receive, from the source processing service, the portion of the data, wherein the subject processing service standardizes the portion of the data received from the source processing service.
12 . The system of claim 8 , wherein the view of the data integrates a subset of the data stored in the first data source and a subset of the data stored in the second data source without modifying the distinct source system formats.
13 . The system of claim 12 , further comprising:
a graphical user interface (GUI) in communication with the management server configured to:
receive a selection of a shop-for-data function; and
generate, based on the selection of the shop-for-data function, the view of the data.
14 . The system of claim 13 , wherein the view of the data comprises a tracking of data lineage comprising one or more of the plurality of data sources and one or more calculations used to generate the view of the data.
15 . A non-transitory computer readable medium storing processor executable instructions that, when executed by at least one processor, cause the at least one processor to:
receive data from a plurality of data sources, wherein at least one data source of the plurality of data sources comprises base metadata; store the data in a distributed file system cluster; generate platform metadata for use in managing the data across a set of data management platform services, wherein the set of data management platform services comprises a data shopping component and a data preparation component; receive a plurality inputs associated with data fields stored in a shopping cart of the data shopping component; generate, based on the received plurality of inputs, a view of data associated with the data fields, wherein the data fields identify at least a first data source and a second data source of the plurality of data sources having distinct source system formats; process, based on the view of the data, one or more commands comprising a join command, a filter command, or a transform command; and generate, based on the one or more commands, a custom dataset comprising at least a portion of the data.
16 . The non-transitory computer readable medium of claim 15 , wherein the set of data management platform services further comprises a load processing service, and the processor executable instructions further cause the at least one processor to:
receive a portion of the data from at least one data source of the plurality of data sources, wherein the load processing service performs a quality control on the portion of the data.
17 . The non-transitory computer readable medium of claim 16 , wherein the set of data management platform services further comprises a source processing service, and the processor executable instructions further cause the at least one processor to:
Receive, from the load processing service, the portion of the data, wherein the source processing service formats and profiles the portion of the data received from the load processing service.
18 . The non-transitory computer readable medium of claim 15 , wherein the view of the data integrates a subset of the data stored in the first data source and a subset of the data stored in the second data source without modifying the distinct source system formats.
19 . The non-transitory computer readable medium of claim 18 , wherein the processor executable instructions further cause the at least one processor to:
receive a selection of a shop-for-data function; and generate, based on the selection of the shop-for-data function, the view of the data.
20 . The non-transitory computer readable medium of claim 19 , wherein the view of the data comprises a tracking of data lineage comprising one or more of the plurality of data sources and one or more calculations used to generate the view of the data.Join the waitlist — get patent alerts
Track US2020125530A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.