US2023153741A1PendingUtilityA1

Continuous data quality assessment and monitoring for big data

Assignee: WALMART APOLLO LLCPriority: Mar 8, 2019Filed: Jan 19, 2023Published: May 18, 2023
Est. expiryMar 8, 2039(~12.5 yrs left)· nominal 20-yr term from priority
G06F 11/3086G06N 5/046G06N 5/025G06Q 10/06395G06F 11/3075G06F 11/3082
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A data quality assessment and monitoring tool addresses inconsistency in large data sets from differing sources, determining data quality attributes such as completeness, conformity, validity, and accuracy. Flexible taxonomies and rollup strategies accommodate diverse business unit needs across a complex enterprise, and provides insight into individual entities’ performance. An exemplary tool comprises a data importer for importing data from a data lake; a rules manager for generating rules and rule sets; a scoring engine for generating data quality scores; a scoring engine for generating data quality scores; a data profiler for producing data quality scores for a plurality of hierarchical data entity units; a hierarchical scoring aggregator for aggregating the data quality scores into first and second tiers of aggregate data quality scores; a persistence component for continually executing an ongoing score rollup strategy; and a reporting component for dynamically outputting the first or second tier aggregate data quality score.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system comprising:
 a data importer, implemented on a processor, that imports data from a data lake;   a rules manager, implemented on the processor, that generates rules and rule sets for the imported data, using dimensions and weights;   a scoring engine, implemented on the processor, that generates data quality scores for the imported data using the rule sets;   a data profiler, implemented on the processor, that operates on data assessment tasks, and uses the scoring engine and the imported data to produce the data quality scores for a plurality of hierarchical data entity units, and to further collate the data quality scores into dimensional scores for the plurality of hierarchical data entity units;   a hierarchical scoring aggregator, implemented on the processor, that aggregates a first set of the data quality scores for a first hierarchical data entity unit of the plurality of hierarchical data entity units into a first tier aggregate data quality score, aggregates a second set of the data quality scores for a second hierarchical data entity unit of the plurality of hierarchical data entity units into another first tier aggregate data quality score, and aggregates the first tier aggregate data quality score for the first hierarchical data entity unit and the first tier aggregate data quality score for the second hierarchical data entity unit into a second tier aggregate data quality score;   a persistence component, implemented on the processor, that continually executes an ongoing score rollup strategy to roll up the first tier aggregate data quality score for the first hierarchical data entity unit and the first tier aggregate data quality score for the second hierarchical data entity unit into the second tier aggregate data quality score; and   a reporting component, implemented on the processor, that dynamically outputs, based on the ongoing score rollup strategy continually executing, one or both of the first tier aggregate data quality score or the second tier aggregate data quality score.   
     
     
         2 . The system of  claim 1 , wherein:
 the hierarchical scoring aggregator:
 aggregates an updated first set of the data quality scores and an updated second set of the data quality scores into updated first tier aggregate data quality scores, and 
 aggregates the second tier aggregate data quality score for the plurality of hierarchical data entity units into a third tier aggregate data quality score, using differently-customized rules and weights for different hierarchical data entity units in the plurality of hierarchical data entity units, 
   the persistence component rolls up the updated first tier aggregate data quality scores into an updated second tier aggregate data quality score, and   the reporting component displays a selected one of the updated first tier aggregate data quality score or the updated second tier aggregate data quality score.   
     
     
         3 . The system of  claim 1 , wherein:
 the scoring engine further generates data metrics for the imported data and the rule sets,   the data imported from the data lake includes raw copies of source data and transformed data, and   the one or dimensional scores measure one or more of accuracy, completeness, conformity, consistency, data decay, duplication, integrity, timeliness, uniqueness, and validity.   
     
     
         4 . The system of  claim 1 , further comprising a job manager that maps the data assessment tasks to the rule sets and the hierarchical data entity units. 
     
     
         5 . The system of  claim 1 , wherein:
 the first tier aggregate data quality score measures an aggregate quality of data in the first tier and the second tier aggregate data quality score measures an aggregate quality of data in the second tier.   
     
     
         6 . The system of  claim 1 , wherein:
 the reporting component further reports aggregate data quality scores for a plurality of different tiers including a first tier and a second tier,   the first tier represents particular data entities, and   the second tier represents product line levels.   
     
     
         7 . The system of  claim 1 , further comprising:
 a data intake node for providing data to the data lake, wherein the data intake node comprises at least one node selected from the list consisting of:   an inventory management system, a retail sales terminal, and a website portal.   
     
     
         8 . A computer-implemented method comprising:
 importing, by a data importer, data from a data lake;   generating, by a rules manager, rules and rule sets for the imported data using dimensions and weights;   generating, by a scoring engine, data quality scores for the imported data using the rule sets;   producing, by a data profiler, the data quality scores for a plurality of hierarchical data entity units using the imported data;   collating, by the data profiler, the data quality scores into dimensional scores for the plurality of hierarchical data entity units;   aggregating, by a hierarchical scoring aggregator, a first set of the data quality scores for a first hierarchical data entity unit of the plurality of hierarchical data entity units into a first tier aggregate data quality score;   aggregating, by the hierarchical scoring aggregator, a second set of the data quality scores for a second hierarchical data entity unit of the plurality of hierarchical data entity units into another first tier aggregate data quality score;   aggregating, by the hierarchical scoring aggregator, the first tier aggregate data quality score for the first hierarchical data entity unit and the first tier aggregate data quality score for the second hierarchical data entity unit into a second tier aggregate data quality score;   continually executing, by a persistence component, an ongoing score rollup strategy to roll up the first tier aggregate data quality score for the first hierarchical data entity unit and the first tier aggregate data quality score for the second hierarchical data entity unit into the second tier aggregate data quality score; and   based on the ongoing score rollup strategy continually executing, dynamically outputting, by a reporting component, one or both of the first tier aggregate data quality score or the second tier aggregate data quality score.   
     
     
         9 . The computer-implemented method of  claim 8 , further comprising:
 aggregating, by the hierarchical scoring aggregator, an updated first set of the data quality scores and an updated second set of the data quality scores into updated first tier aggregate data quality scores; and   aggregating, by the hierarchical scoring aggregator, the second tier aggregate data quality score for the plurality of hierarchical data entity units into a third tier aggregate data quality score, using differently-customized rules and weights for different hierarchical data entity units in the plurality of hierarchical data entity units.   
     
     
         10 . The computer-implemented method of  claim 9 , further comprising:
 rolling up, by the persistence component, the updated first tier aggregate data quality scores into an updated second tier aggregate data quality score; and   displaying, by the reporting component, a selected one of the updated first tier aggregate data quality score or the updated second tier aggregate data quality score.   
     
     
         11 . The computer-implemented method of  claim 8 , further comprising:
 generating, by the scoring engine, data metrics for the imported data and the rule sets,   wherein the data imported from the data lake includes raw copies of source data and transformed data, and   wherein the one or dimensional scores measure one or more of accuracy, completeness, conformity, consistency, data decay, duplication, integrity, timeliness, uniqueness, and validity.   
     
     
         12 . The computer-implemented method of  claim 8 , further comprising mapping, by a job manager, data assessment tasks to the rule sets and the hierarchical data entity units. 
     
     
         13 . The computer-implemented method of  claim 8 , wherein:
 the first tier aggregate data quality score measures an aggregate quality of data in the first tier and the second tier aggregate data quality score measures an aggregate quality of data in the second tier.   
     
     
         14 . The computer-implemented method of  claim 8 , further comprising:
 reporting, by the reporting component, aggregate data quality scores for a plurality of different tiers including a first tier and a second tier,   wherein the first tier represents particular data entities, and   wherein the second tier represents product line levels.   
     
     
         15 . The computer-implemented method of  claim 8 , further comprising:
 providing, by a data intake node, data to the data lake, wherein the data intake node comprises at least one node selected from the list consisting of an inventory management system, a retail sales terminal, and a website portal.   
     
     
         16 . A computer storage device having computer-executable instructions stored thereon, which, on execution by a processor, cause the processor to:
 import, by a data importer implemented on the processor, data from a data lake;   generate, by a rules manager implemented on the processor, rules and rule sets for the imported data using dimensions and weights;   generate, by a scoring engine implemented on the processor, data quality scores for the imported data using the rule sets;   produce, by a data profiler implemented on the processor, the data quality scores for a plurality of hierarchical data entity units using the imported data;   collate, by the data profiler, the data quality scores into dimensional scores for the plurality of hierarchical data entity units;   aggregate, by a hierarchical scoring aggregator implemented on the processor, a first set of the data quality scores for a first hierarchical data entity unit of the plurality of hierarchical data entity units into a first tier aggregate data quality score;   aggregate, by the hierarchical scoring aggregator, a second set of the data quality scores for a second hierarchical data entity unit of the plurality of hierarchical data entity units into another first tier aggregate data quality score;   aggregate, by the hierarchical scoring aggregator, the first tier aggregate data quality score for the first hierarchical data entity unit and the first tier aggregate data quality score for the second hierarchical data entity unit into a second tier aggregate data quality score;   continually execute, by a persistence component implemented on the processor, an ongoing score rollup strategy to roll up the first tier aggregate data quality score for the first hierarchical data entity unit and the first tier aggregate data quality score for the second hierarchical data entity unit into the second tier aggregate data quality score; and   based on the ongoing score rollup strategy continually executing, dynamically output, by a reporting component implemented on the processor, one or both of the first tier aggregate data quality score or the second tier aggregate data quality score.   
     
     
         17 . The one or more computer storage devices of  claim 16 , wherein the instructions further cause the processor to:
 aggregate, by the hierarchical scoring aggregator, an updated first set of the data quality scores and an updated second set of the data quality scores into updated first tier aggregate data quality scores; and   aggregate, by the hierarchical scoring aggregator, the second tier aggregate data quality score for the plurality of hierarchical data entity units into a third tier aggregate data quality score, using differently-customized rules and weights for different hierarchical data entity units in the plurality of hierarchical data entity units.   
     
     
         18 . The one or more computer storage devices of  claim 17 , wherein the instructions further cause the processor to:
 roll up, by the persistence component, the updated first tier aggregate data quality scores into an updated second tier aggregate data quality score; and   display, by the reporting component, a selected one of the updated first tier aggregate data quality score or the updated second tier aggregate data quality score.   
     
     
         19 . The one or more computer storage devices of  claim 16 , wherein the instructions further cause the processor to:
 generate, by the scoring engine, data metrics for the imported data and the rule sets,   wherein the data imported from the data lake includes raw copies of source data and transformed data, and   wherein the one or dimensional scores measure one or more of accuracy, completeness, conformity, consistency, data decay, duplication, integrity, timeliness, uniqueness, and validity.   
     
     
         20 . The one or more computer storage devices of  claim 16 , wherein the instructions further cause the processor to:
 report, by the reporting component, aggregate data quality scores for a plurality of different tiers including a first tier and a second tier,   wherein the first tier represents particular data entities, and   wherein the second tier represents product line levels.

Join the waitlist — get patent alerts

Track US2023153741A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.