Analyzing large-scale data processing jobs
Abstract
Methods, systems, and apparatus for data analysis in a distributed computing system by accessing data stored at a first processing zone associated with a distributed data processing job, detecting information identifying a particular child job associated with the distributed data processing job, comparing the identifying information to data stored at a second processing zone, and identifying an additional child job as associated with the distributed data processing job based on a result of the comparison. The methods, systems and apparatus are further for correlating particular output data associated with the particular child job and additional output data associated with the additional child job for the distributed data processing job, determining performance data for the distributed data processing job based on the output data associated with each of the particular child job and the additional child job, and providing for display the performance data for the distributed data processing job.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
obtaining, at data processing hardware of a distributed computing system, data associated with a parent job that has executed on the distributed computing system; determining, by the data processing hardware, one or more child jobs associated with the parent job, the one or more child jobs distributed across two or more processing zones of the distributed computing system; identifying, by the data processing hardware, performance data for the one or more child jobs associated with the parent job, the performance data comprising a status of each of the one or more child jobs; receiving, at the data processing hardware, from a user of the distributed computing system, a user selection selecting performance data of at least one of the one or more child jobs; and displaying, by the data processing hardware, at a display in communication with the data processing hardware, the performance data of the at least one of the one or more child jobs selected by the user selection.
2 . The method of claim 1 , wherein identifying performance data for the one or more child jobs associated with the parent job comprises detecting, from data stored in a first storage device of one of the two or more processing zones, identifying information that identifies the one or more child jobs associated with the parent job.
3 . The method of claim 2 , wherein identifying performance data for the one or more child jobs associated with the parent job further comprises determining that the identifying information that identifies the one or more child jobs and second identifying information stored in a second storage device of a different one of the two or more processing zones share a common prefix.
4 . The method of claim 1 , wherein the one or more child jobs were created by the parent job.
5 . The method of claim 1 , further comprising correlating, by the data processing hardware, first output data associated with a first child job of the one or more child jobs and second output data associated with a second child job of the one or more child jobs.
6 . The method of claim 5 , further comprising determining, by the data processing hardware, performance data for the parent job based on the first output data associated with the first child job and the second output data associated with the second child job.
7 . The method of claim 6 , further comprising
determining, by the data processing hardware, that the performance data satisfies performance criteria; and in response to determining that the performance data satisfies the performance criteria, triggering, by the data processing hardware, an action to be performed.
8 . The method of claim 7 , wherein triggering the action to be performed comprises providing a notification based on a result of comparing performance data for the parent job to the performance criteria.
9 . The method of claim 1 , wherein the performance data for the one or more child jobs comprises one or more of: a running time; memory usage; CPU time; disk usage; a relationship between each child job and the parent job; or one or more counters associated with the parent job.
10 . The method of claim 1 , wherein displaying the performance data of the at least one of the one or more child jobs selected by the user selection comprises displaying a interactive hierarchical structure.
11 . A system comprising:
data processing hardware of a distributed computing system; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
obtaining data associated with a parent job that has executed on the distributed computing system;
determining one or more child jobs associated with the parent job, the one or more child jobs distributed across two or more processing zones of the distributed computing system;
identifying performance data for the one or more child jobs associated with the parent job, the performance data comprising a status of each of the one or more child jobs;
receiving from a user of the distributed computing system, a user selection selecting performance data of at least one of the one or more child jobs; and
displaying at a display in communication with the data processing hardware, the performance data of the at least one of the one or more child jobs selected by the user selection.
12 . The system of claim 11 , wherein identifying performance data for the one or more child jobs associated with the parent job comprises detecting, from data stored in a first storage device of one of the two or more processing zones, identifying information that identifies the one or more child jobs associated with the parent job.
13 . The system of claim 12 , wherein identifying performance data for the one or more child jobs associated with the parent job further comprises determining that the identifying information that identifies the one or more child jobs and second identifying information stored in a second storage device of a different one of the two or more processing zones share a common prefix.
14 . The system of claim 11 , wherein the one or more child jobs were created by the parent job.
15 . The system of claim 11 , wherein the operations further comprise correlating first output data associated with a first child job of the one or more child jobs and second output data associated with a second child job of the one or more child jobs.
16 . The system of claim 15 , wherein the operations further comprise determining performance data for the parent job based on the first output data associated with the first child job and the second output data associated with the second child job.
17 . The system of claim 16 , wherein the operations further comprise:
determining that the performance data satisfies performance criteria; and in response to determining that the performance data satisfies the performance criteria, triggering an action to be performed.
18 . The system of claim 17 , wherein triggering the action to be performed comprises providing a notification based on a result of comparing performance data for the parent job to the performance criteria.
19 . The system of claim 11 , wherein the performance data comprises one or more of: a running time; memory usage; CPU time; disk usage; a relationship between each child job and the parent job; or one or more counters associated with the parent job.
20 . The system of claim 11 , wherein displaying the performance data of the at least one of the one or more child jobs selected by the user selection comprises displaying a interactive hierarchical structure.Join the waitlist — get patent alerts
Track US2021064505A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.