US2023259436A1PendingUtilityA1

Systems and methods for monitoring application health in a distributed architecture

Assignee: TORONTO DOMINION BANKPriority: Jul 10, 2020Filed: Apr 25, 2023Published: Aug 17, 2023
Est. expiryJul 10, 2040(~13.9 yrs left)· nominal 20-yr term from priority
G06F 11/3006G06F 11/327G06F 11/0772G06F 11/079G06F 11/323G06F 11/0751G06F 11/3055G06F 11/3476G06F 2201/875
52
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A computing device configured for monitoring and analyzing health of a distributed computer system having a plurality of interconnected system components. The computing device tracks communication between the system components and monitors for an alert indicating an error in the communication in the distributed computer system. In response to the error, the computing device receives a health log from each of the system components defining an aggregate health log being in a standardized format indicating messages communicated between the system components. The computing device further receives network infrastructure information defining relationships between the system components and characterizing dependency information; and, automatically determines, based on the aggregate health log and the network infrastructure information, a particular component originating the error and associated dependent components from the system components affected.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computing device for monitoring health of a distributed computer system, the computing device having a processor coupled to a memory, the memory storing instructions which when executed by the processor configure the computing device to:
 track communication between a plurality of interconnected system components of the distributed computer system and monitor for an alert indicating an error in the communication, responsive to detecting the error:
 receive a health log from each of the system components, each said health log being in a standardized format indicating messages communicated between the system components; 
 capture, from each said health log, common key identifiers for tracing a route of the messages communicated for a transaction having the error; 
 receive network infrastructure information defining relationships for connectivity between the system components, the relationships characterizing dependency information between the system components; and 
 automatically determine, based on applying the network infrastructure information to the health logs including the common key identifiers and further mapping to a set of health monitoring rules comprising data integrity information, a particular component of the system components originating the error and associated dependent components affected. 
   
     
     
         2 . The computing device of  claim 1 , further comprising obtaining the health monitoring rules from a data store wherein the data integrity information is for pre-defined communications between the system components, the instructions configuring the computing device to apply the set of health monitoring rules for verifying whether each said health log comprising the common key identifiers complies with the data integrity information. 
     
     
         3 . The computing device of  claim 2 , wherein the health monitoring rules are further defined based on historical error patterns, derived from historical health logs, for the distributed computer system associating a set of traffic flows potentially occurring for the messages communicated between the system components as derived from respective common key identifiers in the historical health logs to a corresponding error type for the error pattern. 
     
     
         4 . The computing device of  claim 1 , wherein the common key identifiers, common to the system components communicating in a particular transaction, link a particular task to the messages communicated for that particular task and depict the route of the messages communicated between the system components for the particular task having the error. 
     
     
         5 . The computing device of  claim 1  wherein the common key identifiers further identify the system components and types of events or messages communicated for each transaction. 
     
     
         6 . The computing device of  claim 1 , wherein the common key identifiers link two or more parties affecting a transaction. 
     
     
         7 . The computing device of  claim 1 , wherein the common key identifiers comprise key metadata that interconnects the system components via an entity function role. 
     
     
         8 . The computing device of  claim 1 , wherein the instructions configure the computing device to modify the common key identifiers each time it is processed or communicated by one of the system components to identify a path taken by the messages. 
     
     
         9 . The computing device of  claim 1 , wherein the instructions further configure the computing device to:
 determine from the dependency information indicating which of the system components are dependent on one another for operations performed in the distributed computer system, an impact of the error originated by the particular component on the associated dependent components.   
     
     
         10 . The computing device of  claim 9 , wherein the instructions further configure the computing device to perform, upon detecting the alert:
 displaying the alert on a user interface of a client application for the device, the alert based on the particular component originating the error.   
     
     
         11 . The computing device of  claim 10 , further comprising: displaying on the user interface along with the alert, the associated dependent components to the particular component originating the error. 
     
     
         12 . The device of  claim 1 , wherein the standardized format comprises a JSON format. 
     
     
         13 . The device of  claim 1 , wherein the system components are APIs (application programming interfaces) on one or more connected computing devices and the health log is an API log for logging activity for the respective API in communication with other APIs and related to the error. 
     
     
         14 . The device of  claim 1 , wherein the processor configuring the computing device to automatically determine origin of the error further comprises: comparing each of the health logs in an aggregate health log to the other health logs in response to the relationships in the network infrastructure information. 
     
     
         15 . A method implemented by a computing device, the method for monitoring health of a distributed computer system, the method comprising:
 tracking communication between a plurality of interconnected system components of the distributed computer system and monitor for an alert indicating an error in the communication, responsive to detecting the error:
 receiving a health log from each of the system components, each said health log being in a standardized format indicating messages communicated between the system components; 
 capturing, from each said health log, common key identifiers for tracing a route of the messages communicated for a transaction having the error; 
 receiving network infrastructure information defining relationships for connectivity between the system components, the relationships characterizing dependency information between the system components; and 
 automatically determining, based on applying the network infrastructure information to the health logs including the common key identifiers and further mapping to a set of health monitoring rules comprising data integrity information, a particular component of the system components originating the error and associated dependent components affected. 
   
     
     
         16 . The method of  claim 15 , further comprising obtaining the health monitoring rules from a data store wherein the data integrity information is for pre-defined communications between the system components, the set of health monitoring rules being applied for verifying whether each said health log comprising the common key identifiers complies with the data integrity information. 
     
     
         17 . The method of  claim 16 , wherein the health monitoring rules are further defined based on historical error patterns, derived from historical health logs, for the distributed computer system associating a set of traffic flows potentially occurring for the messages communicated between the system components as derived from respective common key identifiers in the historical health logs to a corresponding error type for the error pattern. 
     
     
         18 . The method of  claim 15 , wherein the common key identifiers, common to the system components communicating in a particular transaction, link a particular task to the messages communicated for that particular task and depict the route of the messages communicated between the system components for the particular task having the error. 
     
     
         19 . The method of  claim 15 , wherein the common key identifiers further identify the system components and types of events or messages communicated for each transaction. 
     
     
         20 . The method of  claim 19  wherein the common key identifiers link two or more parties affecting a transaction. 
     
     
         21 . The method of  claim 15 , wherein the common key identifiers comprise key metadata that interconnects the system components via an entity function role. 
     
     
         22 . The method of  claim 15 , further comprising updating the common key identifiers each time it is processed or communicated by one of the system components to identify a path taken by the messages. 
     
     
         23 . The method of  claim 15 , further comprising:
 determining from the dependency information indicating which of the system components are dependent on one another for operations performed in the distributed computer system, an impact of the error originated by the particular component on the associated dependent components.   
     
     
         24 . The method of  claim 23 , further comprising upon detecting the alert:
 displaying the alert on a user interface of a client application for the device, the alert based on the particular component originating the error.   
     
     
         25 . The method of  claim 24 , further comprising: displaying on the user interface along with the alert, the associated dependent components to the particular component. 
     
     
         26 . The method of  claim 25 , wherein the standardized format comprises a JSON format. 
     
     
         27 . The method of  claim 15 , wherein the system components are APIs (application programming interfaces) on one or more connected computing devices and the health log is an API log for logging activity for the respective API in communication with other APIs and related to the error. 
     
     
         28 . The method of  claim 15 , wherein automatically determining origination of the further comprises: comparing each of the health logs in an aggregate health log to the other health logs in response to the relationships in the network infrastructure information. 
     
     
         29 . A computer readable medium comprising a non-transitory device storing instructions and data, which when executed by a processor of a computing device, the processor coupled to a memory, configure the computing device to:
 track communication between a plurality of interconnected system components of a distributed computer system and monitor for an alert indicating an error in the communication in the distributed computer system, and upon detecting the error:
 receive a health log from each of the system components, each said health log being in a standardized format indicating messages communicated between the system components; 
 capture, from each said health log, common key identifiers for tracing a route of the messages communicated for a transaction having the error; 
 receive network infrastructure information defining relationships for connectivity between the system components, the relationships characterizing dependency information between the system components; and 
 automatically determine, based on applying the network infrastructure information to the health logs including the common key identifiers and further mapping to a set of health monitoring rules comprising data integrity information, a particular component of the system components originating the error and associated dependent components affected.

Join the waitlist — get patent alerts

Track US2023259436A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.