US2008189225A1PendingUtilityA1

Method and System for Predicting Causes of Network Service Outages Using Time Domain Correlation

Assignee: HERRING DAVIDPriority: Nov 28, 2000Filed: Mar 24, 2008Published: Aug 7, 2008
Est. expiryNov 28, 2020(expired)· nominal 20-yr term from priority
H04L 41/064H04L 43/106H04L 43/10H04L 43/00H04L 43/0817H04L 41/5003H04L 41/5025H04L 41/20G06Q 30/0283H04L 41/147
40
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and systems are described for predicting the likely causes of service outages using only time information, and for predicting and the likely costs of service outages. The likely causes are found by defining a narrow likely cause window around an outage based on service quality and/or service usage data, and correlating service events to the likely cause window in the time domain to find a probability distribution for the events. The likely costs are found by measuring usage loss and duration for a given point during an outage and using cost component functions of the time and usage to extrapolate over the outage. These cause and cost predictions supply service administrators with tools for making more informed decisions about allocation of resources in preventing and correcting service outages.

Claims

exact text as granted — not AI-modified
1 . A method for quantifying the effect of an outage in a service over a first period of time, the method comprising:
 measuring usage of the service over time;   defining a cost of outage time window comprising the first time period and a second time period following the first time period; and   computing a cost of outage as the difference between the measured service usage during the cost of outage time window with service usage measured during a comparison window, the comparison window being substantially equal in time to that of the cost of outage time window and reflecting a similar period of service activity as that of the cost of outage time window without having a service outage.   
     
     
         2 . The method of  claim 1 , comprising determining the second period of time to be a time in which the measured service usage returns to within a given percentage of a normal service usage. 
     
     
         3 . The method of  claim 1 , comprising determining the second period of time to be the shorter of (1) a time in which the measured service usage returns to within a given percentage of a normal service usage and (2) a maximum time period. 
     
     
         4 . The method of  claim 1 , wherein computing the cost of outage comprises computing the difference in units of service usage. 
     
     
         5 . The method of  claim 4 , wherein the service is a communication service conveying a plurality of messages, the method comprising computing the cost of outage in numbers of messages conveyed. 
     
     
         6 . The method of  claim 4 , wherein the service is a network server providing data items in response to requests therefor, the method comprising computing the cost of outage in numbers of requests received or data items provided by a server on the network. 
     
     
         7 . The method of  claim 4 , comprising converting the computed units of cost of service outage to a monetary value. 
     
     
         8 . The method of  claim 7 , wherein converting the computed units of cost of service outage comprises multiplying the units of cost of service outage by a first monetary value per unit of usage. 
     
     
         9 . The method of  claim 1 , comprising comparing the cost of outage to a second cost of outage value for a different service and prioritizing the outages based on the compared costs. 
     
     
         10 . The method of  claim 1 , comprising computing the difference between the monitored service usage following the cost of outage time window and a normal service usage level to thereby measure a long term effect of the service outage. 
     
     
         11 . A method for quantifying the effect of an outage in a service, the method comprising:
 measuring usage amounts of the service during a period of the service outage and a second period following the service outage;   comparing the measured usage amounts to normal usage amounts measured under similar service conditions for a similar period of time where no service outage occurs; and   determining a level of loss of service due to the service outage based on the comparison.   
     
     
         12 . The method of  claim 11 , comprising defining the second period as the shorter of a time period in which measured service usage amounts return to within a given range of normal usage amounts and a predefined maximum time period. 
     
     
         13 . The method of  claim 11 , wherein measuring service usage amounts comprises measuring service usage amounts in terms of units of service usage. 
     
     
         14 . The method of  claim 13 , wherein the service is a communication service conveying a plurality of messages, comprising measuring service usage amounts in terms of number of messages conveyed by the system. 
     
     
         15 . The method of  claim 13 , wherein the service is a network server providing data items in response to requests therefor, comprising measuring service usage amounts in terms of numbers of requests received or data items provided by a server on the network. 
     
     
         16 . The method of  claim 11 , wherein determining the level of service loss comprises determining that substantially no loss of service occurred due to the outage based on the measured service usage amounts and normal service usage amounts being substantially equal. 
     
     
         17 . The method of  claim 11 , comprising:
 measuring service usage amounts during a third period following the second period;   comparing the measured third period service usage amounts to normal usage amounts measured under similar service conditions for a similar period of time; and   determining a long term effect on the service due to the service outage based on the comparison.   
     
     
         18 . A computer readable medium storing program code which, when executed, causes a computer to perform a method for quantifying the effect of an outage in a service over a first period of time, the method comprising:
 measuring service usage over time;   defining a cost of outage time window comprising the first time period and a second time period; and   computing a cost of outage as the difference between the measured level of service usage during the cost of outage time window with a level of usage in a comparison window, the comparison window being substantially equal in time to the cost of outage time window and reflecting a similar period of service activity as the cost of outage time window without having a service outage.   
     
     
         19 . A method for predicting a cost of an outage of a service, the method comprising:
 measuring time duration for and service usage during the outage;   comparing the measured usage amounts to normal usage amounts measured under similar service conditions for a similar period of time where no service outage occurs, to thereby determine a usage loss amount; and   computing a predicted cost of the outage based at least upon a cost component, the cost component comprising a function of the measured time of the outage and measured usage loss amount.   
     
     
         20 . The method of  claim 19 , comprising measuring service usage on an ongoing basis and detecting the onset of the service outage using the measured service usage. 
     
     
         21 . The method of  claim 20 , wherein detecting the onset of the service outage comprises detecting a step change in service usage. 
     
     
         22 . The method of  claim 19 , comprising monitoring quality of the service and detecting the onset of a service outage based upon the service quality. 
     
     
         23 . The method of  claim 22 , wherein monitoring service quality comprises monitoring service quality through periodic polling of the service quality, and wherein detecting the onset of a service outage comprises detecting the outage onset as bounded by a polled point of a first, working state and a polled point of a second, non-working state. 
     
     
         24 . The method of  claim 19 , wherein computing the predicted cost of the outage comprises using a service demand cost component representing an affect on service usage based upon the duration of an outage. 
     
     
         25 . The method of  claim 24 , wherein using the service demand cost component comprises multiplying the measured usage loss by a usage loss curve which is a function of time duration of an outage and represents a predicted percentage usage due to an outage based on time duration of the outage. 
     
     
         26 . The method of  claim 25 , comprising generating the usage loss curve using historical data derived from prior service outages. 
     
     
         27 . The method of  claim 19 , wherein computing the predicted cost of the outage comprises using a customer retention cost component representing a number or percentage of customers lost due to the outage. 
     
     
         28 . The method of  claim 19 , wherein computing the predicted cost of the outage comprises using an agreement penalty component representing penalties arising under one or more service agreements due to a service outage. 
     
     
         29 . The method of  claim 19 , wherein computing the predicted cost of the outage comprises computing the cost in units of service usage. 
     
     
         30 . The method of  claim 29 , comprising converting the computed units of predicted cost to a monetary value by multiplying the units of predicted cost by a first monetary value per unit of usage. 
     
     
         31 . The method of  claim 19 , comprising comparing the predicted cost of service outage to a second predicted cost of outage value for a different service and prioritizing the outages based on the compared costs. 
     
     
         32 . A network monitoring system comprising:
 a usage meter for measuring usage of a service on the network;   an event detector for detecting network events and times at which the network events occur;   a probable cause engine, coupled to receive data from the usage meter and the event detector, for determining which of the network events detected by the event detector is the most likely cause of a service outage based at least in part of the relations of the detected network event times to a service change time window, the service change time window encompassing at least part of an occurrence of the service outage in the network; and   a costing engine, coupled to receive data from the usage meter, for predicting the cost of the service outage.   
     
     
         33 . Computer readable media comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for quantifying the effect of an outage in a service, the method comprising:
 measuring usage amount of the service during a period of the service outage and a second period following the service outage;   comparing the measured usage amounts to normal usage amounts measured under similar service conditions for a similar period of time where no service outage occurs; and   determining a level of loss of service due to the service outage based on the comparison.   
     
     
         34 . The computer readable media of  claim 33  comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for quantifying the effect of an outage in a service, the method comprising defining the second period as the shorter of a time period in which measured service usage amounts return to within a given range of normal usage amounts and a predefined maximum time period. 
     
     
         35 . The computer readable media of  claim 33  comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for quantifying the effect of an outage in a service, wherein measuring service usage amounts comprises measuring service usage amounts in terms of units of service usage. 
     
     
         36 . The computer readable media of  claim 35  comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for quantifying the effect of an outage in a service, wherein the service is a communications service conveying a plurality of messages, comprising measuring service usage amounts in terms of number of messages conveyed by the system. 
     
     
         37 . The computer readable media of  claim 35  comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for quantifying the effect of an outage in a service, wherein the service is a network server providing data items in response to requests therefore, comprising measuring service usage amounts in terms of number of requests received or data items provided by a server on the network. 
     
     
         38 . The computer readable media of  claim 33  comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for quantifying the effect of an outage in a service, wherein determining the level of service loss comprises determining that substantially no loss of service occurred due to the outage based on the measured service usage amounts and normal service usage amounts being substantially equal. 
     
     
         39 . The computer readable media of  claim 33  comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for quantifying the effect of an outage in a service, the method comprising:
 measuring service usage amounts during a third period following the second period;   comparing the measured third period service usage amounts to normal usage amounts measured under similar conditions for a similar period of time; and   determining a long term effect on the service due to the service outage based on the comparison.   
     
     
         40 . Computer readable media comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for predicting a cost of an outage of a service, the method comprising:
 measuring time duration for and service usage during the outage;   comparing the measured usage amounts to normal usage amounts measured under similar service conditions for s similar period of time where no service outage occurs, to thereby determine a usage loss amount; and   computing a predicted cost of the outage based at least upon a cost component. the cost component comprising a function of the measured time of the outage and measured usage loss amount.

Join the waitlist — get patent alerts

Track US2008189225A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.