System and method for fault identification in an electronic system based on context-based alarm analysis
Abstract
A fault identification system consisting of multiple reasoning engines and the blackboard analyzes alarm information and the associated contextual information to identify faults. The contextual information associated with an alarm is derived by analyzing the alarm along four spaces, namely, transaction-space, function-space, execution-space, and signal-space. The reasoning engines associated with these spaces infer and/or validate the occurrences of faults. Transaction reasoning engine, using the associated knowledge repository, processes the generated alarms to infer and validate faults. Monitor reasoning engine, using the associated knowledge repository, processes domain specific monitor variables to infer faults. Execution reasoning engine, using the associated knowledge repository, processes execution specific monitor variables to infer and validate faults. Function reasoning engine, using the associated knowledge repository, reasons to infer and validate faults. Signal reasoning engine, using the associated knowledge repository, processes hardware specific and environment variables to infer and validate faults. Global reasoning engine moderates the inferences and validations by other reasoning engines to provide consolidated fault inference. The invention also provides a process, “design for diagnosis,” for designing electronic systems with maximum emphasis on fault diagnosis.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A fault identification system, for efficiently identifying the faults occurring in a core electronic system based on the analysis of the observed alarm information and the state of hardware and software subsystems, comprising of means for reducing the ambiguity and complexity arising due to the enormity of the alarms generated by the core electronic system and further comprising of:
(a) a subsystem, TRE, for processing the alarms using the analysis of transaction-space related contextual information; (b) a subsystem, FRE, for analyzing the function-space related contextual information; (c) a subsystem, ERE, for analyzing the execution-space related contextual information; (d) a subsystem, SRE, for analyzing the signal-space related contextual information; (e) a subsystem, MRE, for analyzing monitor variable information; (f) a subsystem, GRE, for identifying faults based on the moderation of results posted by other subsystems; (g) a subsystem, BB, to facilitate collaboration among the subsystems; and (h) a subsystem, CIC, for the collecting alarm and associated contextual space information in terms of four dimensional spaces, namely, Transaction-space, Function-space, Execution-space and Signal-space.
2 . The system of claim 1 , wherein said TRE subsystem, comprises of a procedure for transaction-wise alarm processing.
3 . The system of claim 2 , wherein said TRE subsystem further comprises of a procedure for usecase-wise alarm processing.
4 . The system of claim 2 , wherein said TRE subsystem further comprises of a procedure to use inter-relation within alarms as alarm maps and groups of alarms with temporal relation as annotations for alarm processing.
5 . The system of claim 2 , wherein said TRE subsystem further comprises of a procedure for analyzing monitor variables specific to an alarm wherein the behavior of the alarm-specific monitor variable provides support for the inference of a fault or occurrence of the alarm.
6 . The system of claim 2 , wherein said TRE subsystem further comprises of a procedure to use a set of rules, associated with each annotation as pre- and post-condition, and post-action for the annotation, for alarm processing.
7 . The system of claim 2 , wherein said TRE subsystem further comprises of a procedure to use the knowledge repository of plurality of information comprising of transaction information, usecase information, alarm maps, annotation information, pre- and post-conditions, and post-actions associated with each of the annotations, and AMV data and associated rules for alarm processing.
8 . The system of claim 2 , wherein said TRE subsystem further comprises of means for online processing of observed alarms in a transaction to derive segments of alarm sequences.
9 . The system of claim 8 , wherein said means for online processing of observed alarms to derive segments of alarm sequences further comprises of means to identify the missing alarms in a derived segment by comparing with the annotation associated with the transaction and resolving the ambiguity arising out of missing of alarms by analyzing alarm specific monitor variables along with specified rules.
10 . The system of claim 2 , wherein said TRE subsystem further comprises of means to infer the occurrence of a fault based on the analysis of annotations identified during alarm processing along with their pre- and post-conditions, and post-actions.
11 . The system of claim 2 , wherein said TRE subsystem further comprises of means to validate the occurrence of a fault inferred by other subsystems, based on the analysis of AMVs associated with the inferred fault.
12 . The system of claim 1 , wherein said FRE subsystem, further comprises a procedure to use the function space information by identifying the function space element associated with a generated alarm in a transaction.
13 . The system of claim 12 , wherein said FRE subsystem further comprises of a procedure to use the knowledge repository of plurality of information comprising of functions-rules information and monitor function information for the inference and validation of faults.
14 . The system of claim 12 , wherein said FRE subsystem further comprises of a procedure to collect plurality of information from the core electronic system through a software interface and provide the same to blackboard subsystem.
15 . The system of claim 12 , wherein said FRE subsystem further comprises of means to infer faults based on the analysis of results of a monitor function implemented in the core electronic system for the purposes of assessing the behavior of the corresponding critical function.
16 . The system of claim 12 , wherein said FRE subsystem further comprises of means to validate the occurrence of a fault inferred by other subsystems based on the analysis of learned rules that correlate the function-alarm associations with fault occurrences.
17 . The system of claim 16 , wherein said FRE subsystem further comprises of means to learn rules for correlating function-alarm associations with faults based on the positive examples on rectification of the identified faults.
18 . The system of claim 1 , wherein said SRE subsystem, comprises of a procedure to use the signal space information by identifying the signal space elements associated with an alarm for the inference and validation of faults.
19 . The system of claim 18 , wherein said SRE subsystem further comprises of a procedure to use the knowledge repository of plurality of information comprising of resource information, component information, hardware specific signatures, hardware specific monitor variable information and hardware specific rules for the inference and validation of faults.
20 . The system of claim 18 , wherein said SRE subsystem further comprises of a procedure to collect plurality of information from the core electronic system through a software interface and provide the same to blackboard subsystem.
21 . The system of claim 18 , wherein said SRE subsystem further comprises of means to infer faults based on the analysis of hardware monitor variables and environmental monitor variables along with the associated rules.
22 . The system of claim 18 , wherein said SRE subsystem further comprises of means to validate the occurrence of a fault inferred by other subsystems based on the aging analysis of the hardware components associated with the inferred fault.
23 . The system of claim 1 , wherein said MRE subsystem, comprises of a procedure to use domain specific monitor variables associated with faults, the associated unique signatures and the associated rules for the inference of faults.
24 . The system of claim 23 , wherein said MRE subsystem further comprises of a procedure to use the knowledge repository of plurality of information comprising of domain specific monitor variable data, domain specific signatures and domain specific rules for the inference of faults.
25 . The system of claim 1 , wherein said ERE subsystem, comprises of a procedure to use the execution space information by identifying the execution space element associated with an alarm for the inference and validation of faults.
26 . The system of claim 25 , wherein said ERE subsystem further comprises of a procedure to use the knowledge repository of plurality of information comprising of execution specific signatures, execution specific monitor variable information and execution specific rules for the inference and validation of faults.
27 . The system of claim 25 , wherein said ERE subsystem further comprises of means to infer faults based on the analysis of execution monitor variables along with plurality of associated rules.
28 . The system of claim 25 , wherein said ERE subsystem further comprises of means to validate the occurrence of a fault inferred by other subsystems based on the comparison of trends of execution specific monitor variables with the learned signatures for the corresponding execution specific monitor variables.
29 . The system of claim 25 , wherein said ERE subsystem further comprises of means to learn a set of signatures for each execution specific monitor variable based on the positive examples on rectification of the identified faults.
30 . The system of claim 1 , wherein said GRE subsystem, comprises of means to moderate the inferences and validations posted by various subsystems to derive a consolidated fault inference.
31 . The system of claim 30 , wherein said GRE subsystem further comprises of a procedure to use the knowledge repository of plurality of information comprising of fault information and inference-validation table to derive consolidated fault inferences.
32 . The system of claim 30 , wherein said GRE subsystem further comprises of means to learn a correction factor for the inferences made by the various subsystems based on positive and negative examples on rectification/rejection of the identified faults.
33 . The system of claim 1 , wherein said CIC subsystem, comprises of means to collect contextual information related to alarms in terms of transaction, function, and usecase information and to collect various monitor variables in the system.
34 . An apparatus, for efficiently identifying the faults occurring in a core electronic system based on the analysis of observed alarm information and the state of hardware and software subsystems comprising of means for reducing the ambiguity and complexity arising due to the enormity of the alarms generated by the system, comprising of:
(a) a hardware subsystem for performing the identification of faults in the core electronic system; (b) a hardware subsystem for collecting the specified monitor variables from the core electronic subsystem; (c) a software subsystem for collecting the plurality of information from core electronic subsystem. (d) a software subsystem for performing the identification of faults in the core electronic system;
35 . The apparatus of claim 34 , wherein said hardware subsystem for performing identification of faults comprises of:
(a) A processor of appropriate capacity; (b) Memory devices of appropriate capacity; and (c) Interface subsystem for interacting with the core electronic system and with the knowledge repositories.
36 . The apparatus of claim 34 , wherein said hardware subsystem for collecting the specified monitor variables from the core electronic subsystem comprises of sensor appropriately located in the core electronic system.
37 . The apparatus of claim 36 , further comprises of hardware devices to facilitate the collection of internally defined monitor variables from meta-components in the core electronic system wherein a meta-component has been suitably design to provide the internally defined monitor variables to the hardware devices.
38 . The apparatus of claim 34 , wherein said software subsystem for collecting the plurality of information from the core electronic subsystem comprises of software agents implemented as part of software components of the core electronic system.
39 . The apparatus of claim 34 , wherein said software subsystem for performing the identification of faults in the core electronic system comprises of software to process the information, collected from the software agents, using the knowledge repositories.
40 . A method for efficiently identifying the faults occurring in a core electronic system based on the analysis of observed alarm information and the state of hardware and software subsystems comprising of means for reducing the ambiguity and complexity arising due to the enormity of the alarms generated by the system, comprising the step of diagnosis-oriented designing of the electronic system for reducing the ambiguity and complexity arising due to the enormity of the alarms generated by the system.
41 . The method of claim 40 further comprises of one of the steps as the identification of components and meta-components information comprising of aging parameters and operating conditions; resource hierarchy; and hardware specific monitor variables and environmental variables along with associated rules and signatures wherein the said identification is based on the analysis of system specification and component data by a group system designers.
42 . The method of claim 41 further comprises of a step to use the simulation results, and test and operation data of a prototype by a group of system analysts to derive appropriate rules and signatures.
43 . The method of claim 40 further comprises of one of the steps as the identification of the faults and fault-component inter-relations of the core electronic system based on the failure mode analysis by a group of experts.
44 . The method of claim 40 further comprises of one of the steps as the identification of usecases, transactions of a usecase, alarms and annotations associated with each transaction based on the software specification, software design and function graphs by a group of software design specialists.
45 . The method of claim 44 further comprises of one of the steps as the identification of pre- and post-conditions, and post-actions associated with annotations; and identification of alarm specific monitor variable along with associated signatures and rules based on transaction and resource hierarchies.
46 . The method of claim 40 further comprises of one of the steps as the identification of domain specific monitor variables along with associated signatures and rules and identification and designing of monitor functions for critical functions based on system specification and system design by a group of system designers.
47 . The method of claim 40 further comprises of one of the steps as the identification of execution specific monitor variables along with associated signatures and rules based on software execution environment by a group of software specialists.Join the waitlist — get patent alerts
Track US2004010733A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.