Self-testing and -repairing fault-tolerance infrastructure for computer systems
Abstract
ASICs or like fabrication-preprogrammed hardware provide controlled power and recovery signals to a computing system that is made up of commercial, off-the-shelf components—and that has its own conventional hardware and software fault-protection systems, but these are vulnerable to failure due to external and internal events, bugs, human malice and operator error. The computing system preferably includes processors and programming that are diverse in design and source. The hardware infrastructure uses triple modular redundancy to test itself as well as the computing system, and to remove failed elements—powering up and loading data into spares. The hardware is very simplified in design and programs, so that bugs can be thoroughly rooted out. Communications between the protected system and the hardware are protected by very simple circuits with duplex redundancy.
Claims
exact text as granted — not AI-modified1 - 66 . (canceled)
67 . Apparatus for deterring failure of an entire computing system, said computing system being distinct from the apparatus and including at least one processor; wherein the apparatus comprises:
a hardware network of components, having no software, and having no firmware except programs optionally held in an optional unalterable read-only memory; a terminals of the network for connection to the computing system; and fabrication-preprogrammed hardware circuits of the network for guarding the entire computing system, including the at least one processor, from failure.
68 . The apparatus of claim 67 , further comprising:
an unalterable read-only memory holding programs for operation as firmware.
69 . The apparatus of claim 67 , wherein:
each of the at least one processor is a Pentium or IBM G5 processor; or a more-recent processor, or equivalent; and said circuits guard the at least one Pentium or IBM processor, or more-recent processor, or equivalent, against failure.
70 . The apparatus of claim 67 , particularly for use with a system that is capable of generating an error signal in event of incipient failure, and is capable of responding to a recovery signal; and wherein:
at least one of the network terminals is connected to receive at least one error signal generated by each module of the entire system, respectively, including the at least one processor, in event of incipient failure of the module; at least one of the network terminals is connected to provide at least one recovery signal to the respective module upon receipt of the error signal; and the apparatus further comprises means for automatically responding to the at least one error signal by generating the at least one recovery signal for guarding each module of the system against failure.
71 . The apparatus of claim 67 , wherein:
the network is generic in that it can accommodate any computing system whose modules can issue respective error messages and handle respective recovery commands.
72 . The apparatus of claim 67 , wherein:
the circuits are not capable of running any application program.
73 . Apparatus for deterring failure of an entire computing system, said computing system being distinct from the apparatus and including at least one processor; wherein the apparatus comprises:
a hardware network of components, having no software, and having no firmware except programs optionally held in an optional unalterable read-only memory; terminals of the network for connection to the computing system; and an unalterable read-only memory holding programs for operation as firmware, to guard the entire computing system, including the at least one processor, from failure.
74 . Fault-tolerant apparatus comprising:
a computing system, including at least one processor; a hardware network of components, having no software, and having no firmware except optionally programs held in an optional unalterable read-only memory; terminals of the network for connection to the computing system; and fabrication-preprogrammed hardware circuits of the network for guarding the entire computing system, including the at least one processor, from failure.
75 . Apparatus for deterring failure of an entire computing system, including all processor chips, memory chips, and other computing-system modules (other than communications modules) that are present in the system, wherein the computing system optionally includes plural mutually redundant modules; said apparatus comprising:
a network of components having terminals for connection to the system, wherein the network is constructed to be initially and permanently distinct from the computing system including all redundant modules if present; and circuits of the network for operating programs to deter failure of the entire computing system, including the chips and other system modules that are present; the circuits further comprising portions for identifying failure of any of the circuits and correcting for the identified failure.
76 . The apparatus of claim 75 , wherein each of the program-operating circuits is:
a fabrication-preprogrammed circuit, or
an unalterable read-only memory holding programs for operation as firmware.
77 . The apparatus of claim 76 , wherein:
to guard the entire system from failure, said circuits receive from the system error messages warning of incipient failure, and issue recovery commands to the system.
78 . The apparatus of claim 76 , particularly for use with a computing system that has at least one hardware subsystem for generating an error signal; and wherein:
the circuits comprise portions for reacting to the response of that hardware subsystem.
79 . The apparatus of claim 76 , wherein:
the network is an infrastructure that can accommodate any computing system that can issue an error message and handle a recovery command.
80 . The apparatus of claim 75 , wherein:
the circuits do not and cannot operate any application program; and except for receiving error messages from the computing system, the circuits are not controlled by any associated host computer that is capable of running any application program.
81 . Fault-tolerant apparatus comprising:
a computing system, that has multiple modules; a network of components having terminals for connection to the system, wherein the network is constructed to be initially and permanently distinct from the computing system including all of the modules; and circuits of the network for operating programs to guard the entire system, including all of the multiple modules, from failure; the circuits comprising portions for identifying failure of any of the circuits and correcting for the identified failure.
82 . Apparatus for deterring failure of a computing system, said system having multiple modules, each module including at least one hardware subsystem for generating an error message of the module about incipient failure; said apparatus comprising:
a network of components having terminals for connection to the system; and circuits of the network for operating programs to guard the system from failure; the circuits comprising portions for reacting to the error message of the hardware subsystem.
83 . The apparatus of claim 82 , wherein:
in response to the error message, the circuits guard the entire system, including all of the multiple modules, from failure.
84 . The apparatus of claim 82 , wherein:
the network can accommodate any system that can issue at least one error message and handle at least one recovery command.
85 . The apparatus of claim 82 , wherein said circuits:
are not capable of operating any application program; and are not controlled by any associated host computer that is capable of running any application program.
86 . The apparatus of claim 82 , particularly for use with a computing system that has at least one subsystem for genersting a response of the subsystem to failure, and that also has at least one subsystem for receiving recovery commands; and wherein:
the circuits comprise portions for interposing analysis and a corrective reaction between the response-generating subsystem and the command-receiving subsystem.
87 . Fault-tolerant apparatus comprising:
a computing system that has multiple modules, other than components devoted to intercomputer communications, each of said multiple modules including:
at least one respective hardware subsystem for generating an error message of the subsystem about incipient failure, and
each said hardware subsystem comprising processor chips and memory chips;
a network of components having terminals for connection to the system; and circuits of the network for operating firmware programs to guard the computing system, including all the multiple modules and the chips, other than components devoted to intercomputer communications, from failure; the circuits comprising portions for reacting to the error message of the hardware subsystem.
88 . Apparatus for deterring failure of an entire computing system that is distinct from the apparatus and that has plural generally parallel diverse computing channels and has at least one application-data input module, and at least one processor for running an application program; said apparatus comprising:
a network of components having terminals for connection to the system; and fabrication-preprogrammed circuits of the network for operating programs to guard against failure of the entire system, including (a) every one of the parallel computing channels, and (b) every application-data input module and (c) every application-program processor; wherein the network is constructed to be initially and permanently distinct from the computing system including (a) every one of the parallel computing channels, and every application-data input module and (b) every application-program processor, and (c) every parallel computing channel; the circuits comprising portions for comparing computational results from the parallel channels.
89 . A fault-tolerant apparatus comprising:
an entire computing system, including plural generally parallel computing channels, and at least one application-data input module, and at least one processor for running application programs; a network of components having terminals for connection to the computing system; and fabrication-preprogrammed circuits of the network for operating programs to guard against failure the entire computing system, including (a) every one of the parallel computing channels, and (b) every application-data input module, and (c) every application-program processor; wherein the network is constructed to be initially and permanently distinct from the computing system including (a) every one of the parallel computing channels, and every application-data input module and (b) every application-program processor, and (c) every parallel computing channel; the circuits comprising portions for comparing computational results from the parallel channels.
90 . The apparatus of claim 89 , wherein:
the circuits receive error messages from the computing system; the circuits return recovery messages to the computing system; and except for the two functions just recited, the circuits are not controlled by any associated host computer that is capable of running any application program.
91 . The apparatus of claim 90 wherein, to guard against failure of the entire system, including the computing channels and at least one input module and the at least one processor:
the circuits receive from the computing system error messages warning of incipient failure and issue recovery commands to the computing system.
92 . An infrastructure for a computing system that has at least one computing node (“C-node”) for running at least one application program; said infrastructure being for guarding the system against failure, and comprising:
at least one monitoring node (“M-node”) for monitoring the condition of the at least one C-node by waiting for an error signal, indicating incipient failure, from the at least one C-node and responding to the error signal by sending a recovery command to the at least one C-node; and at least one adapter node (“A-node”) for transmitting the error signal and recovery command between the at least one C-node and at least one M-node; and wherein: the at least one M-node is manufactured, and remains, wholly distinct from the at least one C-node; and the at least one M-node cannot, and does not, run any application program.
93 . The infrastructure of claim 92 , particularly for use with a computing system that has plural C-nodes; and further comprising:
a decision-making node (“D-node”) for comparing output data generated by the plural C-nodes and reporting to the at least one M-node any discrepancy between the output data; and wherein: the at least one M-node analyzes the D-node reporting, and based thereon arbitrates among the C-nodes.
94 . A fault-tolerant apparatus, comprising:
a computing system that has at least one computing node (“C-node”) for running at least one application program; an infrastructure guarding the computing system against failure and comprising: at least one monitoring node (“M-node”) for monitoring the condition of the at least one C-node by waiting for an error signal, indicating incipient failure, from the at least one C-node and responding to the error signal by sending a recovery command to the at least one C-node; and at least one adapter node (“A-node”) for transmitting the error signal and recovery command between the at least one C-node and at least one M-node; and wherein: the at least one M-node is manufactured, and remains, wholly distinct from the at least one C-node; and the at least one M-node cannot, and does not, run any application program.Join the waitlist — get patent alerts
Track US2010218035A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.