Technologies for auto-discovery of fault domains
Abstract
A compute device to auto-discover power system fault domains within a computer network is provided. The compute device includes a fault domain manager to receive sled data for at least one sled of a plurality of sleds in the computer network, the sled data including sled identification data and sled health data, parse the sled identification data to identify at least one power zone wherein each power zone includes a subset of the plurality of sleds, generate a fault domain mapping using the at least one identified power zone. The compute device also includes a distributed processing adapter to convert the generated fault domain mapping into a consumable fault domain mapping that is consumable by a distributed processing software system, and provide the consumable fault domain mapping to the distributed processing software system.
Claims
exact text as granted — not AI-modified1 . A compute device to auto-discover power system fault domains within a computer network, the compute device comprising:
one or more processors; and one or more memory devices having stored therein a plurality of instructions that, when executed by the one or more processors, cause the compute device to:
receive sled data for at least one sled of a plurality of sleds in the computer network, the sled data including sled identification data and sled health data;
parse the sled identification data to identify at least one power zone wherein each power zone includes a subset of the plurality of sleds;
generate a fault domain mapping using the at least one identified power zone;
convert the generated fault domain mapping into a consumable fault domain mapping that is consumable by a distributed processing software system; and
provide the consumable fault domain mapping to the distributed processing software system.
2 . The compute device of claim 1 , wherein to receive the sled data comprises to receive an advertisement generated by the at least one sled via a link layer discovery protocol (LLDP), wherein LLDP is configured to identify one or more relationships between one or more sleds of the plurality of sleds, and wherein the sled health data includes one or more of a number of sockets, a number of cores, a number of drives, a latency metric, a reliability metric, and a memory capacity metric.
3 . The compute device of claim 1 , wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to identify, within the sled identification data, a first top-of-rack switch identifier for a first top-of-rack switch that is communicatively coupled to the at least one sled of the plurality of sleds.
4 . The compute device of claim 3 , wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to identify, within the sled identification data, a second top-of-rack switch identifier for a second top-of-rack switch that is distinct from the first top-of-rack switch, wherein the second-top-of-rack switch is communicatively coupled to the at least one sled.
5 . The compute device of claim 4 , wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to determine that the first top-of-rack switch and the second top-of-rack switch are communicatively coupled to a rack compute device.
6 . The compute device of claim 5 , wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to determine that the rack compute device is communicatively coupled to at least one power source device, wherein the at least one power source device is associated with the at least one power zone.
7 . The compute device of claim 5 , wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to determine that the rack compute device is communicatively coupled to at least one other power source device, wherein the at least one other power source device is associated with at least one other power zone.
8 . The compute device of claim 1 , wherein the computer network includes a composed node, wherein the composed node includes the at least one sled, and wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to:
determine that the at least one sled corresponds to the at least one identified power zone; and identify that the composed node corresponds to the at least one identified power zone within the fault domain mapping.
9 . The compute device of claim 1 , wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to identify, within the sled identification data, a datacenter identifier for a datacenter, wherein the datacenter includes at least one manager computer that manages at least one sled of the plurality of sleds, and wherein the manager computer is communicatively coupled to the at least one sled.
10 . The compute device of claim 1 , wherein the one or more memory devices have stored therein a plurality of instructions that, when executed by the one or more processors, further cause the compute device to:
determine a data format corresponding to the distributed processing software system; select a plug-in application based on the data format; and convert the generated fault domain mapping into a consumable fault domain mapping by executing the selected plug-in application.
11 . The compute device of claim 1 , wherein the distributed processing software system includes at least one of Apache Hadoop, Apache Cassandra, Ceph, and MongoDB.
12 . One or more computer-readable storage media comprising a plurality of instructions stored thereon that, when executed by a compute device cause the compute device to:
receive sled data for at least one sled of a plurality of sleds in the computer network, the sled data including sled identification data and sled health data; parse the sled identification data to identify at least one power zone wherein each power zone includes a subset of the plurality of sleds; generate a fault domain mapping using the at least one identified power zone; convert the generated fault domain mapping into a consumable fault domain mapping that is consumable by a distributed processing software system; and provide the consumable fault domain mapping to the distributed processing software system.
13 . The one or more computer-readable storage media of claim 12 , wherein the plurality of instructions further cause the compute device to receive an advertisement generated by the at least one sled via a link layer discovery protocol (LLDP), wherein LLDP is configured to identify one or more relationships between one or more sleds of the plurality of sleds, and wherein the sled health data includes one or more of a number of sockets, a number of cores, a number of drives, a latency metric, a reliability metric, and a memory capacity metric.
14 . The one or more computer-readable storage media of claim 12 , further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to identify, within the sled identification data, a first top-of-rack switch identifier for a first top-of-rack switch that is communicatively coupled to the at least one sled of the plurality of sleds.
15 . The one or more computer-readable storage media of claim 12 , further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to identify, within the sled identification data, a second top-of-rack switch identifier for a second top-of-rack switch that is distinct from the first top-of-rack switch, wherein the second-top-of-rack switch is communicatively coupled to the at least one sled.
16 . The one or more computer-readable storage media of claim 15 , further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to determine that the first top-of-rack switch and the second top-of-rack switch are communicatively coupled to a rack compute device.
17 . The one or more computer-readable storage media of claim 16 , further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to determine that the rack compute device is communicatively coupled to at least one power source device, wherein the at least one power source device is associated with the at least one power zone.
18 . The one or more computer-readable storage media of claim 16 , further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to determine that the rack compute device is communicatively coupled to at least one other power source device, wherein the at least one other power source device is associated with at least one other power zone.
19 . The one or more computer-readable storage media of claim 12 , wherein the computer network includes a composed node, wherein the composed node includes the at least one sled, and further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to:
determine that the at least one sled corresponds to the at least one identified power zone; and identify that the composed node corresponds to the at least one identified power zone within the fault domain mapping.
20 . The one or more computer-readable storage media of claim 12 , further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to identify, within the sled identification data, a datacenter identifier for a datacenter, wherein the datacenter includes at least one manager computer that manages at least one sled of the plurality of sleds, and wherein the manager computer is communicatively coupled to the at least one sled.
21 . The one or more computer-readable storage media of claim 12 , further comprising a plurality of instructions stored thereon that, when executed by the compute device cause the compute device to:
determine a data format corresponding to the distributed processing software system; select a plug-in application based on the data format; and convert the generated fault domain mapping into a consumable fault domain mapping by executing the selected plug-in application.
22 . The one or more computer-readable storage media of claim 12 , wherein the distributed processing software system includes at least one of Apache Hadoop, Apache Cassandra, Ceph, and MongoDB.
23 . A compute device of automatically discovering power system fault domains within a computer network, the compute device comprising:
circuitry for receiving sled data for at least one sled of a plurality of sleds in the computer network, the sled data including sled identification data and sled health data; means for parsing the sled identification data to identify at least one power zone wherein each power zone includes a subset of the plurality of sleds; means for generating a fault domain mapping using the at least one identified power zone; means for converting the generated fault domain mapping into a consumable fault domain mapping that is consumable by a distributed processing software system; and means for providing the consumable fault domain mapping to the distributed processing software system.
24 . A method of automatically discovering power system fault domains within a computer network, the method comprising:
receiving sled data for at least one sled of a plurality of sleds in the computer network, the sled data including sled identification data and sled health data; parsing the sled identification data to identify at least one power zone wherein each power zone includes a subset of the plurality of sleds; generating a fault domain mapping using the at least one identified power zone; converting the generated fault domain mapping into a consumable fault domain mapping that is consumable by a distributed processing software system; and providing the consumable fault domain mapping to the distributed processing software system.
25 . The method of claim 24 , wherein the computer network includes a composed node, wherein the composed node includes the at least one sled, and further comprising:
determining that the at least one sled corresponds to the at least one identified power zone; and identifying that the composed node corresponds to the at least one identified power zone within the fault domain mapping.
26 . The method of claim 24 , further comprising:
determining a data format corresponding to the distributed processing software system; selecting a plug-in application based on the data format; and converting the generated fault domain mapping into a consumable fault domain mapping by executing the selected plug-in application.Join the waitlist — get patent alerts
Track US2019068466A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.