US2006036896A1PendingUtilityA1

Method and system for consistent cluster operational data in a server cluster using a quorum of replicas

Assignee: MICROSOFT CORPPriority: Mar 26, 1999Filed: Aug 12, 2005Published: Feb 16, 2006
Est. expiryMar 26, 2019(expired)· nominal 20-yr term from priority
G06F 11/182G06F 11/1482G06F 11/1662G06F 11/181G06F 11/2023G06F 11/2035G06F 11/1425
49
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and system for increasing server cluster availability by requiring at a minimum only one node and a quorum replica set of replica members to form and operate a cluster. Replica members, independent from the nodes, maintain cluster operational data. A cluster operates when one node possesses a majority of replica members, which ensures that any new or surviving cluster includes consistent cluster operational data via at least one replica member from the immediately prior cluster. Arbitration provides exclusive ownership by one node of the replica members, including at cluster formation, and when the owning node fails. Arbitration uses a fast mutual exclusion algorithm and a reservation mechanism to challenge for and defend the exclusive reservation of each member. A quorum replica set algorithm brings members online and offline with data consistency, including updating unreconciled replica members, and ensures consistent read and update operations.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method, comprising: 
 maintaining cluster operational data on a replica set comprising a plurality of replica members that are each independent of any node of a server cluster;    representing the cluster at a node if the number of replica members controlled by the node comprises at least a majority of the total number of replica members configured to operate in the cluster; and    determining which of the replica members of the replica set has operational data that is most updated, and replicating at least some of that operational data to the other replica members of the replica set.    
     
     
         2 . The method of  claim 1  wherein determining which of the replica members of the replica set has the most updated operational data includes, maintaining an epoch number in association with each replica member.  
     
     
         3 . The method of  claim 2  wherein a value of each epoch number indicates a relative state of the cluster operational data on its respective replica member, and wherein determining which of the replica members of the replica set has operational data that is most updated includes determining which of the epoch numbers from each member is largest.  
     
     
         4 . The method of  claim 3  wherein at least two members have epoch numbers that equal the largest epoch number, and wherein determining which of the replica members of the replica set has the most updated operational data includes, maintaining a sequence number in association with the cluster operational data, and determining a largest sequence number from the replica members that have epoch numbers that equal the largest.  
     
     
         5 . The method of  claim 1  further comprising, evaluating a last record logged on a replica member to which data is being replicated, against at least one record of the replicated data, to determine whether to discard the last record.  
     
     
         6 . The method of  claim 5  further comprising, evaluating a second-to-last record logged on the replica member to which data is being replicated, against at least one record of the replicated data, to determine whether to discard the second-to-last record.  
     
     
         7 . The method of  claim 1  further comprising, detecting availability of a new replica member that is configured to operate in the cluster, and reconciling the cluster operational data of the new replica member.  
     
     
         8 . The method of  claim 1  further comprising, detecting the unavailability of a replica member that was operational, determining whether the majority of replica members still exists, and if not, halting updates to the cluster configuration data and executing a recovery process to attempt to obtain control of a majority of replica members.  
     
     
         9 . The method of  claim 1  wherein maintaining the cluster operational data includes storing information indicative of a total number of replica members configured in the cluster, and/or storing state of at least one other storage device of the cluster.  
     
     
         10 . The method of  claim 1  wherein the node controls the majority of replica members by arbitrating for exclusive ownership of each member.  
     
     
         11 . The method of  claim 10  wherein arbitrating for exclusive ownership includes executing a mutual exclusion algorithm.  
     
     
         12 . The method of  claim 1  wherein the node controls the majority of replica members by arbitrating for exclusive ownership of each member, including, issuing a reset command.  
     
     
         13 . A computer-implemented method, comprising: 
 storing cluster operational data on a replica set of at least one replica member, each replica member being independent from any node;    at a first node, arbitrating with at least two other nodes for control of the replica set, the arbitration being performed for each replica member and comprising, attempting to obtain a right to exclusively reserve that replica member, and if the attempt is successful, exclusively reserving that replica member; and    representing a server cluster at the first node if the replica set is controlled by the first node and has consistent cluster operational data with respect to a previous cluster.    
     
     
         14 . The method of  claim 13  wherein the replica set comprises a plurality of replica members, and wherein the replica set is controlled and has consistent cluster operational data with respect to the previous cluster when a majority of replica members is exclusively reserved.  
     
     
         15 . The method of  claim 13  wherein attempting to obtain a right to exclusively reserve that replica member includes, executing a mutual exclusion algorithm.  
     
     
         16 . The method of  claim 13  wherein attempting to obtain a right to exclusively reserve that replica member includes, attempting to write a unique identifier to a location on the replica member, delaying, and reading from the location to determine whether the unique identifier is unchanged.  
     
     
         17 . The method of  claim 13  wherein arbitration is performed by challenging at the first node for ownership of the replica set when the first node does not represent the cluster.  
     
     
         18 . The method of  claim 13  further comprising, defending exclusive ownership of the replica set at the first node after control of the cluster is achieved.  
     
     
         19 . The method of  claim 13  further comprising determining which of the replica members of the replica set has cluster operational data that is most updated, and replicating that operational data to the other replica members of the replica set.  
     
     
         20 . The method of  claim 13  wherein arbitrating for each replica member includes breaking a reservation of the replica member by another node.  
     
     
         21 . The method of  claim 13  wherein arbitrating for each replica member includes, issuing a reset command for the replica member, delaying for a period of time, and issuing a reserve command for the replica member.  
     
     
         22 . A computer-readable medium having computer-executable instructions, comprising: 
 representing a cluster by obtaining exclusive control of a majority of replica members in an available set of replica members;    detecting a status change of one replica member with respect to the available set; and    taking action in response to the changed status to ensure that the replica members are consistent with respect to any update logged thereto.    
     
     
         23 . The computer-readable medium of  claim 22  wherein taking action in response to the changed status includes running a recovery process to make the replica members consistent, including increasing an epoch number maintained on each available replica member.  
     
     
         24 . The computer-readable medium of  claim 22  wherein detecting a status change includes detecting that the one replica member is offline, and wherein taking action in response to the changed status includes determining whether a majority of replica members still exists, and if a majority of replica members does not still exist, preventing updates from being written to replica members that remain available.

Join the waitlist — get patent alerts

Track US2006036896A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.