US2011173169A1PendingUtilityA1

Methods To Perform Disk Writes In A Distributed Shared Disk System Needing Consistency Across Failures

Assignee: ORACLE INT CORPPriority: Mar 7, 2001Filed: Mar 23, 2011Published: Jul 14, 2011
Est. expiryMar 7, 2021(expired)· nominal 20-yr term from priority
Y10S707/99953G06F 12/0804G06F 12/0866G06F 12/0815G06F 11/1469Y10S707/99952G06F 12/0817G06F 11/1471Y10S707/99954
38
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Techniques are provided for managing caches in a system with multiple caches that may contain different copies of the same data item. Specifically, techniques are provided for coordinating the write-to-disk operations performed on such data items to ensure that older versions of the data item are not written over newer versions, and to reduce the amount of processing required to recover after a failure. Various approaches are provided in which a master is used to coordinate with the multiple caches to cause a data item to be written to persistent storage. Techniques are also provided for managing checkpoints associated with the caches, where the checkpoints are used to determine the position at which to begin processing recovery logs in the event of a failure.

Claims

exact text as granted — not AI-modified
1 . A method for managing versions of a data item, the method comprising the steps of:
 when a dirty version of a data item is transferred from a first node to a second node while a being-written version of the data item is being written to persistent storage, performing the steps of:
 communicating version information about the being-written version to the second node; and 
 based on the version information, the second node preventing any version of the data item that belongs to a first set of versions from being merged with any version of the data item that belongs to a second set of versions; 
 wherein the first set of versions includes all versions of the data item within the second node that are at least as old as the being-written version; and 
 wherein the second set of versions includes versions of the data item within the second node that are newer than the being-written version. 
   
     
     
         2 . The method of  claim 1  wherein the step of communicating is performed by a master assigned to said data item. 
     
     
         3 . The method of  claim 1  wherein:
 the second node includes a plurality of versions in said first set; and 
 the second node merges said plurality of versions. 
 
     
     
         4 . The method of  claim 1  further comprising the steps of:
 informing the second node when the being-written version has been successfully written to persistent storage; and 
 after the second node has been informed that the being-written version has been successfully written to persistent storage, allowing said second node to discard all versions in said first set of versions. 
 
     
     
         5 . The method of  claim 3  further comprising the steps of:
 informing the second node when the being-written version has been successfully written to persistent storage; and 
 after the second node has been informed that the being-written version has been successfully written to persistent storage, allowing said second node to discard a merged version created by merging said plurality of versions. 
 
     
     
         6 . A method for managing past images of a data item, the method comprising the steps of:
 estimating a likelihood that a first past version of a data item will soon be written to persistent storage or covered by a write to persistent storage;   if the estimated likelihood is exceeds a particular threshold, then storing a second past version of the data item separate from the first past version of the data item; and   if the estimated likelihood falls below a particular threshold, then merging the second past version of the data item with the first past version of the data item.   
     
     
         7 . The method of  claim 6  wherein the step of estimating is based on a comparison between a time associated with the first past version of the data item and a time associated with a recent entry in a redo log file. 
     
     
         8 . The method of  claim 6  wherein the step of estimating is based on a comparison between a time associated with the first past version of the data item and a time associated with an entry at the head of a checkpoint queue. 
     
     
         9 . A computer-readable medium carrying instructions for managing versions of a data item, the instructions comprising instructions for performing the steps of:
 when a dirty version of a data item is transferred from a first node to a second node while a being-written version of the data item is being written to persistent storage, performing the steps of:
 communicating version information about the being-written version to the second node; and 
 based on the version information, the second node preventing any version of the data item that belongs to a first set of versions from being merged with any version of the data item that belongs to a second set of versions; 
 wherein the first set of versions includes all versions of the data item within the second node that are at least as old as the being-written version; and 
 wherein the second set of versions includes versions of the data item within the second node that are newer than the being-written version. 
   
     
     
         10 . The computer-readable medium of  claim 9  wherein the step of communicating is performed by a master assigned to said data item. 
     
     
         11 . The computer-readable medium of  claim 9  wherein:
 the second node includes a plurality of versions in said first set; and 
 the second node merges said plurality of versions. 
 
     
     
         12 . The computer-readable medium of  claim 9  further comprising instructions for performing the steps of:
 informing the second node when the being-written version has been successfully written to persistent storage; and 
 after the second node has been informed that the being-written version has been successfully written to persistent storage, allowing said second node to discard all versions in said first set of versions. 
 
     
     
         13 . The computer-readable medium of  claim 11  further comprising instructions for performing the steps of:
 informing the second node when the being-written version has been successfully written to persistent storage; and 
 after the second node has been informed that the being-written version has been successfully written to persistent storage, allowing said second node to discard a merged version created by merging said plurality of versions. 
 
     
     
         14 . A computer-readable medium carrying instructions for managing past images of a data item, the instructions comprising instructions for performing the steps of:
 estimating a likelihood that a first past version of a data item will soon be written to persistent storage or covered by a write to persistent storage;   if the estimated likelihood is exceeds a particular threshold, then storing a second past version of the data item separate from the first past version of the data item; and   if the estimated likelihood falls below a particular threshold, then merging the second past version of the data item with the first past version of the data item.   
     
     
         15 . The computer-readable medium of  claim 14  wherein the step of estimating is based on a comparison between a time associated with the first past version of the data item and a time associated with a recent entry in a redo log file. 
     
     
         16 . The computer-readable medium of  claim 14  wherein the step of estimating is based on a comparison between a time associated with the first past version of the data item and a time associated with an entry at the head of a checkpoint queue.

Join the waitlist — get patent alerts

Track US2011173169A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.