Methods and devices for data recovery after hang detection
Abstract
A processing system includes a driver and an accelerated processing unit including a processor. The processor is configured to initiate a status check of wavefronts being executed by the accelerated processing unit responsive to receiving a status inquiry from the driver. Responsive to the status check indicating a hang, the processor is configured to employ a machine learning algorithm to selectively extract data from one or more registers of the accelerated processing unit. For example, in some cases, the one or more registers are local to one or more compute units of the accelerated processing unit. The processor is further configured to export the data from the accelerated processing unit prior to the accelerated processing unit being reset.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor configured to:
initiate a status check of wavefronts being executed by an accelerated processing unit; and responsive to the status check indicating a hang, employ a machine learning algorithm to selectively extract data from one or more registers of the accelerated processing unit.
2 . The processor of claim 1 , wherein the one or more registers are local to one or more compute units of the accelerated processing unit.
3 . The processor of claim 1 , further configured to:
export the data from the accelerated processing unit prior to a reset of the accelerated processing unit being triggered.
4 . The processor of claim 1 , wherein selectively extracting the data from the one or more registers of the accelerated processing unit comprises:
identifying one or more compute units in a processing pipeline of the accelerated processing unit responsible for the hang; and extracting data from at least one register of the one or more compute units in the processing pipeline of the accelerated processing unit.
5 . The processor of claim 4 , wherein the identifying of the one or more compute units in the processing pipeline of the accelerated processing unit responsible for the hang comprises monitoring an output of each of a plurality of compute units comprising the one or more compute units.
6 . The processor of claim 5 , wherein the output of the one or more compute units is indicative that the one or more compute units are responsible for the hang.
7 . The processor of claim 4 , wherein selectively extracting data from one or more registers of the accelerated processing unit comprises, in a first stage, prioritizing extracting data from the at least one register of the one or more compute units in the processing pipeline over other compute units of the plurality of compute units in the processing pipeline.
8 . The processor of claim 7 , wherein selectively extracting data from one or more registers of the accelerated processing unit comprises, in a second stage after the first stage, prioritizing extracting data from registers of neighboring compute units of the one or more compute units in the processing pipeline over remaining compute units of the plurality of compute units in the processing pipeline.
9 . The processor of claim 1 , wherein the status check comprises sampling an output of the accelerated processing unit over a period of time.
10 . The processor of claim 1 , further configured to:
initiate the status check responsive to receiving a status inquiry from a driver associated with the accelerated processing unit in response to a timer expiring, the timer triggered based on a last receipt of data by the driver from the accelerated processing unit.
11 . A processing system comprising:
a driver; and an accelerated processing unit comprising a processor configured to:
initiate a status check of wavefronts being executed by the accelerated processing unit responsive to receiving a status inquiry from the driver;
responsive to the status check indicating a hang, employ a machine learning algorithm to selectively extract data from one or more registers of the accelerated processing unit; and
export the data from the accelerated processing unit.
12 . The processing system of claim 11 , the driver configured to:
initiate a timer based on a last receipt of data from the accelerated processing unit; and send the status inquiry to the accelerated processing unit responsive to the timer expiring.
13 . The processing system of claim 11 , the accelerated processing unit configured to export the data from the accelerated processing unit to the driver prior to the driver initiating a reset of the accelerated processing unit.
14 . The processing system of claim 11 , wherein selectively extracting the data from the one or more registers of the accelerated processing unit comprises:
identifying one or more compute units in a processing pipeline of the accelerated processing unit responsible for the hang; and extracting data from at least one register of the one or more compute units in the processing pipeline of the accelerated processing unit.
15 . The processing system of claim 14 ,
wherein the identifying of the one or more compute units in the processing pipeline of the accelerated processing unit responsible for the hang comprises monitoring an output of each of a plurality of compute units comprising the one or more compute units, wherein the output of the one or more compute units is indicative that the one or more compute units are responsible for the hang.
16 . The processing system of claim 15 , wherein selectively extracting data from one or more registers of the accelerated processing unit comprises, in a first stage, prioritizing extracting data from the at least one register of the one or more compute units in the processing pipeline over other compute units of the plurality of compute units in the processing pipeline.
17 . The processing system of claim 16 , wherein selectively extracting data from one or more registers of the accelerated processing unit comprises, in a second stage after the first stage, prioritizing extracting data from registers of neighboring compute units of the one or more compute units in the processing pipeline over remaining compute units of the plurality of compute units in the processing pipeline.
18 . The processing system of claim 11 , the driver configured to be updated based on the data exported from the accelerated processing unit.
19 . A method comprising:
initiating, by a processor, a status check of wavefronts being executed by an accelerated processing unit; and responsive to the status check indicating a hang, employing, by the processor, a machine learning algorithm to selectively extract data from one or more registers of the accelerated processing unit.
20 . The method of claim 19 , wherein selectively extracting the data from the one or more registers of the accelerated processing unit comprises:
identifying one or more compute units in a processing pipeline of the accelerated processing unit responsible for the hang; and extracting data from at least one register of the one or more compute units in the processing pipeline of the accelerated processing unit.Join the waitlist — get patent alerts
Track US2026030088A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.