US2022043884A1PendingUtilityA1
System and method for an optimized winograd convolution accelerator
Est. expiryAug 7, 2037(~11 yrs left)· nominal 20-yr term from priority
G06N 3/063G06N 3/044G06N 3/045G06N 3/0464G06F 17/144G06F 15/80G06F 17/15G06F 17/16
57
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
One embodiment provides a compute apparatus to perform machine learning operations, the compute apparatus comprising a hardware accelerator including a compute unit to perform a Winograd convolution, the compute unit configurable to perform the Winograd convolution for a first kernel size using a transform associated with a second kernel size.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . An apparatus to perform machine learning operations, the apparatus comprising: a hardware accelerator to perform a convolution operation associated with a first kernel based on a transform associated with a second kernel.
22 . The apparatus of claim 21 , wherein the convolution operation comprises a Winograd convolution operation, and wherein the transform comprises a Winograd transform.
23 . The apparatus of claim 21 , wherein the hardware accelerator comprises one or more registers to store one or more values associated with the first and second kernels.
24 . The apparatus of claim 21 , wherein the hardware accelerator is further to write input data and kernel data to memory associated with the hardware accelerator.
25 . The apparatus of claim 21 , wherein the convolution operation is performed for multiple kernel strides associated with the first and second kernels.
26 . A method comprising:
performing, by a hardware accelerator of a computing device, a convolution operation associated with a first kernel based on a transform associated with a second kernel.
27 . The method of claim 26 , wherein the convolution operation comprises a Winograd convolution operation, and wherein the transform comprises a Winograd transform.
28 . The method of claim 26 , wherein the hardware accelerator comprises one or more registers to store one or more values associated with the first and second kernels.
29 . The method of claim 26 , wherein the hardware accelerator is further to write input data and kernel data to memory associated with the hardware accelerator.
30 . The method of claim 26 , wherein the convolution operation is performed for multiple kernel strides associated with the first and second kernels.
31 . A computer-readable medium having stored thereon instructions which, when executed, cause a computing device to perform operations comprising:
performing, by a hardware accelerator of the computing device, a convolution operation associated with a first kernel based on a transform associated with a second kernel.
32 . The computer-readable medium of claim 31 , wherein the convolution operation comprises a Winograd convolution operation, and wherein the transform comprises a Winograd transform.
33 . The computer-readable medium of claim 31 , wherein the hardware accelerator comprises one or more registers to store one or more values associated with the first and second kernels.
34 . The computer-readable medium of claim 31 , wherein the hardware accelerator is further to write input data and kernel data to memory associated with the hardware accelerator.
35 . The computer-readable medium of claim 31 , wherein the convolution operation is performed for multiple kernel strides associated with the first and second kernels.Join the waitlist — get patent alerts
Track US2022043884A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.