US2017358053A1PendingUtilityA1

Parallel processor with integrated correlation and convolution engine

Assignee: NVIDIA CORPPriority: Jan 8, 2013Filed: Aug 28, 2017Published: Dec 14, 2017
Est. expiryJan 8, 2033(~6.5 yrs left)· nominal 20-yr term from priority
G06F 17/153G06T 1/20
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for performing computer algorithms. The system includes a graphics pipeline operable to perform graphics processing and an engine operable to perform at least one of a correlation determination and a convolution determination for the graphics pipeline. The graphics pipeline is further operable to execute general computing tasks. The engine comprises a plurality of functional units operable to be configured to perform at least one of the correlation determination and the convolution determination. In one embodiment, the engine is coupled to the graphics pipeline. The system further includes a configuration module operable to configure the engine to perform at least one of the correlation determination and the convolution determination.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system comprising:
 a memory; and   a graphics processing unit (GPU) coupled to said memory and comprising:
 a streaming processor; 
 a graphics pipeline associated with said streaming processor and operable to perform graphics processing, wherein said graphics pipeline is further operable to execute general computing tasks; and 
 a computation engine coupled said streaming processor and said graphics pipeline and dedicated to at least one of correlation computations and convolution computations, wherein said computation engine is configured to perform at least one of a correlation computation and a convolution computation in conjunction with said streaming processor. 
   
     
     
         2 . The system of  claim 1 , wherein said GPU further comprises another streaming processor coupled to said computation engine, wherein said computation engine is further configured to perform at least one of correlation computations and convolution computations for said streaming processor and said another streaming processor on an on-demand basis based on respective workloads of said streaming processor and said another streaming processor. 
     
     
         3 . The system of  claim 1 , wherein said streaming processor is operable to offload said at least one of said correlation computation and said convolution computation to said computation engine. 
     
     
         4 . The system of  claim 1 , wherein said computation engine further comprises a plurality of execution units configurable to perform said at least one of said correlation computation and said convolution computation based on an application programming interface (API) call. 
     
     
         5 . The system of  claim 4 , wherein said computation engine further comprises a configuration module configured to:
 determine a data flow between said plurality of execution units;   load data received from said streaming processor to said computation engine; and   select between a fixed configuration and a flexible configuration for said computation engine to perform said at least one of said correlation computation and said convolution computation.   
     
     
         6 . The system of  claim 4 , wherein each of said plurality of execution units comprises circuit elements and is configured to perform one of a plurality of functions related to said at least one of said correlation computation and said convolution computation. 
     
     
         7 . The system of  claim 1 , wherein said streaming processor is configured to switch to process a first thread while waiting for a second thread to be performed by said computation engine, wherein said first and said second threads are different. 
     
     
         8 . The system of  claim 1 , wherein said computation engine is configured to compute a template matching process by:
 precomputing a template average;   computing an internal image; and   using summed area tables to determine an average of an image tile.   
     
     
         9 . The system of  claim 1 , wherein said computation engine is further configured to compute a filter bank function with a single read operation of source image pixels. 
     
     
         10 . The system of  claim 1 , wherein said computation engine is configured to provide pixel computation units at a computation precision of 8-bits or lower. 
     
     
         11 . A method of accelerating computations in a graphics processing unit (GPU) by using a computation engine, said method comprising:
 at said computation engine, receiving a request to perform at least one of a correlation computation and a convolution computation;   determining a configuration of a plurality of execution units in said computation engine, wherein said configuration corresponds to at least one of said correlation computation and said convolution computation;   at said computation engine, performing at least one of said correlation computation and said convolution computation based on said configuration to generate a result; and   supplying aid result of said at least one of said correlation computation and said convolution computation to a processing unit of said GPU.   
     
     
         12 . The method of  claim 11 , wherein said determining said configuration comprises one of:
 determining a data flow between said plurality of execution units;   loading data received from said processing unit of said GPU to said computation engine; and   selecting between a fixed configuration and a flexible configuration for said computation engine to perform said at least one of said correlation computation and convolution computation.   
     
     
         13 . The method of  claim 11  further comprising said processing unit offloading said at least one of said correlation computation and said convolution computation to said computation engine. 
     
     
         14 . The method of  claim 11 , wherein said performing at least one of said correlation computation and said convolution computation comprises computing a template matching process by:
 precomputing a template average;   computing an internal image; and   using summed area tables to determine an average of an image tile.   
     
     
         15 . The method of  claim 11 , wherein said performing at least one of said correlation computation and said convolution computation comprises computing a filter bank function with a single read operation of source image pixels. 
     
     
         16 . The method of  claim 11  further comprising, at said computation engine, performing at least one of a correlation computation and a convolution computation for another processing unit of said GPU, wherein said computation engine is used on an on-demand basis based on workloads of said processing unit and said another processing unit. 
     
     
         17 . The method of  claim 11 , wherein said performing said at least one of said correlation computation and said convolution computation comprises one of:
 precomputing a portion of at least one of said correlation computation and said convolution computation; and   postcomputing a portion of at least one of said correlation computation and said convolution computation.   
     
     
         18 . The method of  claim 11  further comprising said processing unit switching to process a first thread while waiting for a second thread to be performed by said computation engine, wherein said first thread and said second thread are different. 
     
     
         19 . The method of  claim 11 , wherein said request is received from a graphics pipeline of said GPU. 
     
     
         20 . The method of  claim 19  further comprising said graphics pipeline precomputing or postcomputing a portion of said at least one of said correlation computation and said convolution computation.

Join the waitlist — get patent alerts

Track US2017358053A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.