US2021326504A1PendingUtilityA1

METHODS, SYSTEMS AND APPARATUS TO IMPROVE FPGA PIPELINE EMULATION EFFICIENCY ON CPUs

Assignee: INTEL CORPPriority: Jun 28, 2017Filed: Dec 26, 2020Published: Oct 21, 2021
Est. expiryJun 28, 2037(~10.9 yrs left)· nominal 20-yr term from priority
G06F 30/331
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, apparatus, systems and articles of manufacture are disclosed to improve FPGA pipeline emulation efficiency on CPUs. An example disclosed apparatus includes a loop detector to identify a register shift loop in field programmable gate array (FPGA) code, an unroller to shift and store pipeline stages in the register shift loop to a temporary unroll array, an intermediate canceller to cancel out intermediate load and store values of the temporary unroll array to retain last shifted values of the pipeline stages, and a propagator to improve emulation efficiency of the FPGA code by generating a scalar loop of the retained last shifted values for a vectorization input.

Claims

exact text as granted — not AI-modified
1 . An apparatus to improve emulation efficiency, the apparatus comprising:
 a loop detector to identify a register shift loop in field programmable gate array (FPGA) code;   an unroller to shift and store one or more pipeline stages in the register shift loop to a temporary unroll array;   an intermediate canceller to cancel out one or more intermediate load and store values of the temporary unroll array to retain one or more last shifted values of the one or more pipeline stages; and   a propagator to improve emulation efficiency of the FPGA code by generating a scalar loop of the retained one or more last shifted values for a vectorization input.   
     
     
         2 .- 24 . (canceled) 
     
     
         25 . The apparatus as defined in  claim 1 , wherein the loop detector is to identify the register shift loop by detecting an instance of a double-nested for-loop in the FPGA code. 
     
     
         26 . The apparatus as defined in  claim 1 , wherein the propagator is to generate the scalar loop for a target central processing unit (CPU) capable of multi-way parallelization. 
     
     
         27 . The apparatus as defined in  claim 1 , further including a DEFs/USEs identifier to identify data-dependency parameters associated with the register shift loop. 
     
     
         28 . The apparatus as defined in  claim 27 , further including an instance remover to remove the data-dependency parameters from the register shift loop. 
     
     
         29 . The apparatus as defined in  claim 27 , wherein the DEFs/USEs identifier is to determine whether the data-dependency parameters associated with the register shift loop are located in at least one other loop of the FPGA code. 
     
     
         30 . The apparatus as defined in  claim 1 , further including a single instruction multiple data (SIMD) code generator to generate target code for a target processor, the target code to remove for-loops associated with data-dependency parameters associated with FPGA pipeline register shifting of the target processor. 
     
     
         31 . A computer-implemented method to improve emulation efficiency, the method comprising:
 identifying, by executing a computer instruction by a processor, a register shift loop in field programmable gate array (FPGA) code;   shifting and storing, by executing a computer instruction by the processor, one or more pipeline stages in the register shift loop to a temporary unroll array;   cancelling out, by executing a computer instruction by the processor, one or more intermediate load and store values of the temporary unroll array to retain one or more last shifted values of the one or more pipeline stages; and   improving emulation efficiency of the FPGA code, by executing a computer instruction by the processor, by generating a scalar loop of the retained one or more last shifted values for a vectorization input.   
     
     
         32 . The method as defined in  claim 31 , further including identifying the register shift loop by detecting an instance of a double-nested for-loop in the FPGA code. 
     
     
         33 . The method as defined in  claim 31 , further including generating the scalar loop for a target central processing unit (CPU) capable of multi-way parallelization. 
     
     
         34 . The method as defined in  claim 31 , further including identifying data-dependency parameters associated with the register shift loop. 
     
     
         35 . The method as defined in  claim 34 , further including removing the data-dependency parameters from the register shift loop. 
     
     
         36 . The method as defined in  claim 34 , further including determining whether the data-dependency parameters associated with the register shift loop are located in at least one other loop of the FPGA code. 
     
     
         37 . The method as defined in  claim 31 , further including generating target code for a target processor, the target code to remove for-loops associated with data-dependency parameters associated with FPGA pipeline register shifting by the target processor. 
     
     
         38 . A non-transitory computer-readable medium comprising instructions that, when executed, cause one or more processors to, at least:
 identify a register shift loop in field programmable gate array (FPGA) code;   shift and store one or more pipeline stages in the register shift loop to a temporary unroll array;   cancel out one or more intermediate load and store values of the temporary unroll array to retain one or more last shifted values of the one or more pipeline stages; and   improve emulation efficiency of the FPGA code by generating a scalar loop of the retained one or more last shifted values for a vectorization input.   
     
     
         39 . The computer-readable medium as defined in  claim 38 , wherein the instructions, when executed, further cause the one or more processors to identify the register shift loop by detecting an instance of a double-nested for-loop in the FPGA code. 
     
     
         40 . The computer-readable medium as defined in  claim 38 , wherein the instructions, when executed, further cause the one or more processors to generate the scalar loop for a target central processing unit (CPU) capable of multi-way parallelization. 
     
     
         41 . The computer-readable medium as defined in  claim 38 , wherein the instructions, when executed, further cause the one or more processors to identify data-dependency parameters associated with the register shift loop. 
     
     
         42 . The computer-readable medium as defined in  claim 41 , wherein the instructions, when executed, further cause the one or more processors to remove the data-dependency parameters from the register shift loop. 
     
     
         43 . The computer-readable medium as defined in  claim 41 , wherein the instructions, when executed, further causes the one or more processors to determine whether the data-dependency parameters associated with the register shift loop are located in at least one other loop of the FPGA code.

Join the waitlist — get patent alerts

Track US2021326504A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.