US2019213006A1PendingUtilityA1

Multi-functional execution lane for image processor

Assignee: GOOGLE LLCPriority: Dec 4, 2015Filed: Jan 18, 2019Published: Jul 11, 2019
Est. expiryDec 4, 2035(~9.4 yrs left)· nominal 20-yr term from priority
G06F 7/57G06F 9/30014G06F 9/3001G06F 15/80G06F 9/3887G06F 9/3885G06F 9/3893G06F 9/3877
61
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An apparatus is described that includes an execution unit having a multiply add computation unit, a first ALU logic unit and a second ALU logic unit. The ALU unit is to perform first, second, third and fourth instructions. The first instruction is a multiply add instruction. The second instruction is to perform parallel ALU operations with the first and second ALU logic units operating simultaneously to produce different respective output resultants of the second instruction. The third instruction is to perform sequential ALU operations with one of the ALU logic units operating from an output of the other of the ALU logic units to determine an output resultant of the third instruction. The fourth instruction is to perform an iterative divide operation in which the first ALU logic unit and the second ALU logic unit operate during to determine first and second division resultant digit values.

Claims

exact text as granted — not AI-modified
1 .- 19 . (canceled) 
     
     
         20 . An image processor comprising an array of processing units, wherein each processing unit of the array of processing units comprises:
 four input ports and two output ports; and   a first arithmetic-logic unit (ALU) and a second ALU configured to perform a double-width ALU operation, during which:
 the first ALU is configured to receive data from a first pair of input ports, to perform a first full-width ALU operation to compute (i) a lower half result of the double-width ALU operation and (ii) a carry term, to provide the lower half result of the double-width ALU operation to one of the two output ports, and to provide the carry term to the second ALU, and 
 the second ALU is configured to receive data from a second pair of input ports and receive the carry term from the first ALU, to perform a second full-width ALU operation to compute an upper half result of the double-width ALU operation, and to provide the upper half result of the double-width ALU operation to another of the two output ports. 
   
     
     
         21 . The image processor of  claim 20 , wherein the each processing unit is configured to perform the second full-width ALU operation after the first full-width ALU operation is complete. 
     
     
         22 . The image processor of  claim 21 , wherein each processing unit has a carry line between the first ALU and the second ALU to provide the carry term to the second ALU. 
     
     
         23 . The image processor of  claim 22 , wherein the second ALU is configured to perform the second full-width ALU operation only upon receiving the carry term on the carry line. 
     
     
         24 . The image processor of  claim 20 , wherein the first ALU and the second ALU of each processing unit are further configured to perform four half-width ALU operations at least partially in parallel, during which:
 the first ALU and the second ALU are each configured to receive input operands from a respective pair of input ports, to perform a first-half width operation on a lower half of each of the input operands, to perform a second half-width operation on an upper half of each of the input operands, and to write a result to a respective one of the two output ports.   
     
     
         25 . The image processor of  claim 20 , wherein the first ALU and the second ALU of each processing unit are further configured to perform a fused operation comprising a second operation performed serially on the result of a first operation, during which:
 the first ALU is configured to receive data from the first pair of input ports, to perform the first operation, and to provide a result of the first operation to the second ALU; and   the second ALU is configured to receive data from one input port of the second pair of input ports and to receive the result of the first operation from the first ALU, to perform the second operation, and to provide a result of the second operation to one of the two output ports.   
     
     
         26 . The image processor of  claim 25 , wherein the first operation and the second operation are different. 
     
     
         27 . A method implemented by a processing unit of an image processing comprising an array of processing units, the method comprising:
 performing, by a first arithmetic-logic unit (ALU) and a second ALU of the processing unit, a double-width ALU operation using data received at a first pair of input ports and a second pair of input ports of the processing unit, including:
 receiving, by the first ALU, data from the first pair of input ports of the processing unit, 
 performing, by the first ALU, a first full-width ALU operation using the data from the first pair of input ports to compute a lower half result of the double-width ALU operation and a carry term, 
 providing, by the first ALU, the lower half result of the double-width ALU operation to one of two output ports of the processing unit, 
 providing, by the first ALU, the carry term to the second ALU, 
 receiving, by the second ALU, data from the second pair of input ports of the processing unit, 
 receiving, by the second ALU, the carry term from the first ALU, 
 performing, by the second ALU, a second full-width ALU operation using the data from the second pair of input ports and the carry term to compute an upper half result of the double-width ALU operation, and 
 providing, by the second ALU, the upper half result of the double-width ALU operation to another of the two output ports. 
   
     
     
         28 . The method of  claim 27 , wherein performing the second full-width ALU operation comprises performing the second full-width ALU operation after the first full-width ALU operation is complete. 
     
     
         29 . The method of  claim 28 , wherein each processing unit has a carry line between the first ALU and the second ALU to provide the carry term to the second ALU. 
     
     
         30 . The method of  claim 29 , wherein performing the second full-width ALU operation performing the second full-width ALU operation only upon receiving the carry term on the carry line. 
     
     
         31 . The method of  claim 27 , further comprising:
 performing, by the first ALU and the second ALU, four half-width ALU operations at least partially in parallel, including:   receiving, by the first ALU and the second ALU, respective input operands from a respective pair of input ports,   performing, by the first ALU and the second ALU, a first-half width operation on a lower half of each of the input operands,   performing, by the first ALU and the second ALU, a second half-width operation on an upper half of each of the input operands, and   writing, by the first ALU and the second ALU, a result to a respective one of the two output ports.   
     
     
         32 . The method of  claim 27 , further comprising:
 performing, by the first ALU and the second ALU, a fused operation comprising a second operation performed serially on the result of a first operation, including:   receiving, by the first ALU, data from the first pair of input ports,   performing, by the first ALU, the first operation, and   providing, by the first ALU, a result of the first operation to the second ALU,   receiving, by the second ALU, data from one input port of the second pair of input ports,   receiving, by the second ALU, the result of the first operation from the first ALU,   performing, by the second ALU, the second operation, and   providing, by the second ALU, a result of the second operation to one of the two output ports.   
     
     
         33 . The method of  claim 32 , wherein the first operation and the second operation are different. 
     
     
         34 . An image processor comprising an array of processing units, wherein each processing unit of the array of processing units is configured to perform a double-width ALU operation, wherein each processing unit comprises:
 four input ports and two output ports; and   means for performing a first full-width ALU operation using data received at a first pair of the input ports to write a lower half result of the double-width ALU operation to a first output port and to generate a carry term; and   means for performing a second full-width ALU operation using the carry term and data received at a second pair of the input ports and to write an upper half result of the double-width ALU operation to a second output port.   
     
     
         35 . The image processor of  claim 34 , wherein each processing unit is configured to perform the second full-width ALU operation after the first full-width ALU operation is complete. 
     
     
         36 . The image processor of  claim 35 , wherein each processing unit has a carry line between the means for performing the first full-width ALU operation and the means for performing the second full-width ALU operation. 
     
     
         37 . The image processor of  claim 36 , wherein the means for performing the second full-width ALU operation is configured to perform the second full-width ALU operation only upon receiving the carry term on the carry line.

Join the waitlist — get patent alerts

Track US2019213006A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.