US2022206796A1PendingUtilityA1

Multi-functional execution lane for image processor

Assignee: GOOGLE LLCPriority: Dec 4, 2015Filed: Mar 10, 2022Published: Jun 30, 2022
Est. expiryDec 4, 2035(~9.3 yrs left)· nominal 20-yr term from priority
G06F 7/57G06F 9/3887G06F 9/3001G06F 15/80G06F 9/3893G06F 9/30014G06F 9/3877G06F 9/3885
70
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An apparatus is described that includes an execution unit having a multiply add computation unit, a first ALU logic unit and a second ALU logic unit. The ALU unit is to perform first, second, third and fourth instructions. The first instruction is a multiply add instruction. The second instruction is to perform parallel ALU operations with the first and second ALU logic units operating simultaneously to produce different respective output resultants of the second instruction. The third instruction is to perform sequential ALU operations with one of the ALU logic units operating from an output of the other of the ALU logic units to determine an output resultant of the third instruction. The fourth instruction is to perform an iterative divide operation in which the first ALU logic unit and the second ALU logic unit operate during to determine first and second division resultant digit values.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A processing unit comprising:
 a first pair of input ports, a second pair of input ports, and two output ports;   a first arithmetic-logic unit (ALU) configured to receive data from the first pair of input ports; and   a second ALU configured to receive data from the second pair of input ports,   wherein the processing unit is configured to decode instructions that have a single ALU opcode to select between parallel and sequential operation of the first ALU and the second ALU, wherein when a first value of the ALU opcode is received, the processing unit is configured to operate the first ALU and the second ALU in sequence, and wherein when a second value of the ALU opcode is received, the processing unit is configured to operate the first ALU and the second ALU in parallel.   
     
     
         3 . The processing unit of  claim 2 , wherein when the second value of the ALU opcode is received, the second ALU receives data from an output of the first ALU. 
     
     
         4 . The processing unit of  claim 2 , wherein the value of the ALU opcode specifies whether the first ALU and the second ALU perform the same operations or different operations. 
     
     
         5 . The processing unit of  claim 2 , wherein the value of the ALU opcode specifies that the first ALU and the second ALU are to perform four operations in parallel. 
     
     
         6 . The processing unit of  claim 2 , wherein the value of the ALU opcode specifies a fused operation during which the first ALU generates a result to be used as input to the second ALU. 
     
     
         7 . The processing unit of  claim 2 , wherein the value of the ALU opcode specifies a double width ALU operation during which the first ALU generates a carry term to be used by the second ALU. 
     
     
         8 . The processing unit of  claim 2 , wherein the value of the ALU opcode specifies that the first ALU and the second ALU are to perform two full-width operations in parallel. 
     
     
         9 . The processing unit of  claim 2 , wherein the value of the ALU opcode specifies that the first ALU and the second ALU are to perform four half-width operations in parallel. 
     
     
         10 . A method implemented by a processing unit of an image processor comprising an array of processing units, each processing unit comprising a first pair of input ports, a second pair of input ports, two output ports, a first arithmetic-logic unit (ALU) configured to receive data from the first pair of input ports, and a second ALU configured to receive data from the second pair of input ports, the method comprising:
 decoding instructions that have a single ALU opcode to select between parallel and sequential operation of the first ALU and the second ALU, including:
 operating the first ALU and the second ALU in sequence when a first value of the ALU opcode is received; and
 operating the first ALU and the second ALU in parallel when a second value of the ALU opcode is received. 
 
   
     
     
         11 . The method of  claim 10 , wherein when the second value of the ALU opcode is received, the second ALU receives data from an output of the first ALU. 
     
     
         12 . The method of  claim 10 , wherein the value of the ALU opcode specifies whether the first ALU and the second ALU perform the same operations or different operations. 
     
     
         13 . The method of  claim 10 , wherein the value of the ALU opcode specifies that the first ALU and the second ALU are to perform four operations in parallel. 
     
     
         14 . The method of  claim 10 , wherein the value of the ALU opcode specifies a fused operation during which the first ALU generates a result to be used as input to the second ALU. 
     
     
         15 . The method of  claim 10 , wherein the value of the ALU opcode specifies a double width ALU operation during which the first ALU generates a carry term to be used by the second ALU. 
     
     
         16 . The method of  claim 10 , wherein the value of the ALU opcode specifies that the first ALU and the second ALU are to perform two full-width operations in parallel. 
     
     
         17 . The method of  claim 10 , wherein the value of the ALU opcode specifies that the first ALU and the second ALU are to perform four half-width operations in parallel. 
     
     
         18 . An image processor comprising an array of processing units, wherein each processing unit comprises:
 a first pair of input ports, a second pair of input ports, and two output ports;   a first arithmetic-logic unit (ALU) configured to receive data from the first pair of input ports; and   a second ALU configured to receive data from the second pair of input ports,   wherein each processing unit is configured to decode instructions that have a single ALU opcode to select between parallel and sequential operation of the first ALU and the second ALU, wherein when a first value of the ALU opcode is received, the processing unit is configured to operate the first ALU and the second ALU in sequence, and wherein when a second value of the ALU opcode is received, the processing unit is configured to operate the first ALU and the second ALU in parallel.   
     
     
         19 . The image processor of  claim 18 , wherein when the second value of the ALU opcode is received, the second ALU receives data from an output of the first ALU. 
     
     
         20 . The image processor of  claim 18 , wherein the value of the ALU opcode specifies whether the first ALU and the second ALU perform the same operations or different operations. 
     
     
         21 . The image processor of  claim 18 , wherein the value of the ALU opcode specifies that the first ALU and the second ALU are to perform.

Join the waitlist — get patent alerts

Track US2022206796A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.