Floating-point supportive pipeline for emulated shared memory architectures
Abstract
A processor architecture arrangement for emulated shared memory (ESM) architectures, including a number of multithreaded processors each provided with interleaved inter-thread pipeline and a plurality of functional units for carrying out arithmetic and logical operations on data, wherein the pipeline includes at least two operatively parallel pipeline branches, first pipeline branch includes a first sub-group of said plurality of functional units, such as ALUs (arithmetic logic unit), arranged for carrying out integer operations, and second pipeline branch includes a second, non-overlapping sub-group of said plurality of functional units, such as FPUs (floating point unit), arranged for carrying out floating point operations, and further wherein one or more of the functional units of at least said second sub-group arranged for floating point operations are located operatively in parallel with the memory access segment of the pipeline.
Claims
exact text as granted — not AI-modified1 .- 10 . (canceled)
11 . A processor architecture arrangement for an emulated shared memory (ESM) architecture, the processor architecture arrangement comprising:
a number of multi-threaded processors each provided with an interleaved inter-thread pipeline, multiple threads being executed in cyclic, interleaved manner by each pipeline so that a number of other threads are executed while a thread references a common physically distributed and logically shared data memory of the ESM architecture, and a plurality of functional units for carrying out arithmetic and logical operations on data, wherein each interleaved inter-thread pipeline includes at least two operatively parallel pipeline branches, a first pipeline branch having a first sub-group of said plurality of functional units which are configured to and arranged for carrying out integer operations, and a second pipeline branch having a second, non-overlapping, sub-group of said plurality of functional units which are configured to and arranged for carrying out floating point operations, and wherein the first and the second pipeline branches each comprise a plurality of segments, the segments being connected in series, and wherein each segment comprises one or more of said functional units, wherein the plurality of segments of the first and second pipeline branches are parallel with each other such that parallel segments of the first and second pipeline branches begin after a same latency from a beginning of the interleaved inter-thread pipeline, wherein at least one of the functional units of the second sub-group which is configured to and arranged for performing floating point operations is located operatively in parallel with a memory access segment and operates concurrently with the memory access segment after a same latency from the beginning of the interleaved inter-thread pipeline, such that pipeline stages of associated floating-point operations are executed simultaneously with parallel pipeline stages of a memory unit taking care of pending access of shared data memory, while pipeline stages of a first segment of floating point functional units are operated in parallel with corresponding stages of a first segment of arithmetic and logic functional units and pipeline stages of a second segment of floating point functional units are operating in parallel with corresponding stages of a second segment of arithmetic and logic functional units, and wherein no overlapping exists between the functional units included in the first sub-group and the functional units included in the second sub-group, and wherein at least one of the floating point functional has a longer execution latency than at least one of the integer units.
12 . The processor architecture arrangement according to claim 11 , wherein at least one of the functional units of the first sub-group is located operatively in parallel with the memory access segment of the interleaved inter-thread pipeline.
13 . The processor architecture arrangement according to claim 11 , wherein at least two or more of the functional units of the second sub-group in the second branch are chained together in a chain, and wherein a chained functional unit may pass an operation result to a subsequent functional unit in the chain as an operand.
14 . The processor architecture arrangement according to claim 11 , wherein a number of functional units in at least one of said first and/or branch or said second branch are functionally positioned before a memory, where some functional units are in parallel, and some functional units are after the memory access segment.
15 . The processor architecture arrangement according to claim 11 , wherein at least two functional units of the second sub-group are mutually of different complexity in terms of operation execution latency.
16 . The processor architecture arrangement of claim 15 , wherein a functional unit associated with longer latency is logically located in parallel with an end portion of the memory access segment.
17 . The processor architecture arrangement according to claim 11 , wherein at least one functional unit is controllable through a number of operation selection fields of instruction words.
18 . The processor architecture arrangement according to claim 11 , wherein a number of operands for a functional unit can be determined in an operand select stage of the interleaved inter-thread pipeline in accordance with a number of operand selection fields given in an instruction word.
19 . The processor architecture arrangement according to claim 11 , wherein the second sub-group of functional units in said second branch includes at least one functional unit configured to execute at least one floating-point operation selected from the group consisting of: addition, subtraction, multiplication, division, comparison, transformation from integer to floating point, transformation from floating point to integer, square root, logarithm, and exponentiation.
20 . The processor architecture arrangement according to claim 11 , wherein a first functional unit of said second sub-group is configured to execute a plurality of floating point operations, and a second functional unit of said second sub-group is configured to execute one or more other floating point operations.Join the waitlist — get patent alerts
Track US2024004666A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.