US2023334008A1PendingUtilityA1
Execution engine for executing single assignment programs with affine dependencies
Assignee: STILLWATER SUPERCOMPUTING INCPriority: May 27, 2008Filed: Jun 19, 2023Published: Oct 19, 2023
Est. expiryMay 27, 2028(~1.8 yrs left)· nominal 20-yr term from priority
Inventors:Erwinus Theodorus Leonardus Omtzigt
G06F 15/825G06F 15/17381G06F 15/8023
76
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The execution engine is a new organization for a digital data processing apparatus, suitable for highly parallel execution of structured fine-grain parallel computations. The execution engine includes a memory for storing data and a domain flow program, a controller for requesting the domain flow program from the memory, and further for translating the program into programming information, a processor fabric for processing the domain flow programming information and a crossbar for sending tokens and the programming information to the processor fabric.
Claims
exact text as granted — not AI-modified1 . A computing device comprising:
a memory for storing data and a domain flow program; a controller for requesting the data and the domain flow program from the memory; and a processor fabric for processing the data and the domain flow program via a plurality of processing elements.
2 . (canceled)
3 . (canceled)
4 . (canceled)
5 . (canceled)
6 . (canceled)
7 . (canceled)
8 . The device of claim 1 wherein for sparse matrices, index structures that use hierarchical blocks to enable n-bit indices are used to minimize memory bandwidth and maximize performance for a DRAM.
9 . The device of claim 1 wherein the controller is further configured for presenting a read request to a memory controller which translates the read request to a memory request and returns the data to the controller.
10 . (canceled)
11 . (canceled)
12 . (canceled)
13 . The device of claim 1 further comprising a crossbar configured for routing the data to rows or columns in the processor fabric.
14 . The device of claim 13 wherein the processor fabric is further configured for producing output data.
15 . The device of claim 14 wherein the output data is written to the memory by traversing through the crossbar according to a data structure descriptor definition and whereupon the data are is presented to a memory controller which writes the data into the memory.
16 - 30 . (canceled)
31 . The device of claim 1 wherein the processor fabric is programmable.
32 . The device of claim 1 wherein the processor fabric receives the data, executes instructions on the data, and produces output data.
33 . The device of claim 1 wherein the data comprises multidimensional data and is processed under control of the domain flow program which is a single assignment program.
34 . The device of claim 33 wherein the plurality of processing elements comprises a regular network of processing elements capable of executing the single assignment program.
35 . The device of claim 1 wherein the plurality of processing elements comprise a matrix of processing elements.
36 . A method comprising:
storing, in a memory of a device, data and a domain flow program; requesting, with a controller, the data and the domain flow program from the memory; and processing, with a processor fabric, the data and the domain flow program via a plurality of elements.
37 . The method of claim 36 wherein for sparse matrices, index structures that use hierarchical blocks to enable n-bit indices are used to minimize memory bandwidth and maximize performance for a DRAM.
38 . The method of claim 36 wherein the controller is further configured for presenting a read request to a memory controller which translates the read request to a memory request and returns the data to the controller.
39 . The method of claim 36 further comprising routing, with a crossbar, the data to rows or columns in the processor fabric.
40 . The method of claim 39 wherein the processor fabric is further configured for producing output data.
41 . The method of claim 40 wherein the output data is written to the memory by traversing through the crossbar according to a data structure descriptor definition and whereupon the data is presented to a memory controller which writes the data into the memory.
42 . The method of claim 36 wherein the processor fabric is programmable.
43 . The method of claim 36 wherein the processor fabric receives the data, executes instructions on the data, and produces output data.
44 . The method of claim 36 wherein the data comprises multidimensional data and is processed under control of the domain flow program which is a single assignment program.
45 . The method of claim 36 wherein the plurality of elements comprise a matrix of processing elements.
46 . The method of claim 36 wherein the plurality of elements comprise a regular network of processing elements.
47 . A processor fabric for processing domain flow programming information comprising:
a first set of processing elements; and a second set of processing elements which communicate with the first set of processing elements, wherein the first set of processing elements and the second set of processing elements are configured for processing data and a domain flow program.
48 . The processor fabric of claim 47 wherein the first set of processing elements and the second set of processing elements process an instruction set architecture configured for a specific class of algorithms.
49 . The processor fabric of claim 48 wherein the specific class of algorithms comprise hashing algorithms.
50 . The processor fabric of claim 48 wherein the specific class of algorithms comprise optimizations for interpolations and resampling.
51 . The processor fabric of claim 48 wherein the plurality of elements comprise a regular network of processing elements.
52 . The processor fabric of claim 47 wherein the first set of processing elements and the second set of processing elements are further configured for receiving multidimensional data and producing output data.
53 . The processor fabric of claim 52 wherein the output data is written to a memory by traversing through a crossbar to a memory controller which writes the data into the memory.
54 . The processor fabric of claim 47 wherein the first set of processing elements and the second set of processing elements recognize a spatial tag of a computational event and take action under control of the domain flow program.
55 . The processor fabric of claim 47 wherein the domain flow program evolves as multi-dimensional data match up with the first set of processing elements and the second set of processing elements and produce new multi-dimensional data.
56 . The processor fabric of claim 47 wherein the first set of processing elements and the second set of processing elements each include a content addressable memory, an instruction scheduling/dispatch queue, one or more functional units, and a router configured to generate affine routing vectors.
57 . The processor fabric of claim 47 wherein the processor fabric manages and maintains global operators.
58 . The processor fabric of claim 47 wherein the processor fabric implements testing and reconfiguring to identify and isolate a faulty processing element of the first set of processing elements or the second set of processing elements or a faulty storage element.
59 . The processor fabric of claim 47 wherein the processor fabric is implemented on a chip with one or more additional processor fabrics which are able to communicate data to each other.
60 . The processor fabric of claim 47 wherein the first set of processing elements and the second set of processing elements each comprise a Single Instruction Multiple Data (SIMD) unit configured to process a plurality of operations per instruction.Join the waitlist — get patent alerts
Track US2023334008A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.