Methods and systems for executing vectorized pythagorean tuple instructions
Abstract
Disclosed embodiments relate generally to computer processor architecture, and, more specifically, to methods and systems for executing vectorized Pythagorean tuple instructions. In one example, a processor includes fetch circuitry to fetch an instruction having an opcode, an order, a destination identifier, and N source identifiers, N being equal to the order, and the order being one of two, three, and four, decode circuitry to decode the fetched instruction, and execution circuitry, for each element of the identified destination, to generate N squares by squaring each corresponding element of the N identified sources and generate a sum of the N squares and previous contents of the element.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor comprising:
fetch circuitry to fetch an instruction having an opcode, an order, a destination identifier, and N source identifiers, N being equal to the order, and the order being one of two, three, and four; decode circuitry to decode the fetched instruction; and execution circuitry, for each element of the identified destination, to:
generate N squares by squaring each corresponding element of the N identified sources; and
generate a sum of the N squares and previous contents of the element.
2 . The processor of claim 1 , wherein the execution circuitry uses a chain of N two-way fused multiply adders to generate the N squares and the sum.
3 . The processor of claim 1 , wherein the execution circuitry uses N two-input multipliers to generate the N squares in parallel, and uses a N-plus-one-input adder to generate the sum.
4 . The processor of claim 1 , wherein the order is specified by one of the opcode, an opcode prefix, an opcode suffix, and an immediate.
5 . The processor of claim 1 , wherein each element of the identified destination and the N identified sources comprises a fixed size, the instruction further comprising a precision operand to specify the fixed size.
6 . The processor of claim 1 , wherein each element of the identified destination and the N identified sources comprises a floating point value.
7 . The processor of clam 1 , wherein the instruction further comprises a writemask, the writemask being a multi-bit value with each bit to control, for each element of the identified destination, whether the sum is stored to the element.
8 . The processor of claim 1 , wherein the destination identifier and the N source identifiers each specifies a vector register having a vector length, wherein the vector length is selected from a group consisting of 128 bits, 256 bits, and 512 bits, and wherein the instruction further specifies the vector length using one of the opcode, a prefix to the opcode, and an immediate.
9 . The processor of claim 1 , wherein the identified destination is zeroed after reset.
10 . The processor of claim 1 , wherein the execution circuit is to execute the decoded instruction over multiple cycles, processing a subset of the elements of the identified destination on each cycle.
11 . A method comprising:
fetching, using fetch circuitry, an instruction having an opcode, an order, a destination identifier, and N source identifiers, N being equal to the order, and the order being one of two, three, and four; decoding, using decode circuitry, the fetched instruction; and executing, by execution circuitry, to, for each element of the identified destination:
generate N squares by squaring each corresponding element of the N identified sources; and
generate a sum of the N squares and previous contents of the element.
12 . The method of claim 11 , further comprising using, by the execution circuit, a chain of N two-way fused multiply adders to generate the N squares and the sum.
13 . The method of claim 11 , further comprising using, by the execution circuit, N two-input multipliers to generate the N squares in parallel, and a N-plus-one-input adder to generate the sum.
14 . The method of claim 11 , wherein the order is specified by one the opcode, an opcode prefix, an opcode suffix, and an immediate.
15 . The method of claim 11 , wherein each element of the identified destination and the N identified sources comprises a fixed size, the instruction further comprising a precision operand to specify the fixed size.
16 . An apparatus comprising:
means for fetching an instruction having an opcode, an order, a destination identifier, and N source identifiers, N being equal to the order, and the order being one of two, three, and four; means for decoding the fetched instruction; and means for executing to, for each element of the identified destination:
generate N squares by squaring each corresponding element of the N identified sources; and
generate a sum of the N squares and previous contents of the element.
17 . The apparatus of claim 16 , wherein the means for executing uses a chain of N two-way fused multiply adders to generate the N squares and the sum.
18 . The apparatus of claim 16 , wherein the means for executing uses N two-input multipliers to generate the N squares in parallel, and uses a N-plus-one-input adder to generate the sum.
19 . The apparatus of claim 16 , wherein the order is specified by one the opcode, an opcode prefix, an opcode suffix, and an immediate.
20 . The apparatus of claim 16 , wherein each element of the identified destination and the N identified sources comprises a fixed size, the instruction further comprising a precision operand to specify the fixed size.
21 . A non-transitory computer-readable medium containing instructions that, when execute by a processor, cause the processor to:
fetch, using fetch circuitry, an instruction having an opcode, an order, a destination identifier, and N source identifiers, N being equal to the order, and the order being one of two, three, and four; decode, using decode circuitry, the fetched instruction; and execute, by execution circuitry, to, for each element of the identified destination:
generate N squares by squaring each corresponding element of the N identified sources; and
generate a sum of the N squares and previous contents of the element.
22 . The non-transitory computer-readable medium of claim 21 , further comprising using, by the execution circuit, a chain of N two-way fused multiply adders to generate the N squares and the sum.
23 . The non-transitory computer-readable medium of claim 21 , further comprising using, by the execution circuit, N two-input multipliers to generate the N squares in parallel, and a N-plus-one-input adder to generate the sum.
24 . The non-transitory computer-readable medium of claim 21 , wherein the order is specified by one the opcode, an opcode prefix, an opcode suffix, and an immediate.
25 . The non-transitory computer-readable medium of claim 21 , wherein each element of the identified destination and the N identified sources comprises a fixed size, the instruction further comprising a precision operand to specify the fixed size.Join the waitlist — get patent alerts
Track US2019102199A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.