Apparatus and method for multiply, add/subtract, and accumulate of packed data elements
Abstract
An apparatus and method for performing dual concurrent multiplications, subtraction/addition, and accumulation of packed data elements. For example one embodiment of a processor comprises: a decoder to decode an instruction to generate a decoded instruction; a first source register to store first and second packed data elements; a second source register to store third and fourth packed data elements; execution circuitry to execute the decoded instruction, the execution circuitry comprising: multiplier circuitry to multiply the first and third packed data elements to generate a first temporary product and to concurrently multiply the second and fourth packed data elements to generate a second temporary product, the first through fourth packed data elements all being a first width; circuitry to negate the first temporary product to generate a negated first product; adder circuitry to add the first negated product to a first accumulated packed data element from a third source register to generate a first result, the first result being a second width which is at least twice as large as the first width; the adder circuitry to concurrently add the second temporary product to a second accumulated packed data element to generate a second result of the second width; the first and second results to be stored in specified first and second data element positions within a destination register.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor comprising:
a decoder to decode an instruction that specifies a first register, a second register, and a third register to generate a decoded instruction; the first register to store first and second packed data elements, both of which being signed; the second register to store third and fourth packed data elements, both of which being signed; and execution circuitry to execute the decoded instruction, the execution comprising:
multiplying the first and third packed data elements to generate a first temporary product and multiplying the second and fourth packed data elements to generate a second temporary product, the first through fourth packed data elements all being a first width;
negating the first temporary product to generate a negated first product;
adding the negated first product to a first accumulated packed data element from the third register to generate a first result, the first result being a second width which is at least twice as large as the first width;
adding the second temporary product to a second accumulated packed data element from the third register to generate a second result of the second width; and
storing the first and second results in specified first and second data element positions within the third register.
2 . The processor of claim 1 , wherein the first through fourth packed data elements comprise doublewords having a width of 32 bits and the first and second accumulated packed data elements and first and second results comprise quadwords having a width of 64 bits.
3 . The processor of claim 2 , wherein the first register comprises a 128-bit register to store first and second quadwords, the first and second doublewords selected from upper 32-bits of each of the first and second quadwords, respectively, and wherein the second register comprises a 128-bit register to store third and fourth quadwords, the third and fourth doublewords selected from upper 32-bits of each of the third and fourth quadwords, respectively.
4 . The processor of claim 1 , wherein the first and second results are signed values.
5 . The processor of claim 1 , wherein the execution further comprises shifting the first and second results by a specified amount to generate first and second shifted results, wherein a number of most significant bits of the first and second results are to be selected and written to the same number of least significant bit positions within the third register.
6 . The processor of claim 5 , wherein shifting the first and second results by the specified amount is responsive to another instruction.
7 . The processor of claim 5 , wherein shifting the first and second results by the specified amount comprises shifting the first and second results left and inserting zeros in bit locations made available by the shifting of the first and second results.
8 . The processor of claim 1 , wherein the execution further comprises rounding and/or saturating the first and second results.
9 . The processor of claim 1 , wherein the execution further comprises setting a saturation flag.
10 . The processor of claim 1 , wherein the first packed data element is stored in a lower bit positions than the second packed data element in the first register, the third packed data element is stored in a lower bit position than the fourth packed data element in the second register, and the first result is stored in a lower bit position than the second result in the third register.
11 . A non-transitory machine-readable medium having program code stored thereon which, when executed by a machine, is capable of causing the machine to perform:
decoding an instruction that specifies a first register, a second register, and a third register to generate a decoded instruction; and executing the decoded instruction, wherein the execution comprises:
storing first and second packed data elements that are signed in the first register;
storing third and fourth packed data elements that are signed in the second register;
multiplying the first and third packed data elements to generate a first temporary product and multiplying the second and fourth packed data elements to generate a second temporary product, the first through fourth packed data elements all being a first width;
negating the first temporary product to generate a negated first product;
adding the negated first product to a first accumulated packed data element from the third register to generate a first result, the first result being a second width which is at least twice as large as the first width;
adding the second temporary product to a second accumulated packed data element from the third register to generate a second result of the second width; and
storing the first and second results in specified first and second data element positions within the third register.
12 . The non-transitory machine-readable medium of claim 11 , wherein the first through fourth packed data elements comprise doublewords having a width of 32 bits and the first and second accumulated packed data elements and first and second results comprise quadwords having a width of 64 bits.
13 . The non-transitory machine-readable medium of claim 11 , where the execution further comprising shifting the first and second results by a specified amount to generate first and second shifted results, wherein a number of most significant bits of the first and second results are to be selected and written to the same number of least significant bit positions within the third register.
14 . The non-transitory machine-readable medium of claim 11 , wherein the execution further comprises setting a saturation flag.
15 . The non-transitory machine-readable medium of claim 11 , wherein the first packed data element is stored in a lower bit positions than the second packed data element in the first register, the third packed data element is stored in a lower bit position than the fourth packed data element in the second register, and the first result is stored in a lower bit position than the second result in the third register.
16 . A method comprising:
decoding an instruction that specifies a first register, a second register, and a third register to generate a decoded instruction; and executing the decoded instruction, wherein the execution comprises:
storing first and second packed data elements that are signed in the first register;
storing third and fourth packed data elements that are signed in the second register;
multiplying the first and third packed data elements to generate a first temporary product and multiplying the second and fourth packed data elements to generate a second temporary product, the first through fourth packed data elements all being a first width;
negating the first temporary product to generate a negated first product;
adding the negated first product to a first accumulated packed data element from the third register to generate a first result, the first result being a second width which is at least twice as large as the first width;
adding the second temporary product to a second accumulated packed data element from the third register to generate a second result of the second width; and
storing the first and second results in specified first and second data element positions within the third register.
17 . The method of claim 16 , wherein the first through fourth packed data elements comprise doublewords having a width of 32 bits and the first and second accumulated packed data elements and first and second results comprise quadwords having a width of 64 bits.
18 . The method of claim 17 , wherein the first register comprises a 128-bit register to store first and second quadwords, the first and second doublewords selected from upper 32-bits of each of the first and second quadwords, respectively, and wherein the second register comprises a 128-bit register to store third and fourth quadwords, the third and fourth doublewords selected from upper 32-bits of each of the third and fourth quadwords, respectively.
19 . The method of claim 16 , wherein the first and second results are signed values.
20 . The method of claim 16 , wherein the execution further comprises shifting the first and second results by a specified amount to generate first and second shifted results, wherein a number of most significant bits of the first and second results are to be selected and written to the same number of least significant bit positions within the third register.Join the waitlist — get patent alerts
Track US2021357215A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.