Processor for performing multiply-add operations on packed data
Abstract
A method and apparatus for including in a processor instructions for performing multiply-subtract operations on packed data. In one embodiment, a processor is coupled to a memory. The memory has stored therein a first packed data and a second packed data. The processor performs operations on data elements in said first packed data and said second packed data to generate a third packed data in response to receiving an instruction. At least two of the data elements in this third packed data storing the result of performing multiply-subtract operations on data elements in the first and second packed data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor comprising:
a memory to store one or more instructions including a multiply-subtract instruction, one or more execution units, and a register file including a plurality of registers to store packed data, including a first data element (A1), a second data element (A2), a third data element (A3), and a fourth data element (A4) in a first register and a fifth data element (B1), a sixth data element (B2), a seventh data element (B3), and an eighth (B4) data element in a second register; wherein the one or more execution units, in response to performing the multiply-subtract instruction, are to generate a first intermediate result equal to the first data element multiplied by the fifth data element (IR1=A1×B1), generate a second intermediate result equal to the second data element multiplied by the sixth data element (IR2=A2×B2), generate a third intermediate result equal to the third data element multiplied by the seventh data element (IR3=A3×B3), generate a fourth intermediate result equal to the fourth data element multiplied by the eighth data element (IR4=A4×B4), generate a first result equal to the first intermediate result minus the second intermediate result (R1=IR1−IR2=[(A1×B1)−(A2×B2)]) and a second result equal to the third intermediate result minus the fourth intermediate result (R2=IR3−IR4=[(A3×B3)−(A4×B4)]).
2 . The processor of claim 1 , wherein in response to performing the multiply-subtract instruction a result packed data is stored into said register file without arithmetically combining said first and second results.
3 . The processor of claim 2 , wherein storing said result packed data includes storing said result packed data as an operand for use by another instruction.
4 . The processor of claim 1 , wherein said first and second results each contain more bits than each of said first, second, third, fourth, fifth, sixth, seventh and eighth data elements.
5 . The processor of claim 4 , wherein said first and second results each contain twice as many bits as each of said first, second, third, fourth, fifth, sixth, seventh and eighth data elements.
6 . A computer-implemented method responsive to the execution of a single instruction comprising the steps of:
performing the following steps in response to executing said single instruction, multiplying together a first value and a second value to generate a first intermediate result, multiplying together a third value and a fourth value to generate a second intermediate result, multiplying together a fifth value and a sixth value to generate a third intermediate result, multiplying together a seventh value and an eighth value to generate a fourth intermediate result, subtracting said first intermediate result and said second intermediate result to generate a first data element in a first packed data, subtracting said third intermediate result and said fourth intermediate result to generate a second data element in said first packed data; storing said first packed data in a first storage area.
7 . The method of claim 1 , wherein execution of said single instruction is completed without arithmetically combining said first and second data elements of said first packed data.
8 . The method of claim 1 , wherein said storing said first packed data includes storing said first packed data as an operand for use by another instruction.
9 . The method of claim 1 , wherein said first and second data elements of said first packed data each contain more bits than each of said first, second, third, fourth, fifth, sixth, seventh and eighth values.
10 . The method of claim 1 , wherein said first and second data elements of said first packed data each contain twice as many bits as each of said first, second, third, fourth, fifth, sixth, seventh and eighth values.Join the waitlist — get patent alerts
Track US2013262836A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.