SSE2 scalar and packed arithmetic

SSE2 scalar and packed arithmetic

x86-64 guarantees the SSE2 instruction set. In 64-bit mode it provides sixteen registers named xmm0 through xmm15, each 128 bits wide. These registers hold bits; an instruction decides whether those bits represent one floating-point value or several values side by side.

This lesson draws every XMM register from its low bits to its high bits, left to right:

             low bits                                      high bits
xmm0  [ lane 0: bits 0..63       | lane 1: bits 64..127        ]

Each 64-bit region above is a lane when an instruction treats the register as two binary64 values. Lane 0 is the low lane and lane 1 is the high lane.

Reading the instruction suffixes

In the instruction families used here, the final letters say how many values the instruction uses and which format it gives those bits:

suffixmeaninglanes used in one XMM register
ssscalar binary32one 32-bit value in the low bits
sdscalar binary64one 64-bit value in the low lane
pspacked binary32four 32-bit lanes
pdpacked binary64two 64-bit lanes

Scalar means one value. Packed means several independent lanes. For example, addsd adds the low binary64 values, while addpd adds both pairs of binary64 lanes.

The examples begin with binary64, so their arithmetic mnemonics end in sd or pd.

Scalar arithmetic and the high half

The two-operand scalar SSE2 instructions used here read the low lane of each operand, write their answer to the destination's low lane, and preserve the destination's high 64 bits. This first calculation starts with memory loads whose high halves become zero:

Here is the complete addsd change:

before: xmm0 = [ low 1.5  | high 0 ]
        xmm1 = [ low 2.25 | high 0 ]
after:  xmm0 = [ low 3.75 | high 0 ]

Only the low lanes take part in the addition. The high lane shown in the result is the old high lane of xmm0.

The source form of movsd determines what happens to that high half:

  • movsd xmm0, [value] loads eight bytes into the low lane and clears the high 64 bits.
  • movsd xmm0, xmm1 copies the low lane and preserves the high 64 bits already in xmm0.

This example gives the preserved half a visible value:

The register-to-register move changes xmm0 like this:

before: xmm0 = [ low 1.0 | high 99.0 ]
        xmm1 = [ low 4.0 | high 0    ]
after:  xmm0 = [ low 4.0 | high 99.0 ]

subsd, mulsd, divsd, and sqrtsd follow the same destination rule: they replace the low binary64 lane and preserve the destination's high 64 bits.

Packed loads and arithmetic

Packed binary64 instructions use both 64-bit lanes. With dq 1.0, 2.0, the first qword is at the lower address and becomes low lane 0. The next qword becomes high lane 1:

increasing address  ---->
memory left:  [ qword 1.0 | qword 2.0 ]
xmm0:         [ low 1.0   | high 2.0  ]

movupd loads or stores all 128 bits and has no alignment requirement. addpd then adds corresponding lanes independently:

lanexmm0 beforexmm1xmm0 after
low lane 01.010.011.0
high lane 12.020.022.0

mulpd has the same lane-by-lane shape, with multiplication in place of addition. Neither packed operation carries a value from one lane into the other.

Converting signed integers and binary64 values

Conversion instructions change the representation of a value:

instructionoperation
cvtsi2sd xmm0, raxconvert the signed qword in rax to binary64 in the low lane
cvttsd2si rax, xmm0convert the low binary64 value to a signed qword, truncating toward zero
cvtsd2si rax, xmm0convert the low binary64 value to a signed qword using MXCSR rounding control
movq rax, xmm0copy only the low 64 raw bits, with no numerical conversion

The usual MXCSR rounding mode is nearest, with halfway cases going to the result whose low bit is even. In that mode, cvtsd2si rounds 2.7 to 3. cvttsd2si always truncates toward zero, so it converts 2.7 to 2 regardless of that rounding setting.

Like scalar arithmetic, cvtsi2sd xmm, r64 replaces the low lane and preserves the destination's high 64 bits:

The conversion of 7 changes xmm0 as follows:

before: xmm0 = [ low 1.0 | high 99.0 ]
after:  xmm0 = [ low 7.0 | high 99.0 ]

Decimal 2.7 has an infinite repeating expansion in base two. The assembler stores the nearest binary64 encoding, 0x400599999999999A, under its normal rounding rule. Its exponent makes the normalized value approximately 1.35 x 2^1; the stored fraction bits describe the digits after the leading 1 in that significand. movq lets r10 expose those low-lane bits without treating them as an integer value to convert.

MXCSR in brief

mxcsr holds control and status for SSE floating-point work, including its rounding-control field, exception masks, and recorded exception conditions. The playground begins in the usual default state: nearest-even rounding with the floating-point exceptions masked.

A masked exception records its condition and lets the instruction produce its defined floating-point result. For example, a nonzero finite value divided by zero produces a signed infinity when divide-by-zero is masked. 0.0 / 0.0 is an invalid operation and produces a NaN when invalid operation is masked. The result for a masked exceptional operation depends on the operation and condition; masking does not turn every invalid calculation into the same result.

Comparing low lanes

ucomisd xmm0, xmm1 compares only the low binary64 lanes. It ignores both high lanes and records one of four outcomes in ZF, PF, and CF:

relation of low xmm0 to low xmm1ZFPFCF
greater000
less001
equal100
unordered111

Unordered means at least one operand is a NaN. Test it first with jp. Its row also satisfies the flag conditions used by jb and je, so either of those jumps would misclassify a NaN if it ran first.

This input leaves 2 in r12.

Packed comparisons produce masks

A packed comparison writes a result into every lane. cmppd takes an immediate predicate; the value 1 selects ordered signaling less-than: “left lane is less than right lane.” If that relation is true, the result lane is all one bits. If it is false, the result lane is all zero bits. A quiet NaN makes the lane false and all-zero, and it raises the invalid-operation condition. With the masked default taught above, MXCSR records that condition and execution continues.

lanecomparisonresult bits
low lane 01.0 < 2.0, true0xFFFFFFFFFFFFFFFF
high lane 15.0 < 4.0, false0x0000000000000000

Those all-one and all-zero lanes form a mask. They are bit patterns for selecting later data, so the floating-point display of the result is not useful here.

Floating-point arguments and results

For ordinary non-variadic System V x86-64 calls with simple scalar floating-point arguments, the first floating arguments use xmm0 through xmm7; integer arguments are counted separately in the general-purpose argument registers. A floating-point result is returned in xmm0.

Every XMM register is caller-saved. A called function may change any of them. A caller that needs an XMM value afterward must preserve it in memory, copy its bits elsewhere, or recompute it.

The playground gives _start a 16-byte-aligned rsp, and this code does not move rsp before the call. call pushes the return address, so rsp is eight bytes away from a 16-byte boundary on entry to sum_of_squares, as required by the convention. The function makes no nested calls. xorpd xmm1, xmm1 clears every bit in xmm1, so both of its 64-bit lanes become zero.

Your turn: scalar sum and average

Use all four binary64 values at values. Calculate their sum in xmm1 and their average in xmm0. The exact sum is 18.0 and the exact average is 4.5; the stored divisor is 4.0, so loading the divisor alone cannot produce either requested answer.

The fixed code stores both results and copies the low 64 raw bits into r12 and r13. The checks assert both memory values and both safe register copies.

Show solution

Now perform packed addition and multiplication. Load the two lanes at left and right, put their lane-wise sums in xmm0, and put their lane-wise products in xmm2. The fixed code stores all four answers and copies each low/high qword to a safe general-purpose register.

Show solution