x86-64 extensions

Instructions without a page

Everything else the assembler accepts, under the heading NASM files it under. These have no page of their own: the only description of them we are free to publish is the one line below, and blink implements a part of them.

AMD Enhanced 3DNow! (Athlon) instructions

5 instructions
  • pf2iw — Packed Floating-Point to Integer Word Conversion
  • pfnacc — Packed Floating-Point Negative Accumulate
  • pfpnacc — Packed Floating-Point Positive-Negative Accumulate
  • pi2fw — Packed Integer to Floating-Point Word Conversion
  • pswapd — Packed Swap Doubleword

AMD Lightweight Profiling (LWP) instructions

4 instructions
  • llwpcb
  • lwpins
  • lwpval
  • slwpcb

AMD SSE4A

4 instructions
  • extrq — Extract Field
  • insertq — Insert Field
  • movntsd — Store Scalar Double-Precision Floating-Point Values Using Non-Temporal Hint
  • movntss — Store Scalar Single-Precision Floating-Point Values Using Non-Temporal Hint

AMD XOP and FMA4 instructions (SSE5)

73 instructions
  • vfmaddpd — Fused Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfmaddps — Fused Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfmaddsd — Fused Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfmaddss — Fused Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfmaddsubpd — Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Values
  • vfmaddsubps — Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Values
  • vfmsubaddpd — Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Values
  • vfmsubaddps — Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Values
  • vfmsubpd — Fused Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfmsubps — Fused Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfmsubsd — Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfmsubss — Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfnmaddpd — Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfnmaddps — Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfnmaddsd — Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfnmaddss — Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfnmsubpd — Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfnmsubps — Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfnmsubsd — Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfnmsubss — Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfrczpd — Extract Fraction Packed Double-Precision Floating-Point
  • vfrczps — Extract Fraction Packed Single-Precision Floating-Point
  • vfrczsd — Extract Fraction Scalar Double-Precision Floating-Point
  • vfrczss — Extract Fraction Scalar Single-Precision Floating Point
  • vpcmov — Packed Conditional Move
  • vpcomb — Compare Packed Signed Byte Integers
  • vpcomd — Compare Packed Signed Doubleword Integers
  • vpcomq — Compare Packed Signed Quadword Integers
  • vpcomub — Compare Packed Unsigned Byte Integers
  • vpcomud — Compare Packed Unsigned Doubleword Integers
  • vpcomuq — Compare Packed Unsigned Quadword Integers
  • vpcomuw — Compare Packed Unsigned Word Integers
  • vpcomw — Compare Packed Signed Word Integers
  • vphaddbd — Packed Horizontal Add Signed Byte to Signed Doubleword
  • vphaddbq — Packed Horizontal Add Signed Byte to Signed Quadword
  • vphaddbw — Packed Horizontal Add Signed Byte to Signed Word
  • vphadddq — Packed Horizontal Add Signed Doubleword to Signed Quadword
  • vphaddubd — Packed Horizontal Add Unsigned Byte to Doubleword
  • vphaddubq — Packed Horizontal Add Unsigned Byte to Quadword
  • vphaddubw — Packed Horizontal Add Unsigned Byte to Word
  • vphaddudq — Packed Horizontal Add Unsigned Doubleword to Quadword
  • vphadduwd — Packed Horizontal Add Unsigned Word to Doubleword
  • vphadduwq — Packed Horizontal Add Unsigned Word to Quadword
  • vphaddwd — Packed Horizontal Add Signed Word to Signed Doubleword
  • vphaddwq — Packed Horizontal Add Signed Word to Signed Quadword
  • vphsubbw — Packed Horizontal Subtract Signed Byte to Signed Word
  • vphsubdq — Packed Horizontal Subtract Signed Doubleword to Signed Quadword
  • vphsubwd — Packed Horizontal Subtract Signed Word to Signed Doubleword
  • vpmacsdd — Packed Multiply Accumulate Signed Doubleword to Signed Doubleword
  • vpmacsdqh — Packed Multiply Accumulate Signed High Doubleword to Signed Quadword
  • vpmacsdql — Packed Multiply Accumulate Signed Low Doubleword to Signed Quadword
  • vpmacssdd — Packed Multiply Accumulate with Saturation Signed Doubleword to Signed Doubleword
  • vpmacssdqh — Packed Multiply Accumulate with Saturation Signed High Doubleword to Signed Quadword
  • vpmacssdql — Packed Multiply Accumulate with Saturation Signed Low Doubleword to Signed Quadword
  • vpmacsswd — Packed Multiply Accumulate with Saturation Signed Word to Signed Doubleword
  • vpmacssww — Packed Multiply Accumulate with Saturation Signed Word to Signed Word
  • vpmacswd — Packed Multiply Accumulate Signed Word to Signed Doubleword
  • vpmacsww — Packed Multiply Accumulate Signed Word to Signed Word
  • vpmadcsswd — Packed Multiply Add Accumulate with Saturation Signed Word to Signed Doubleword
  • vpmadcswd — Packed Multiply Add Accumulate Signed Word to Signed Doubleword
  • vpperm — Packed Permute Bytes
  • vprotb — Packed Rotate Bytes
  • vprotd — Packed Rotate Doublewords
  • vprotq — Packed Rotate Quadwords
  • vprotw — Packed Rotate Words
  • vpshab — Packed Shift Arithmetic Bytes
  • vpshad — Packed Shift Arithmetic Doublewords
  • vpshaq — Packed Shift Arithmetic Quadwords
  • vpshaw — Packed Shift Arithmetic Words
  • vpshlb — Packed Shift Logical Bytes
  • vpshld — Packed Shift Logical Doublewords
  • vpshlq — Packed Shift Logical Quadwords
  • vpshlw — Packed Shift Logical Words

AMD XOP bit operations

9 instructions
  • blcfill — Fill From Lowest Clear Bit
  • blci — Isolate Lowest Clear Bit
  • blcic — Isolate Lowest Set Bit and Complement
  • blcmsk — Mask From Lowest Clear Bit
  • blcs — Set Lowest Clear Bit
  • blsfill — Fill From Lowest Set Bit
  • blsic — Isolate Lowest Set Bit and Complement
  • t1mskc — Inverse Mask From Trailing Ones
  • tzmsk — Mask From Trailing Zeros

AVX no exception conversions

7 instructions
  • vbcstnebf162ps — Load BF16 Element and Convert to FP32 Element With Broadcas
  • vbcstnebf16ps
  • vbcstnesh2ps — Load FP16 Element and Convert to FP32 Element with Broadcast
  • vcvtneebf162ps — Convert Even Elements of Packed BF16 Values to FP32 Values
  • vcvtneeph2ps — Convert Even Elements of Packed FP16 Values to FP32 Values
  • vcvtneobf162ps — Convert Odd Elements of Packed BF16 Values to FP32 Values
  • vcvtneoph2ps — Convert Odd Elements of Packed FP16 Values to FP32 Values

AVX-512 instructions

605 instructions
  • vaddpd — Add Packed Double-Precision Floating-Point Values
  • vaddps — Add Packed Single-Precision Floating-Point Values
  • valignd — Align Doubleword Vectors
  • valignq — Align Quadword Vectors
  • vandnpd — Bitwise Logical AND NOT of Packed Double-Precision Floating-Point Values
  • vandnps — Bitwise Logical AND NOT of Packed Single-Precision Floating-Point Values
  • vandpd — Bitwise Logical AND of Packed Double-Precision Floating-Point Values
  • vandps — Bitwise Logical AND of Packed Single-Precision Floating-Point Values
  • vblendmpd — Blend Packed Double-Precision Floating-Point Vectors Using an OpMask Control
  • vblendmps — Blend Packed Single-Precision Floating-Point Vectors Using an OpMask Control
  • vbroadcastf32x2 — Broadcast Two Single-Precision Floating-Point Elements
  • vbroadcastf32x4 — Broadcast Four Single-Precision Floating-Point Elements
  • vbroadcastf32x8 — Broadcast Eight Single-Precision Floating-Point Elements
  • vbroadcastf64x2 — Broadcast Two Double-Precision Floating-Point Elements
  • vbroadcastf64x4 — Broadcast Four Double-Precision Floating-Point Elements
  • vbroadcasti32x2 — Broadcast Two Doubleword Elements
  • vbroadcasti32x4 — Broadcast Four Doubleword Elements
  • vbroadcasti32x8 — Broadcast Eight Doubleword Elements
  • vbroadcasti64x2 — Broadcast Two Quadword Elements
  • vbroadcasti64x4 — Broadcast Four Quadword Elements
  • vbroadcastsd — Broadcast Double-Precision Floating-Point Element
  • vbroadcastss — Broadcast Single-Precision Floating-Point Element
  • vcmpeq_oqpd
  • vcmpeq_oqps
  • vcmpeq_oqsd
  • vcmpeq_oqss
  • vcmpeq_uqpd
  • vcmpeq_uqps
  • vcmpeq_uspd
  • vcmpeq_usps
  • vcmpeqpd
  • vcmpeqps
  • vcmpfalse_oqpd
  • vcmpfalse_oqps
  • vcmpfalse_ospd
  • vcmpfalse_osps
  • vcmpfalsepd
  • vcmpfalseps
  • vcmpge_oqpd
  • vcmpge_oqps
  • vcmpge_ospd
  • vcmpge_osps
  • vcmpgepd
  • vcmpgeps
  • vcmpgt_oqpd
  • vcmpgt_oqps
  • vcmpgt_ospd
  • vcmpgt_osps
  • vcmpgtpd
  • vcmpgtps
  • vcmple_oqpd
  • vcmple_oqps
  • vcmple_ospd
  • vcmple_osps
  • vcmplepd
  • vcmpleps
  • vcmplt_oqpd
  • vcmplt_oqps
  • vcmplt_ospd
  • vcmplt_osps
  • vcmpltpd
  • vcmpltps
  • vcmpneq_oqpd
  • vcmpneq_oqps
  • vcmpneq_ospd
  • vcmpneq_osps
  • vcmpneq_uqpd
  • vcmpneq_uqps
  • vcmpneq_uspd
  • vcmpneq_usps
  • vcmpneqpd
  • vcmpneqps
  • vcmpnge_uqpd
  • vcmpnge_uqps
  • vcmpnge_uspd
  • vcmpnge_usps
  • vcmpngepd
  • vcmpngeps
  • vcmpngt_uqpd
  • vcmpngt_uqps
  • vcmpngt_uspd
  • vcmpngt_usps
  • vcmpngtpd
  • vcmpngtps
  • vcmpnle_uqpd
  • vcmpnle_uqps
  • vcmpnle_uspd
  • vcmpnle_usps
  • vcmpnlepd
  • vcmpnleps
  • vcmpnlt_uqpd
  • vcmpnlt_uqps
  • vcmpnlt_uspd
  • vcmpnlt_usps
  • vcmpnltpd
  • vcmpnltps
  • vcmpord_qpd
  • vcmpord_qps
  • vcmpord_spd
  • vcmpord_sps
  • vcmpordpd
  • vcmpordps
  • vcmppd — Compare Packed Double-Precision Floating-Point Values
  • vcmpps — Compare Packed Single-Precision Floating-Point Values
  • vcmptrue_uqpd
  • vcmptrue_uqps
  • vcmptrue_uspd
  • vcmptrue_usps
  • vcmptruepd
  • vcmptrueps
  • vcmpunord_qpd
  • vcmpunord_qps
  • vcmpunord_spd
  • vcmpunord_sps
  • vcmpunordpd
  • vcmpunordps
  • vcompresspd — Store Sparse Packed Double-Precision Floating-Point Values into Dense Memory/Register
  • vcompressps — Store Sparse Packed Single-Precision Floating-Point Values into Dense Memory/Register
  • vcvtdq2pd — Convert Packed Dword Integers to Packed Double-Precision FP Values
  • vcvtdq2ps — Convert Packed Dword Integers to Packed Single-Precision FP Values
  • vcvtpd2qq — Convert Packed Double-Precision Floating-Point Values to Packed Quadword Integers
  • vcvtpd2udq — Convert Packed Double-Precision Floating-Point Values to Packed Unsigned Doubleword Integers
  • vcvtpd2uqq — Convert Packed Double-Precision Floating-Point Values to Packed Unsigned Quadword Integers
  • vcvtps2dq — Convert Packed Single-Precision FP Values to Packed Dword Integers
  • vcvtps2pd — Convert Packed Single-Precision FP Values to Packed Double-Precision FP Values
  • vcvtps2qq — Convert Packed Single Precision Floating-Point Values to Packed Singed Quadword Integer Values
  • vcvtps2udq — Convert Packed Single-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Values
  • vcvtps2uqq — Convert Packed Single Precision Floating-Point Values to Packed Unsigned Quadword Integer Values
  • vcvtqq2pd — Convert Packed Quadword Integers to Packed Double-Precision Floating-Point Values
  • vcvtqq2ps — Convert Packed Quadword Integers to Packed Single-Precision Floating-Point Values
  • vcvtsd2usi — Convert Scalar Double-Precision Floating-Point Value to Unsigned Doubleword Integer
  • vcvtss2usi — Convert Scalar Single-Precision Floating-Point Value to Unsigned Doubleword Integer
  • vcvttpd2qq — Convert with Truncation Packed Double-Precision Floating-Point Values to Packed Quadword Integers
  • vcvttpd2udq — Convert with Truncation Packed Double-Precision Floating-Point Values to Packed Unsigned Doubleword Integers
  • vcvttpd2uqq — Convert with Truncation Packed Double-Precision Floating-Point Values to Packed Unsigned Quadword Integers
  • vcvttps2dq — Convert with Truncation Packed Single-Precision FP Values to Packed Dword Integers
  • vcvttps2qq — Convert with Truncation Packed Single Precision Floating-Point Values to Packed Singed Quadword Integer Values
  • vcvttps2udq — Convert with Truncation Packed Single-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Values
  • vcvttps2uqq — Convert with Truncation Packed Single Precision Floating-Point Values to Packed Unsigned Quadword Integer Values
  • vcvttsd2usi — Convert with Truncation Scalar Double-Precision Floating-Point Value to Unsigned Integer
  • vcvttss2usi — Convert with Truncation Scalar Single-Precision Floating-Point Value to Unsigned Integer
  • vcvtudq2pd — Convert Packed Unsigned Doubleword Integers to Packed Double-Precision Floating-Point Values
  • vcvtudq2ps — Convert Packed Unsigned Doubleword Integers to Packed Single-Precision Floating-Point Values
  • vcvtuqq2pd — Convert Packed Unsigned Quadword Integers to Packed Double-Precision Floating-Point Values
  • vcvtuqq2ps — Convert Packed Unsigned Quadword Integers to Packed Single-Precision Floating-Point Values
  • vcvtusi2sd — Convert Unsigned Integer to Scalar Double-Precision Floating-Point Value
  • vcvtusi2ss — Convert Unsigned Integer to Scalar Single-Precision Floating-Point Value
  • vdbpsadbw — Double Block Packed Sum-Absolute-Differences on Unsigned Bytes
  • vdivpd — Divide Packed Double-Precision Floating-Point Values
  • vdivps — Divide Packed Single-Precision Floating-Point Values
  • vexp2pd — Approximation to the Exponential 2^x of Packed Double-Precision Floating-Point Values with Less Than 2^-23 Relative Error
  • vexp2ps — Approximation to the Exponential 2^x of Packed Single-Precision Floating-Point Values with Less Than 2^-23 Relative Error
  • vexpandpd — Load Sparse Packed Double-Precision Floating-Point Values from Dense Memory
  • vexpandps — Load Sparse Packed Single-Precision Floating-Point Values from Dense Memory
  • vextractf32x4 — Extract 128 Bits of Packed Single-Precision Floating-Point Values
  • vextractf32x8 — Extract 256 Bits of Packed Single-Precision Floating-Point Values
  • vextractf64x2 — Extract 128 Bits of Packed Double-Precision Floating-Point Values
  • vextractf64x4 — Extract 256 Bits of Packed Double-Precision Floating-Point Values
  • vextracti32x4 — Extract 128 Bits of Packed Doubleword Integer Values
  • vextracti32x8 — Extract 256 Bits of Packed Doubleword Integer Values
  • vextracti64x2 — Extract 128 Bits of Packed Quadword Integer Values
  • vextracti64x4 — Extract 256 Bits of Packed Quadword Integer Values
  • vextractps — Extract Packed Single Precision Floating-Point Value
  • vfixupimmpd — Fix Up Special Packed Double-Precision Floating-Point Values
  • vfixupimmps — Fix Up Special Packed Single-Precision Floating-Point Values
  • vfixupimmsd — Fix Up Special Scalar Double-Precision Floating-Point Value
  • vfixupimmss — Fix Up Special Scalar Single-Precision Floating-Point Value
  • vfmadd132pd — Fused Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfmadd132ps — Fused Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfmadd213pd — Fused Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfmadd213ps — Fused Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfmadd231pd — Fused Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfmadd231ps — Fused Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfmaddsub132pd — Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Values
  • vfmaddsub132ps — Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Values
  • vfmaddsub213pd — Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Values
  • vfmaddsub213ps — Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Values
  • vfmaddsub231pd — Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Values
  • vfmaddsub231ps — Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Values
  • vfmsub132pd — Fused Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfmsub132ps — Fused Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfmsub213pd — Fused Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfmsub213ps — Fused Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfmsub231pd — Fused Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfmsub231ps — Fused Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfmsubadd132pd — Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Values
  • vfmsubadd132ps — Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Values
  • vfmsubadd213pd — Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Values
  • vfmsubadd213ps — Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Values
  • vfmsubadd231pd — Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Values
  • vfmsubadd231ps — Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Values
  • vfnmadd132pd — Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfnmadd132ps — Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfnmadd213pd — Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfnmadd213ps — Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfnmadd231pd — Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Values
  • vfnmadd231ps — Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Values
  • vfnmsub132pd — Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfnmsub132ps — Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfnmsub213pd — Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfnmsub213ps — Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfnmsub231pd — Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Values
  • vfnmsub231ps — Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Values
  • vfpclasspd — Test Class of Packed Double-Precision Floating-Point Values
  • vfpclassps — Test Class of Packed Single-Precision Floating-Point Values
  • vfpclasssd — Test Class of Scalar Double-Precision Floating-Point Value
  • vfpclassss — Test Class of Scalar Single-Precision Floating-Point Value
  • vgatherdpd — Gather Packed Double-Precision Floating-Point Values Using Signed Doubleword Indices
  • vgatherdps — Gather Packed Single-Precision Floating-Point Values Using Signed Doubleword Indices
  • vgatherpf0dpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hint
  • vgatherpf0dps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hint
  • vgatherpf0qpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hint
  • vgatherpf0qps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hint
  • vgatherpf1dpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hint
  • vgatherpf1dps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hint
  • vgatherpf1qpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hint
  • vgatherpf1qps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hint
  • vgatherqpd — Gather Packed Double-Precision Floating-Point Values Using Signed Quadword Indices
  • vgatherqps — Gather Packed Single-Precision Floating-Point Values Using Signed Quadword Indices
  • vgetexppd — Extract Exponents of Packed Double-Precision Floating-Point Values as Double-Precision Floating-Point Values
  • vgetexpps — Extract Exponents of Packed Single-Precision Floating-Point Values as Single-Precision Floating-Point Values
  • vgetexpsd — Extract Exponent of Scalar Double-Precision Floating-Point Value as Double-Precision Floating-Point Value
  • vgetexpss — Extract Exponent of Scalar Single-Precision Floating-Point Value as Single-Precision Floating-Point Value
  • vgetmantpd — Extract Normalized Mantissas from Packed Double-Precision Floating-Point Values
  • vgetmantps — Extract Normalized Mantissas from Packed Single-Precision Floating-Point Values
  • vgetmantsd — Extract Normalized Mantissa from Scalar Double-Precision Floating-Point Value
  • vgetmantss — Extract Normalized Mantissa from Scalar Single-Precision Floating-Point Value
  • vinsertf32x4 — Insert 128 Bits of Packed Single-Precision Floating-Point Values
  • vinsertf32x8 — Insert 256 Bits of Packed Single-Precision Floating-Point Values
  • vinsertf64x2 — Insert 128 Bits of Packed Double-Precision Floating-Point Values
  • vinsertf64x4 — Insert 256 Bits of Packed Double-Precision Floating-Point Values
  • vinserti32x4 — Insert 128 Bits of Packed Doubleword Integer Values
  • vinserti32x8 — Insert 256 Bits of Packed Doubleword Integer Values
  • vinserti64x2 — Insert 128 Bits of Packed Quadword Integer Values
  • vinserti64x4 — Insert 256 Bits of Packed Quadword Integer Values
  • vmaxpd — Return Maximum Packed Double-Precision Floating-Point Values
  • vmaxph — Return Maximum Packed Half-Precision Floating-Point Values
  • vmaxps — Return Maximum Packed Single-Precision Floating-Point Values
  • vminpd — Return Minimum Packed Double-Precision Floating-Point Values
  • vminph — Return Minimum Packed Half-Precision Floating-Point Values
  • vminps — Return Minimum Packed Single-Precision Floating-Point Values
  • vmovapd — Move Aligned Packed Double-Precision Floating-Point Values
  • vmovaps — Move Aligned Packed Single-Precision Floating-Point Values
  • vmovddup — Move One Double-FP and Duplicate
  • vmovdqa32 — Move Aligned Doubleword Values
  • vmovdqa64 — Move Aligned Quadword Values
  • vmovdqu16 — Move Unaligned Word Values
  • vmovdqu32 — Move Unaligned Doubleword Values
  • vmovdqu64 — Move Unaligned Quadword Values
  • vmovdqu8 — Move Unaligned Byte Values
  • vmovntdq — Store Double Quadword Using Non-Temporal Hint
  • vmovntdqa — Load Double Quadword Non-Temporal Aligned Hint
  • vmovntpd — Store Packed Double-Precision Floating-Point Values Using Non-Temporal Hint
  • vmovntps — Store Packed Single-Precision Floating-Point Values Using Non-Temporal Hint
  • vmovshdup — Move Packed Single-FP High and Duplicate
  • vmovsldup — Move Packed Single-FP Low and Duplicate
  • vmovupd — Move Unaligned Packed Double-Precision Floating-Point Values
  • vmovups — Move Unaligned Packed Single-Precision Floating-Point Values
  • vmulpd — Multiply Packed Double-Precision Floating-Point Values
  • vmulps — Multiply Packed Single-Precision Floating-Point Values
  • vorpd — Bitwise Logical OR of Double-Precision Floating-Point Values
  • vorps — Bitwise Logical OR of Single-Precision Floating-Point Values
  • vpabsb — Packed Absolute Value of Byte Integers
  • vpabsd — Packed Absolute Value of Doubleword Integers
  • vpabsq — Packed Absolute Value of Quadword Integers
  • vpabsw — Packed Absolute Value of Word Integers
  • vpackssdw — Pack Doublewords into Words with Signed Saturation
  • vpacksswb — Pack Words into Bytes with Signed Saturation
  • vpackusdw — Pack Doublewords into Words with Unsigned Saturation
  • vpackuswb — Pack Words into Bytes with Unsigned Saturation
  • vpaddb — Add Packed Byte Integers
  • vpaddd — Add Packed Doubleword Integers
  • vpaddq — Add Packed Quadword Integers
  • vpaddsb — Add Packed Signed Byte Integers with Signed Saturation
  • vpaddsw — Add Packed Signed Word Integers with Signed Saturation
  • vpaddusb — Add Packed Unsigned Byte Integers with Unsigned Saturation
  • vpaddusw — Add Packed Unsigned Word Integers with Unsigned Saturation
  • vpaddw — Add Packed Word Integers
  • vpalignr — Packed Align Right
  • vpandd — Bitwise Logical AND of Packed Doubleword Integers
  • vpandnd — Bitwise Logical AND NOT of Packed Doubleword Integers
  • vpandnq — Bitwise Logical AND NOT of Packed Quadword Integers
  • vpandq — Bitwise Logical AND of Packed Quadword Integers
  • vpavgb — Average Packed Byte Integers
  • vpavgw — Average Packed Word Integers
  • vpblendmb — Blend Byte Vectors Using an OpMask Control
  • vpblendmd — Blend Doubleword Vectors Using an OpMask Control
  • vpblendmq — Blend Quadword Vectors Using an OpMask Control
  • vpblendmw — Blend Word Vectors Using an OpMask Control
  • vpbroadcastb — Broadcast Byte Integer
  • vpbroadcastd — Broadcast Doubleword Integer
  • vpbroadcastmb2q — Broadcast Low Byte of Mask Register to Packed Quadword Values
  • vpbroadcastmw2d — Broadcast Low Word of Mask Register to Packed Doubleword Values
  • vpbroadcastq — Broadcast Quadword Integer
  • vpbroadcastw — Broadcast Word Integer
  • vpcmpb — Compare Packed Signed Byte Values
  • vpcmpd — Compare Packed Signed Doubleword Values
  • vpcmpeqb — Compare Packed Byte Data for Equality
  • vpcmpeqd — Compare Packed Doubleword Data for Equality
  • vpcmpeqq — Compare Packed Quadword Data for Equality
  • vpcmpequb
  • vpcmpequd
  • vpcmpequq
  • vpcmpequw
  • vpcmpeqw — Compare Packed Word Data for Equality
  • vpcmpgeb
  • vpcmpged
  • vpcmpgeq
  • vpcmpgeub
  • vpcmpgeud
  • vpcmpgeuq
  • vpcmpgeuw
  • vpcmpgew
  • vpcmpgtb — Compare Packed Signed Byte Integers for Greater Than
  • vpcmpgtd — Compare Packed Signed Doubleword Integers for Greater Than
  • vpcmpgtq — Compare Packed Data for Greater Than
  • vpcmpgtub
  • vpcmpgtud
  • vpcmpgtuq
  • vpcmpgtuw
  • vpcmpgtw — Compare Packed Signed Word Integers for Greater Than
  • vpcmpleb
  • vpcmpled
  • vpcmpleq
  • vpcmpleub
  • vpcmpleud
  • vpcmpleuq
  • vpcmpleuw
  • vpcmplew
  • vpcmpltb
  • vpcmpltd
  • vpcmpltq
  • vpcmpltub
  • vpcmpltud
  • vpcmpltuq
  • vpcmpltuw
  • vpcmpltw
  • vpcmpneqb
  • vpcmpneqd
  • vpcmpneqq
  • vpcmpnequb
  • vpcmpnequd
  • vpcmpnequq
  • vpcmpnequw
  • vpcmpneqw
  • vpcmpngtb
  • vpcmpngtd
  • vpcmpngtq
  • vpcmpngtub
  • vpcmpngtud
  • vpcmpngtuq
  • vpcmpngtuw
  • vpcmpngtw
  • vpcmpnleb
  • vpcmpnled
  • vpcmpnleq
  • vpcmpnleub
  • vpcmpnleud
  • vpcmpnleuq
  • vpcmpnleuw
  • vpcmpnlew
  • vpcmpnltb
  • vpcmpnltd
  • vpcmpnltq
  • vpcmpnltub
  • vpcmpnltud
  • vpcmpnltuq
  • vpcmpnltuw
  • vpcmpnltw
  • vpcmpq — Compare Packed Signed Quadword Values
  • vpcmpub — Compare Packed Unsigned Byte Values
  • vpcmpud — Compare Packed Unsigned Doubleword Values
  • vpcmpuq — Compare Packed Unsigned Quadword Values
  • vpcmpuw — Compare Packed Unsigned Word Values
  • vpcmpw — Compare Packed Signed Word Values
  • vpcompressd — Store Sparse Packed Doubleword Integer Values into Dense Memory/Register
  • vpcompressq — Store Sparse Packed Quadword Integer Values into Dense Memory/Register
  • vpconflictd — Detect Conflicts Within a Vector of Packed Doubleword Values into Dense Memory/Register
  • vpconflictq — Detect Conflicts Within a Vector of Packed Quadword Values into Dense Memory/Register
  • vpermb — Permute Byte Integers
  • vpermd — Permute Doubleword Integers
  • vpermi2b — Full Permute of Bytes From Two Tables Overwriting the Index
  • vpermi2d — Full Permute of Doublewords From Two Tables Overwriting the Index
  • vpermi2pd — Full Permute of Double-Precision Floating-Point Values From Two Tables Overwriting the Index
  • vpermi2ps — Full Permute of Single-Precision Floating-Point Values From Two Tables Overwriting the Index
  • vpermi2q — Full Permute of Quadwords From Two Tables Overwriting the Index
  • vpermi2w — Full Permute of Words From Two Tables Overwriting the Index
  • vpermilpd — Permute Double-Precision Floating-Point Values
  • vpermilps — Permute Single-Precision Floating-Point Values
  • vpermpd — Permute Double-Precision Floating-Point Elements
  • vpermps — Permute Single-Precision Floating-Point Elements
  • vpermq — Permute Quadword Integers
  • vpermt2b — Full Permute of Bytes From Two Tables Overwriting a Table
  • vpermt2d — Full Permute of Doublewords From Two Tables Overwriting a Table
  • vpermt2pd — Full Permute of Double-Precision Floating-Point Values From Two Tables Overwriting a Table
  • vpermt2ps — Full Permute of Single-Precision Floating-Point Values From Two Tables Overwriting a Table
  • vpermt2q — Full Permute of Quadwords From Two Tables Overwriting a Table
  • vpermt2w — Full Permute of Words From Two Tables Overwriting a Table
  • vpermw — Permute Word Integers
  • vpexpandd — Load Sparse Packed Doubleword Integer Values from Dense Memory/Register
  • vpexpandq — Load Sparse Packed Quadword Integer Values from Dense Memory/Register
  • vpextrb — Extract Byte
  • vpextrw — Extract Word
  • vpgatherdd — Gather Packed Doubleword Values Using Signed Doubleword Indices
  • vpgatherdq — Gather Packed Quadword Values Using Signed Doubleword Indices
  • vpgatherqd — Gather Packed Doubleword Values Using Signed Quadword Indices
  • vpgatherqq — Gather Packed Quadword Values Using Signed Quadword Indices
  • vplzcntd — Count the Number of Leading Zero Bits for Packed Doubleword Values
  • vplzcntq — Count the Number of Leading Zero Bits for Packed Quadword Values
  • vpmadd52huq — Packed Multiply of Unsigned 52-bit Unsigned Integers and Add High 52-bit Products to Quadword Accumulators
  • vpmadd52luq — Packed Multiply of Unsigned 52-bit Integers and Add the Low 52-bit Products to Quadword Accumulators
  • vpmaddubsw — Multiply and Add Packed Signed and Unsigned Byte Integers
  • vpmaddwd — Multiply and Add Packed Signed Word Integers
  • vpmaxsb — Maximum of Packed Signed Byte Integers
  • vpmaxsd — Maximum of Packed Signed Doubleword Integers
  • vpmaxsq — Maximum of Packed Signed Quadword Integers
  • vpmaxsw — Maximum of Packed Signed Word Integers
  • vpmaxub — Maximum of Packed Unsigned Byte Integers
  • vpmaxud — Maximum of Packed Unsigned Doubleword Integers
  • vpmaxuq — Maximum of Packed Unsigned Quadword Integers
  • vpmaxuw — Maximum of Packed Unsigned Word Integers
  • vpminsb — Minimum of Packed Signed Byte Integers
  • vpminsd — Minimum of Packed Signed Doubleword Integers
  • vpminsq — Minimum of Packed Signed Quadword Integers
  • vpminsw — Minimum of Packed Signed Word Integers
  • vpminub — Minimum of Packed Unsigned Byte Integers
  • vpminud — Minimum of Packed Unsigned Doubleword Integers
  • vpminuq — Minimum of Packed Unsigned Quadword Integers
  • vpminuw — Minimum of Packed Unsigned Word Integers
  • vpmovb2m — Move Signs of Packed Byte Integers to Mask Register
  • vpmovd2m — Move Signs of Packed Doubleword Integers to Mask Register
  • vpmovdb — Down Convert Packed Doubleword Values to Byte Values with Truncation
  • vpmovdw — Down Convert Packed Doubleword Values to Word Values with Truncation
  • vpmovm2b — Expand Bits of Mask Register to Packed Byte Integers
  • vpmovm2d — Expand Bits of Mask Register to Packed Doubleword Integers
  • vpmovm2q — Expand Bits of Mask Register to Packed Quadword Integers
  • vpmovm2w — Expand Bits of Mask Register to Packed Word Integers
  • vpmovq2m — Move Signs of Packed Quadword Integers to Mask Register
  • vpmovqb — Down Convert Packed Quadword Values to Byte Values with Truncation
  • vpmovqd — Down Convert Packed Quadword Values to Doubleword Values with Truncation
  • vpmovqw — Down Convert Packed Quadword Values to Word Values with Truncation
  • vpmovsdb — Down Convert Packed Doubleword Values to Byte Values with Signed Saturation
  • vpmovsdw — Down Convert Packed Doubleword Values to Word Values with Signed Saturation
  • vpmovsqb — Down Convert Packed Quadword Values to Byte Values with Signed Saturation
  • vpmovsqd — Down Convert Packed Quadword Values to Doubleword Values with Signed Saturation
  • vpmovsqw — Down Convert Packed Quadword Values to Word Values with Signed Saturation
  • vpmovswb — Down Convert Packed Word Values to Byte Values with Signed Saturation
  • vpmovsxbd — Move Packed Byte Integers to Doubleword Integers with Sign Extension
  • vpmovsxbq — Move Packed Byte Integers to Quadword Integers with Sign Extension
  • vpmovsxbw — Move Packed Byte Integers to Word Integers with Sign Extension
  • vpmovsxdq — Move Packed Doubleword Integers to Quadword Integers with Sign Extension
  • vpmovsxwd — Move Packed Word Integers to Doubleword Integers with Sign Extension
  • vpmovsxwq — Move Packed Word Integers to Quadword Integers with Sign Extension
  • vpmovusdb — Down Convert Packed Doubleword Values to Byte Values with Unsigned Saturation
  • vpmovusdw — Down Convert Packed Doubleword Values to Word Values with Unsigned Saturation
  • vpmovusqb — Down Convert Packed Quadword Values to Byte Values with Unsigned Saturation
  • vpmovusqd — Down Convert Packed Quadword Values to Doubleword Values with Unsigned Saturation
  • vpmovusqw — Down Convert Packed Quadword Values to Word Values with Unsigned Saturation
  • vpmovuswb — Down Convert Packed Word Values to Byte Values with Unsigned Saturation
  • vpmovw2m — Move Signs of Packed Word Integers to Mask Register
  • vpmovwb — Down Convert Packed Word Values to Byte Values with Truncation
  • vpmovzxbd — Move Packed Byte Integers to Doubleword Integers with Zero Extension
  • vpmovzxbq — Move Packed Byte Integers to Quadword Integers with Zero Extension
  • vpmovzxbw — Move Packed Byte Integers to Word Integers with Zero Extension
  • vpmovzxdq — Move Packed Doubleword Integers to Quadword Integers with Zero Extension
  • vpmovzxwd — Move Packed Word Integers to Doubleword Integers with Zero Extension
  • vpmovzxwq — Move Packed Word Integers to Quadword Integers with Zero Extension
  • vpmuldq — Multiply Packed Signed Doubleword Integers and Store Quadword Result
  • vpmulhrsw — Packed Multiply Signed Word Integers and Store High Result with Round and Scale
  • vpmulhuw — Multiply Packed Unsigned Word Integers and Store High Result
  • vpmulhw — Multiply Packed Signed Word Integers and Store High Result
  • vpmulld — Multiply Packed Signed Doubleword Integers and Store Low Result
  • vpmullq — Multiply Packed Signed Quadword Integers and Store Low Result
  • vpmullw — Multiply Packed Signed Word Integers and Store Low Result
  • vpmultishiftqb — Select Packed Unaligned Bytes from Quadword Sources
  • vpmuludq — Multiply Packed Unsigned Doubleword Integers
  • vpord — Bitwise Logical OR of Packed Doubleword Integers
  • vporq — Bitwise Logical OR of Packed Quadword Integers
  • vprold — Rotate Packed Doubleword Left
  • vprolq — Rotate Packed Quadword Left
  • vprolvd — Variable Rotate Packed Doubleword Left
  • vprolvq — Variable Rotate Packed Quadword Left
  • vprord — Rotate Packed Doubleword Right
  • vprorq — Rotate Packed Quadword Right
  • vprorvd — Variable Rotate Packed Doubleword Right
  • vprorvq — Variable Rotate Packed Quadword Right
  • vpsadbw — Compute Sum of Absolute Differences
  • vpscatterdd — Scatter Packed Doubleword Values with Signed Doubleword Indices
  • vpscatterdq — Scatter Packed Quadword Values with Signed Doubleword Indices
  • vpscatterqd — Scatter Packed Doubleword Values with Signed Quadword Indices
  • vpscatterqq — Scatter Packed Quadword Values with Signed Quadword Indices
  • vpshufb — Packed Shuffle Bytes
  • vpshufd — Shuffle Packed Doublewords
  • vpshufhw — Shuffle Packed High Words
  • vpshuflw — Shuffle Packed Low Words
  • vpslld — Shift Packed Doubleword Data Left Logical
  • vpslldq — Shift Packed Double Quadword Left Logical
  • vpsllq — Shift Packed Quadword Data Left Logical
  • vpsllvd — Variable Shift Packed Doubleword Data Left Logical
  • vpsllvq — Variable Shift Packed Quadword Data Left Logical
  • vpsllvw — Variable Shift Packed Word Data Left Logical
  • vpsllw — Shift Packed Word Data Left Logical
  • vpsrad — Shift Packed Doubleword Data Right Arithmetic
  • vpsraq — Shift Packed Quadword Data Right Arithmetic
  • vpsravd — Variable Shift Packed Doubleword Data Right Arithmetic
  • vpsravq — Variable Shift Packed Quadword Data Right Arithmetic
  • vpsravw — Variable Shift Packed Word Data Right Arithmetic
  • vpsraw — Shift Packed Word Data Right Arithmetic
  • vpsrld — Shift Packed Doubleword Data Right Logical
  • vpsrldq — Shift Packed Double Quadword Right Logical
  • vpsrlq — Shift Packed Quadword Data Right Logical
  • vpsrlvd — Variable Shift Packed Doubleword Data Right Logical
  • vpsrlvq — Variable Shift Packed Quadword Data Right Logical
  • vpsrlvw — Variable Shift Packed Word Data Right Logical
  • vpsrlw — Shift Packed Word Data Right Logical
  • vpsubb — Subtract Packed Byte Integers
  • vpsubd — Subtract Packed Doubleword Integers
  • vpsubq — Subtract Packed Quadword Integers
  • vpsubsb — Subtract Packed Signed Byte Integers with Signed Saturation
  • vpsubsw — Subtract Packed Signed Word Integers with Signed Saturation
  • vpsubusb — Subtract Packed Unsigned Byte Integers with Unsigned Saturation
  • vpsubusw — Subtract Packed Unsigned Word Integers with Unsigned Saturation
  • vpsubw — Subtract Packed Word Integers
  • vpternlogd — Bitwise Ternary Logical Operation on Doubleword Values
  • vpternlogq — Bitwise Ternary Logical Operation on Quadword Values
  • vptestmb — Logical AND of Packed Byte Integer Values and Set Mask
  • vptestmd — Logical AND of Packed Doubleword Integer Values and Set Mask
  • vptestmq — Logical AND of Packed Quadword Integer Values and Set Mask
  • vptestmw — Logical AND of Packed Word Integer Values and Set Mask
  • vptestnmb — Logical NAND of Packed Byte Integer Values and Set Mask
  • vptestnmd — Logical NAND of Packed Doubleword Integer Values and Set Mask
  • vptestnmq — Logical NAND of Packed Quadword Integer Values and Set Mask
  • vptestnmw — Logical NAND of Packed Word Integer Values and Set Mask
  • vpunpckhbw — Unpack and Interleave High-Order Bytes into Words
  • vpunpckhdq — Unpack and Interleave High-Order Doublewords into Quadwords
  • vpunpckhqdq — Unpack and Interleave High-Order Quadwords into Double Quadwords
  • vpunpckhwd — Unpack and Interleave High-Order Words into Doublewords
  • vpunpcklbw — Unpack and Interleave Low-Order Bytes into Words
  • vpunpckldq — Unpack and Interleave Low-Order Doublewords into Quadwords
  • vpunpcklqdq — Unpack and Interleave Low-Order Quadwords into Double Quadwords
  • vpunpcklwd — Unpack and Interleave Low-Order Words into Doublewords
  • vpxord — Bitwise Logical Exclusive OR of Packed Doubleword Integers
  • vpxorq — Bitwise Logical Exclusive OR of Packed Quadword Integers
  • vrangepd — Range Restriction Calculation For Packed Pairs of Double-Precision Floating-Point Values
  • vrangeps — Range Restriction Calculation For Packed Pairs of Single-Precision Floating-Point Values
  • vrangesd — Range Restriction Calculation For a pair of Scalar Double-Precision Floating-Point Values
  • vrangess — Range Restriction Calculation For a pair of Scalar Single-Precision Floating-Point Values
  • vrcp14pd — Compute Approximate Reciprocals of Packed Double-Precision Floating-Point Values
  • vrcp14ps — Compute Approximate Reciprocals of Packed Single-Precision Floating-Point Values
  • vrcp14sd — Compute Approximate Reciprocal of a Scalar Double-Precision Floating-Point Value
  • vrcp14ss — Compute Approximate Reciprocal of a Scalar Single-Precision Floating-Point Value
  • vrcp28pd — Approximation to the Reciprocal of Packed Double-Precision Floating-Point Values with Less Than 2^-28 Relative Error
  • vrcp28ps — Approximation to the Reciprocal of Packed Single-Precision Floating-Point Values with Less Than 2^-28 Relative Error
  • vrcp28sd — Approximation to the Reciprocal of a Scalar Double-Precision Floating-Point Value with Less Than 2^-28 Relative Error
  • vrcp28ss — Approximation to the Reciprocal of a Scalar Single-Precision Floating-Point Value with Less Than 2^-28 Relative Error
  • vreducepd — Perform Reduction Transformation on Packed Double-Precision Floating-Point Values
  • vreduceps — Perform Reduction Transformation on Packed Single-Precision Floating-Point Values
  • vreducesd — Perform Reduction Transformation on a Scalar Double-Precision Floating-Point Value
  • vreducess — Perform Reduction Transformation on a Scalar Single-Precision Floating-Point Value
  • vrndscalepd — Round Packed Double-Precision Floating-Point Values To Include A Given Number Of Fraction Bits
  • vrndscaleph — Round Packed Half-Precision Floating-Point Values To Include A Given Number Of Fraction Bits
  • vrndscaleps — Round Packed Single-Precision Floating-Point Values To Include A Given Number Of Fraction Bits
  • vrndscalesd — Round Scalar Double-Precision Floating-Point Value To Include A Given Number Of Fraction Bits
  • vrndscalesh — Round Scalar Half-Precision Floating-Point Value To Include A Given Number Of Fraction Bits
  • vrndscaless — Round Scalar Single-Precision Floating-Point Value To Include A Given Number Of Fraction Bits
  • vrsqrt14pd — Compute Approximate Reciprocals of Square Roots of Packed Double-Precision Floating-Point Values
  • vrsqrt14ps — Compute Approximate Reciprocals of Square Roots of Packed Single-Precision Floating-Point Values
  • vrsqrt14sd — Compute Approximate Reciprocal of a Square Root of a Scalar Double-Precision Floating-Point Value
  • vrsqrt14ss — Compute Approximate Reciprocal of a Square Root of a Scalar Single-Precision Floating-Point Value
  • vrsqrt28pd — Approximation to the Reciprocal Square Root of Packed Double-Precision Floating-Point Values with Less Than 2^-28 Relative Error
  • vrsqrt28ps — Approximation to the Reciprocal Square Root of Packed Single-Precision Floating-Point Values with Less Than 2^-28 Relative Error
  • vrsqrt28sd — Approximation to the Reciprocal Square Root of a Scalar Double-Precision Floating-Point Value with Less Than 2^-28 Relative Error
  • vrsqrt28ss — Approximation to the Reciprocal Square Root of a Scalar Single-Precision Floating-Point Value with Less Than 2^-28 Relative Error
  • vscalefpd — Scale Packed Double-Precision Floating-Point Values With Double-Precision Floating-Point Values
  • vscalefps — Scale Packed Single-Precision Floating-Point Values With Single-Precision Floating-Point Values
  • vscalefsd — Scale Scalar Double-Precision Floating-Point Value With a Double-Precision Floating-Point Value
  • vscalefss — Scale Scalar Single-Precision Floating-Point Value With a Single-Precision Floating-Point Value
  • vscatterdpd — Scatter Packed Double-Precision Floating-Point Values with Signed Doubleword Indices
  • vscatterdps — Scatter Packed Single-Precision Floating-Point Values with Signed Doubleword Indices
  • vscatterpf0dpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hint with Intent to Write
  • vscatterpf0dps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hint with Intent to Write
  • vscatterpf0qpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hint with Intent to Write
  • vscatterpf0qps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hint with Intent to Write
  • vscatterpf1dpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hint with Intent to Write
  • vscatterpf1dps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hint with Intent to Write
  • vscatterpf1qpd — Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hint with Intent to Write
  • vscatterpf1qps — Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hint with Intent to Write
  • vscatterqpd — Scatter Packed Double-Precision Floating-Point Values with Signed Quadword Indices
  • vscatterqps — Scatter Packed Single-Precision Floating-Point Values with Signed Quadword Indices
  • vshuff32x4 — Shuffle 128-Bit Packed Single-Precision Floating-Point Values
  • vshuff64x2 — Shuffle 128-Bit Packed Double-Precision Floating-Point Values
  • vshufi32x4 — Shuffle 128-Bit Packed Doubleword Integer Values
  • vshufi64x2 — Shuffle 128-Bit Packed Quadword Integer Values
  • vshufpd — Shuffle Packed Double-Precision Floating-Point Values
  • vshufps — Shuffle Packed Single-Precision Floating-Point Values
  • vsqrtpd — Compute Square Roots of Packed Double-Precision Floating-Point Values
  • vsqrtps — Compute Square Roots of Packed Single-Precision Floating-Point Values
  • vsubpd — Subtract Packed Double-Precision Floating-Point Values
  • vsubps — Subtract Packed Single-Precision Floating-Point Values
  • vunpckhpd — Unpack and Interleave High Packed Double-Precision Floating-Point Values
  • vunpckhps — Unpack and Interleave High Packed Single-Precision Floating-Point Values
  • vunpcklpd — Unpack and Interleave Low Packed Double-Precision Floating-Point Values
  • vunpcklps — Unpack and Interleave Low Packed Single-Precision Floating-Point Values
  • vxorpd — Bitwise Logical XOR for Double-Precision Floating-Point Values
  • vxorps — Bitwise Logical XOR for Single-Precision Floating-Point Values

AVX-512 mask register instructions

140 instructions
  • addb
  • addd
  • addq
  • addw
  • andb
  • andd
  • andnb
  • andnd
  • andnq
  • andnw
  • andq
  • andw
  • kadd
  • kaddb — ADD Two 8-bit Masks
  • kaddd — ADD Two 32-bit Masks
  • kaddq — ADD Two 64-bit Masks
  • kaddw — ADD Two 16-bit Masks
  • kand
  • kandb — Bitwise Logical AND 8-bit Masks
  • kandd — Bitwise Logical AND 32-bit Masks
  • kandn
  • kandnb — Bitwise Logical AND NOT 8-bit Masks
  • kandnd — Bitwise Logical AND NOT 32-bit Masks
  • kandnq — Bitwise Logical AND NOT 64-bit Masks
  • kandnw — Bitwise Logical AND NOT 16-bit Masks
  • kandq — Bitwise Logical AND 64-bit Masks
  • kandw — Bitwise Logical AND 16-bit Masks
  • kmov
  • kmovb — Move 8-bit Mask
  • kmovd — Move 32-bit Mask
  • kmovq — Move 64-bit Mask
  • kmovw — Move 16-bit Mask
  • knot
  • knotb — NOT 8-bit Mask Register
  • knotd — NOT 32-bit Mask Register
  • knotq — NOT 64-bit Mask Register
  • knotw — NOT 16-bit Mask Register
  • kor
  • korb — Bitwise Logical OR 8-bit Masks
  • kord — Bitwise Logical OR 32-bit Masks
  • korq — Bitwise Logical OR 64-bit Masks
  • kortest
  • kortestb — OR 8-bit Masks and Set Flags
  • kortestd — OR 32-bit Masks and Set Flags
  • kortestq — OR 64-bit Masks and Set Flags
  • kortestw — OR 16-bit Masks and Set Flags
  • korw — Bitwise Logical OR 16-bit Masks
  • kshiftl
  • kshiftlb — Shift Left 8-bit Masks
  • kshiftld — Shift Left 32-bit Masks
  • kshiftlq — Shift Left 64-bit Masks
  • kshiftlw — Shift Left 16-bit Masks
  • kshiftr
  • kshiftrb — Shift Right 8-bit Masks
  • kshiftrd — Shift Right 32-bit Masks
  • kshiftrq — Shift Right 64-bit Masks
  • kshiftrw — Shift Right 16-bit Masks
  • kshl
  • kshlb
  • kshld
  • kshlq
  • kshlw
  • kshr
  • kshrb
  • kshrd
  • kshrq
  • kshrw
  • ktest
  • ktestb — Bit Test 8-bit Masks and Set Flags
  • ktestd — Bit Test 32-bit Masks and Set Flags
  • ktestq — Bit Test 64-bit Masks and Set Flags
  • ktestw — Bit Test 16-bit Masks and Set Flags
  • kunpck
  • kunpckbw — Unpack and Interleave 8-bit Masks
  • kunpckd
  • kunpckdq — Unpack and Interleave 32-bit Masks
  • kunpckq
  • kunpckw
  • kunpckwd — Unpack and Interleave 16-bit Masks
  • kxnor
  • kxnorb — Bitwise Logical XNOR 8-bit Masks
  • kxnord — Bitwise Logical XNOR 32-bit Masks
  • kxnorq — Bitwise Logical XNOR 64-bit Masks
  • kxnorw — Bitwise Logical XNOR 16-bit Masks
  • kxor
  • kxorb — Bitwise Logical XOR 8-bit Masks
  • kxord — Bitwise Logical XOR 32-bit Masks
  • kxorq — Bitwise Logical XOR 64-bit Masks
  • kxorw — Bitwise Logical XOR 16-bit Masks
  • movb
  • movw
  • notb
  • notd
  • notq
  • notw
  • orb
  • ord
  • orq
  • ortest
  • ortestb
  • ortestd
  • ortestq
  • ortestw
  • orw
  • shiftl
  • shiftlb
  • shiftld
  • shiftlq
  • shiftlw
  • shiftr
  • shiftrb
  • shiftrd
  • shiftrq
  • shiftrw
  • shlb
  • shlq
  • shlw
  • shrb
  • shrq
  • shrw
  • testb
  • testd
  • testq
  • testw
  • unpck
  • unpckbw
  • unpckd
  • unpckdq
  • unpckq
  • unpckw
  • unpckwd
  • xnor
  • xnorb
  • xnord
  • xnorq
  • xnorw
  • xorb
  • xord
  • xorq
  • xorw

AVX10.2 BF16 instructions

29 instructions
  • vaddbf16
  • vcmpbf16
  • vcomisbf16
  • vdivbf16
  • vfmadd132bf16
  • vfmadd213bf16
  • vfmadd231bf16
  • vfmsub132bf16
  • vfmsub213bf16
  • vfmsub231bf16
  • vfnmadd132bf16
  • vfnmadd213bf16
  • vfnmadd231bf16
  • vfnmsub132bf16
  • vfnmsub213bf16
  • vfnmsub231bf16
  • vfpclassbf16
  • vgetexpbf16
  • vgetmantbf16
  • vmaxbf16
  • vminbf16
  • vmulbf16
  • vrcpbf16
  • vreducebf16
  • vrndscalebf16
  • vrsqrtbf16
  • vscalefbf16
  • vsqrtbf16
  • vsubbf16

AVX10.2 Compare scalar fp with enhanced eflags instructions

6 instructions
  • vcomxsd
  • vcomxsh
  • vcomxss
  • vucomxsd
  • vucomxsh
  • vucomxss

AVX10.2 Convert instructions

14 instructions
  • vcvt2ph2bf8
  • vcvt2ph2bf8s
  • vcvt2ph2hf8
  • vcvt2ph2hf8s
  • vcvt2ps2phx
  • vcvtbiasph2bf8
  • vcvtbiasph2bf8s
  • vcvtbiasph2hf8
  • vcvtbiasph2hf8s
  • vcvthf82ph
  • vcvtph2bf8
  • vcvtph2bf8s
  • vcvtph2hf8
  • vcvtph2hf8s

AVX10.2 Integer and FP16 VNNI, media new instructions

14 instructions
  • vdpphps
  • vmpsadbw — Compute Multiple Packed Sums of Absolute Difference
  • vpdpbssd — Packed Dot Product of Signed-by-Singed Byte subvectors into Doubleword
  • vpdpbssds — Packed Dot Product of Signed-by-Singed Byte subvectors into Doubleword with Saturation
  • vpdpbsud — Packed Dot Product of Signed-by-Unsinged Byte subvectors into Doubleword
  • vpdpbsuds — Packed Dot Product of Signed-by-Unsinged Byte subvectors into Doubleword with Saturation
  • vpdpbuud — Packed Dot Product of Unsigned-by-Unsinged Byte subvectors into Doubleword
  • vpdpbuuds — Packed Dot Product of Unsigned-by-Unsinged Byte subvectors into Doubleword with Saturation
  • vpdpwsud — Packed Dot Product of Signed-by-Unsigned Word subvectors into Doubleword
  • vpdpwsuds — Packed Dot Product of Signed-by-Unsigned Word subvectors into Doubleword with Saturation
  • vpdpwusd — Packed Dot Product of Unsigned-by-Signed Word subvectors into Doubleword
  • vpdpwusds — Packed Dot Product of Unsigned-by-Signed Word subvectors into Doubleword with Saturation
  • vpdpwuud — Packed Dot Product of Unsigned-by-Unsigned Word subvectors into Doubleword
  • vpdpwuuds — Packed Dot Product of Unsigned-by-Unsigned Word subvectors into Doubleword with Saturation

AVX10.2 MINMAX instructions

7 instructions
  • vminmaxbf16
  • vminmaxpd
  • vminmaxph
  • vminmaxps
  • vminmaxsd
  • vminmaxsh
  • vminmaxss

AVX10.2 Saturating convert instructions

24 instructions
  • vcvtbf162ibs
  • vcvtbf162iubs
  • vcvtph2ibs
  • vcvtph2iubs
  • vcvtps2ibs
  • vcvtps2iubs
  • vcvttbf162ibs
  • vcvttbf162iubs
  • vcvttpd2dqs
  • vcvttpd2qqs
  • vcvttpd2udqs
  • vcvttpd2uqqs
  • vcvttph2ibs
  • vcvttph2iubs
  • vcvttps2dqs
  • vcvttps2ibs
  • vcvttps2iubs
  • vcvttps2qqs
  • vcvttps2udqs
  • vcvttps2uqqs
  • vcvttsd2sis
  • vcvttsd2usis
  • vcvttss2sis
  • vcvttss2usis

AVX512 4-iteration Dot Product

2 instructions
  • v4dpwssd
  • v4dpwssds

AVX512 4-iteration Multiply-Add

4 instructions
  • v4fmaddps
  • v4fmaddss
  • v4fnmaddps
  • v4fnmaddss

AVX512 Bfloat16 instructions

3 instructions
  • vcvtne2ps2bf16 — Convert with Nearest-Even rounding 2 Single-Precision FP vectors into BFloat16 FP vector
  • vcvtneps2bf16 — Convert with Nearest-Even rounding a Single-Precision FP vector into a BFloat16 FP vector
  • vdpbf16ps — Packed Dot Product of BFloat16 FP subvectors into Single-Precision FP values

AVX512 Bit Algorithms

5 instructions
  • vpopcntb — Packed Population Count for Byte Integers
  • vpopcntd — Packed Population Count for Doubleword Integers
  • vpopcntq — Packed Population Count for Quadword Integers
  • vpopcntw — Packed Population Count for Word Integers
  • vpshufbitqmb — Shuffle Bits From Quadword Elements Using Byte Indexes Into Mask

AVX512 mask intersect instructions

2 instructions
  • vp2intersectd
  • vp2intersectq

AVX512 Vector Bit Manipulation Instructions 2

16 instructions
  • vpcompressb — Store Sparse Packed Byte Integer Values into Dense Memory/Register
  • vpcompressw — Store Sparse Packed Word Integer Values into Dense Memory/Register
  • vpexpandb — Load Sparse Packed Byte Integer Values from Dense Memory/Register
  • vpexpandw — Load Sparse Packed Word Integer Values from Dense Memory/Register
  • vpshldd — Concatenate and Shift Packed Doubleword Data Left Logical
  • vpshldq — Concatenate and Shift Packed Quadword Data Left Logical
  • vpshldvd — Concatenate and Variable Shift Packed Doubleword Data Left Logical
  • vpshldvq — Concatenate and Variable Shift Packed Quadword Data Left Logical
  • vpshldvw — Concatenate and Variable Shift Packed Word Data Left Logical
  • vpshldw — Concatenate and Shift Packed Word Data Left Logical
  • vpshrdd — Concatenate and Shift Packed Doubleword Data Right Logical
  • vpshrdq — Concatenate and Shift Packed Quadword Data Right Logical
  • vpshrdvd — Concatenate and Variable Shift Packed Doubleword Data Right Logical
  • vpshrdvq — Concatenate and Variable Shift Packed Quadword Data Right Logical
  • vpshrdvw — Concatenate and Variable Shift Packed Word Data Right Logical
  • vpshrdw — Concatenate and Shift Packed Word Data Right Logical

AVX512 VNNI

4 instructions
  • vpdpbusd — Packed Dot Product of Unsigned-by-Singed Byte subvectors into Doubleword
  • vpdpbusds — Packed Dot Product of Unsigned-by-Singed Byte subvectors into Doubleword with Saturation
  • vpdpwssd — Packed Dot Product of Signed-by-Signed Word subvectors into Doubleword
  • vpdpwssds — Packed Dot Product of Signed-by-Signed Word subvectors into Doubleword with Saturation

BMI1 and BMI2 bit operations

10 instructions
  • andn — Logical AND NOT
  • bextr — Bit Field Extract
  • blsi — Isolate Lowest Set Bit
  • blsmsk — Mask From Lowest Set Bit
  • blsr — Reset Lowest Set Bit
  • bzhi — Zero High Bits Starting with Specified Bit Position
  • lzcnt — Count the Number of Leading Zero Bits
  • pdep — Parallel Bits Deposit
  • pext — Parallel Bits Extract
  • tzcnt — Count the Number of Trailing Zero Bits

Conditional instructions (extensions)

90 instructions
  • cfcmova (APX)
  • cfcmovae (APX)
  • cfcmovb (APX)
  • cfcmovbe (APX)
  • cfcmovc (APX)
  • cfcmove (APX)
  • cfcmovg (APX)
  • cfcmovge (APX)
  • cfcmovl (APX)
  • cfcmovle (APX)
  • cfcmovna (APX)
  • cfcmovnae (APX)
  • cfcmovnb (APX)
  • cfcmovnbe (APX)
  • cfcmovnc (APX)
  • cfcmovne (APX)
  • cfcmovng (APX)
  • cfcmovnge (APX)
  • cfcmovnl (APX)
  • cfcmovnle (APX)
  • cfcmovno (APX)
  • cfcmovnp (APX)
  • cfcmovns (APX)
  • cfcmovnz (APX)
  • cfcmovo (APX)
  • cfcmovp (APX)
  • cfcmovpe (APX)
  • cfcmovpo (APX)
  • cfcmovs (APX)
  • cfcmovz (APX)
  • cmpaexadd (CMPCCXADD)
  • cmpaxadd (CMPCCXADD)
  • cmpbexadd — Compare for Below or Equals and Add (CMPCCXADD)
  • cmpbxadd — Compare for Below and Add (CMPCCXADD)
  • cmpcxadd (CMPCCXADD)
  • cmpexadd (CMPCCXADD)
  • cmpgexadd (CMPCCXADD)
  • cmpgxadd (CMPCCXADD)
  • cmplexadd — Compare for Less or Equals and Add (CMPCCXADD)
  • cmplxadd — Compare for Less and Add (CMPCCXADD)
  • cmpnaexadd (CMPCCXADD)
  • cmpnaxadd (CMPCCXADD)
  • cmpnbexadd — Compare for Not Below or Equals and Add (CMPCCXADD)
  • cmpnbxadd — Compare for Not Below and Add (CMPCCXADD)
  • cmpncxadd (CMPCCXADD)
  • cmpnexadd (CMPCCXADD)
  • cmpngexadd (CMPCCXADD)
  • cmpngxadd (CMPCCXADD)
  • cmpnlexadd — Compare for Not Less or Equals and Add (CMPCCXADD)
  • cmpnlxadd — Compare for Not Less and Add (CMPCCXADD)
  • cmpnoxadd — Compare for Not Overflow and Add (CMPCCXADD)
  • cmpnpxadd — Compare for Not Parity and Add (CMPCCXADD)
  • cmpnsxadd — Compare for Not Sign and Add (CMPCCXADD)
  • cmpnzxadd — Compare for Not Zero and Add (CMPCCXADD)
  • cmpoxadd — Compare for Overflow and Add (CMPCCXADD)
  • cmppexadd (CMPCCXADD)
  • cmppoxadd (CMPCCXADD)
  • cmppxadd — Compare for Parity and Add (CMPCCXADD)
  • cmpsxadd — Compare for Sign and Add (CMPCCXADD)
  • cmpzxadd — Compare for Zero and Add (CMPCCXADD)
  • setaezu (APX)
  • setazu (APX)
  • setbezu (APX)
  • setbzu (APX)
  • setczu (APX)
  • setezu (APX)
  • setgezu (APX)
  • setgzu (APX)
  • setlezu (APX)
  • setlzu (APX)
  • setnaezu (APX)
  • setnazu (APX)
  • setnbezu (APX)
  • setnbzu (APX)
  • setnczu (APX)
  • setnezu (APX)
  • setngezu (APX)
  • setngzu (APX)
  • setnlezu (APX)
  • setnlzu (APX)
  • setnozu (APX)
  • setnpzu (APX)
  • setnszu (APX)
  • setnzzu (APX)
  • setozu (APX)
  • setpezu (APX)
  • setpozu (APX)
  • setpzu (APX)
  • setszu (APX)
  • setzzu (APX)

doc 319433-034 May 2018

4 instructions
  • cldemote — Cache Line Demote
  • movdir64b — MOVe to DIRect store 64 Bytes
  • movdiri — MOVe to DIRect store Integer
  • pconfig

doc 319433-058 June 2025

2 instructions
  • pbndkb
  • prefetchrst2

Extended Page Tables VMX instructions

2 instructions
  • invept
  • invvpid

Flag register instructions (extensions)

2 instructions
  • clac (SMAP)
  • stac (SMAP)

Galois field operations (GFNI)

6 instructions
  • gf2p8affineinvqb — Galois Field (2^8) Affine Inverse Transformation
  • gf2p8affineqb — Galois Field (2^8) Affine Transformation
  • gf2p8mulb — Galois Field Multiply Bytes
  • vgf2p8affineinvqb — Galois Field (2^8) Affine Inverse Transformation
  • vgf2p8affineqb — Galois Field (2^8) Affine Transformation
  • vgf2p8mulb — Galois Field Multiply Bytes

Generic memory operations

6 instructions
  • prefetchit0 — Prefetch Code Into Instruction Caches using IT0 Hint
  • prefetchit1 — Prefetch Code Into Instruction Caches using IT1 Hint
  • prefetchnta — Prefetch Data Into Caches using NTA Hint
  • prefetcht0 — Prefetch Data Into Caches using T0 Hint
  • prefetcht1 — Prefetch Data Into Caches using T1 Hint
  • prefetcht2 — Prefetch Data Into Caches using T2 Hint

Geode (Cyrix) 3DNow! additions

2 instructions
  • pfrcpv
  • pfrsqrtv

History reset

1 instruction
  • hreset

I/O instructions

2 instructions
  • in
  • out

Instructions from ISE doc 319433-040, June 2020

4 instructions
  • enqcmd
  • enqcmds
  • xresldtrk
  • xsusldtrk

Intel Advanced Matrix Extensions (AMX)

44 instructions
  • ldtilecfg — LoaD TILE ConFiGuration
  • sttilecfg — STore TILE ConFiGuration
  • t2rpntlvwz0
  • t2rpntlvwz0rs
  • t2rpntlvwz0rst1
  • t2rpntlvwz0t1
  • t2rpntlvwz1
  • t2rpntlvwz1rs
  • t2rpntlvwz1rst1
  • t2rpntlvwz1t1
  • tcmmimfp16ps — Tile Complex Matrix Multiply IMaginary part of FP16 tiles with Packed Single-precision accumulation
  • tcmmrlfp16ps — Tile Complex Matrix Multiply ReaL part of FP16 tiles with Packed Single-precision accumulation
  • tconjtcmmimfp16ps
  • tconjtfp16
  • tcvtrowd2ps
  • tcvtrowps2bf16h
  • tcvtrowps2bf16l
  • tcvtrowps2phh
  • tcvtrowps2phl
  • tdpbf16ps — Tile Dot Product of BF16 tiles with Packed Single-precision accumulation
  • tdpbf8ps
  • tdpbhf8ps
  • tdpbssd — Tile Dot Product of Signed bytes by Signed bytes with Doubleword accumulation
  • tdpbsud — Tile Dot Product of Signed bytes by Unsigned bytes with Doubleword accumulation
  • tdpbusd — Tile Dot Product of Unsigned bytes by Signed bytes with Doubleword accumulation
  • tdpbuud — Tile Dot Product of Unsigned bytes by Unsigned bytes with Doubleword accumulation
  • tdpfp16ps — Tile Dot Product of FP16 tiles with Packed Single-precision accumulation
  • tdphbf8ps
  • tdphf8ps
  • tileloadd — TILE LOAD Data
  • tileloaddrs
  • tileloaddrst1
  • tileloaddt1 — TILE LOAD Data with T1 caching hint
  • tilemovrow
  • tilerelease — TILE RELEASE register state
  • tilestored — TILE STORE Data
  • tilezero — TILE ZERO data
  • tmmultf32ps
  • ttcmmimfp16ps
  • ttcmmrlfp16ps
  • ttdpbf16ps
  • ttdpfp16ps
  • ttmmultf32ps
  • ttransposed

Intel AES instructions

6 instructions
  • aesdec — Perform One Round of an AES Decryption Flow
  • aesdeclast — Perform Last Round of an AES Decryption Flow
  • aesenc — Perform One Round of an AES Encryption Flow
  • aesenclast — Perform Last Round of an AES Encryption Flow
  • aesimc — Perform the AES InvMixColumn Transformation
  • aeskeygenassist — AES Round Key Generation Assist

Intel AES Key Locker

11 instructions
  • aesdec128kl
  • aesdec256kl
  • aesdecwide128kl
  • aesdecwide256kl
  • aesenc128kl
  • aesenc256kl
  • aesencwide128kl
  • aesencwide256kl
  • encodekey128
  • encodekey256
  • loadiwkey

Intel AVX AES instructions

2 instructions
  • vaesimc — Perform the AES InvMixColumn Transformation
  • vaeskeygenassist — AES Round Key Generation Assist

Intel AVX Carry-Less Multiplication instructions (CLMUL)

21 instructions
  • vfcmulcph — Fused Conjugate Multiply of Complex Packed Half-Precision Floating-Point Values
  • vfmadd132sh
  • vfmadd213sh
  • vfmadd231sh
  • vfmsub132sh — Fused Multiply-Subtract of Scalar Half-Precision Floating-Point Values
  • vfmsub213sh — Fused Multiply-Subtract of Scalar Half-Precision Floating-Point Values
  • vfmsub231sh — Fused Multiply-Subtract of Scalar Half-Precision Floating-Point Values
  • vfmulcph — Fused Fused Multiply of Complex Packed Half-Precision Floating-Point Values
  • vfnmadd132sh
  • vfnmadd213sh
  • vfnmadd231sh
  • vfnmsub132sh — Fused Negative Multiply-Subtract of Scalar Half-Precision Floating-Point Values
  • vfnmsub213sh — Fused Negative Multiply-Subtract of Scalar Half-Precision Floating-Point Values
  • vfnmsub231sh — Fused Negative Multiply-Subtract of Scalar Half-Precision Floating-Point Values
  • vmaxsh — Return Maximum Scalar Half-Precision Floating-Point Value
  • vminsh — Return Minimum Scalar Half-Precision Floating-Point Value
  • vpclmulhqhqdq
  • vpclmulhqlqdq
  • vpclmullqhqdq
  • vpclmullqlqdq
  • vpclmulqdq — Carry-Less Quadword Multiplication

Intel AVX instructions

204 instructions
  • vaddsd — Add Scalar Double-Precision Floating-Point Values
  • vaddss — Add Scalar Single-Precision Floating-Point Values
  • vaddsubpd — Packed Double-FP Add/Subtract
  • vaddsubps — Packed Single-FP Add/Subtract
  • vblendpd — Blend Packed Double Precision Floating-Point Values
  • vblendps — Blend Packed Single Precision Floating-Point Values
  • vblendvpd — Variable Blend Packed Double Precision Floating-Point Values
  • vblendvps — Variable Blend Packed Single Precision Floating-Point Values
  • vbroadcastf128 — Broadcast 128 Bit of Floating-Point Data
  • vcmpeq_ospd
  • vcmpeq_osps
  • vcmpeq_ossd
  • vcmpeq_osss
  • vcmpeq_uqsd
  • vcmpeq_uqss
  • vcmpeq_ussd
  • vcmpeq_usss
  • vcmpeqsd
  • vcmpeqss
  • vcmpfalse_oqsd
  • vcmpfalse_oqss
  • vcmpfalse_ossd
  • vcmpfalse_osss
  • vcmpfalsesd
  • vcmpfalsess
  • vcmpge_oqsd
  • vcmpge_oqss
  • vcmpge_ossd
  • vcmpge_osss
  • vcmpgesd
  • vcmpgess
  • vcmpgt_oqsd
  • vcmpgt_oqss
  • vcmpgt_ossd
  • vcmpgt_osss
  • vcmpgtsd
  • vcmpgtss
  • vcmple_oqsd
  • vcmple_oqss
  • vcmple_ossd
  • vcmple_osss
  • vcmplesd
  • vcmpless
  • vcmplt_oqsd
  • vcmplt_oqss
  • vcmplt_ossd
  • vcmplt_osss
  • vcmpltsd
  • vcmpltss
  • vcmpneq_oqsd
  • vcmpneq_oqss
  • vcmpneq_ossd
  • vcmpneq_osss
  • vcmpneq_uqsd
  • vcmpneq_uqss
  • vcmpneq_ussd
  • vcmpneq_usss
  • vcmpneqsd
  • vcmpneqss
  • vcmpnge_uqsd
  • vcmpnge_uqss
  • vcmpnge_ussd
  • vcmpnge_usss
  • vcmpngesd
  • vcmpngess
  • vcmpngt_uqsd
  • vcmpngt_uqss
  • vcmpngt_ussd
  • vcmpngt_usss
  • vcmpngtsd
  • vcmpngtss
  • vcmpnle_uqsd
  • vcmpnle_uqss
  • vcmpnle_ussd
  • vcmpnle_usss
  • vcmpnlesd
  • vcmpnless
  • vcmpnlt_uqsd
  • vcmpnlt_uqss
  • vcmpnlt_ussd
  • vcmpnlt_usss
  • vcmpnltsd
  • vcmpnltss
  • vcmpord_qsd
  • vcmpord_qss
  • vcmpord_ssd
  • vcmpord_sss
  • vcmpordsd
  • vcmpordss
  • vcmpsd — Compare Scalar Double-Precision Floating-Point Values
  • vcmpss — Compare Scalar Single-Precision Floating-Point Values
  • vcmptrue_uqsd
  • vcmptrue_uqss
  • vcmptrue_ussd
  • vcmptrue_usss
  • vcmptruesd
  • vcmptruess
  • vcmpunord_qsd
  • vcmpunord_qss
  • vcmpunord_ssd
  • vcmpunord_sss
  • vcmpunordsd
  • vcmpunordss
  • vcomisd — Compare Scalar Ordered Double-Precision Floating-Point Values and Set EFLAGS
  • vcomiss — Compare Scalar Ordered Single-Precision Floating-Point Values and Set EFLAGS
  • vcvtpd2dq — Convert Packed Double-Precision FP Values to Packed Dword Integers
  • vcvtpd2ps — Convert Packed Double-Precision FP Values to Packed Single-Precision FP Values
  • vcvtsd2si — Convert Scalar Double-Precision FP Value to Integer
  • vcvtsd2ss — Convert Scalar Double-Precision FP Value to Scalar Single-Precision FP Value
  • vcvtsi2sd — Convert Dword Integer to Scalar Double-Precision FP Value
  • vcvtsi2ss — Convert Dword Integer to Scalar Single-Precision FP Value
  • vcvtss2sd — Convert Scalar Single-Precision FP Value to Scalar Double-Precision FP Value
  • vcvtss2si — Convert Scalar Single-Precision FP Value to Dword Integer
  • vcvttpd2dq — Convert with Truncation Packed Double-Precision FP Values to Packed Dword Integers
  • vcvttsd2si — Convert with Truncation Scalar Double-Precision FP Value to Signed Integer
  • vcvttss2si — Convert with Truncation Scalar Single-Precision FP Value to Dword Integer
  • vdivsd — Divide Scalar Double-Precision Floating-Point Values
  • vdivss — Divide Scalar Single-Precision Floating-Point Values
  • vdppd — Dot Product of Packed Double Precision Floating-Point Values
  • vdpps — Dot Product of Packed Single Precision Floating-Point Values
  • vextractf128 — Extract Packed Floating-Point Values
  • vhaddpd — Packed Double-FP Horizontal Add
  • vhaddps — Packed Single-FP Horizontal Add
  • vhsubpd — Packed Double-FP Horizontal Subtract
  • vhsubps — Packed Single-FP Horizontal Subtract
  • vinsertf128 — Insert Packed Floating-Point Values
  • vinsertps — Insert Packed Single Precision Floating-Point Value
  • vlddqu — Load Unaligned Integer 128 Bits
  • vldmxcsr — Load MXCSR Register
  • vldqqu
  • vmaskmovdqu — Store Selected Bytes of Double Quadword
  • vmaskmovpd — Conditional Move Packed Double-Precision Floating-Point Values
  • vmaskmovps — Conditional Move Packed Single-Precision Floating-Point Values
  • vmaxsd — Return Maximum Scalar Double-Precision Floating-Point Value
  • vmaxss — Return Maximum Scalar Single-Precision Floating-Point Value
  • vminsd — Return Minimum Scalar Double-Precision Floating-Point Value
  • vminss — Return Minimum Scalar Single-Precision Floating-Point Value
  • vmovd — Move Doubleword
  • vmovdqa — Move Aligned Double Quadword
  • vmovdqu — Move Unaligned Double Quadword
  • vmovhlps — Move Packed Single-Precision Floating-Point Values High to Low
  • vmovhpd — Move High Packed Double-Precision Floating-Point Value
  • vmovhps — Move High Packed Single-Precision Floating-Point Values
  • vmovlhps — Move Packed Single-Precision Floating-Point Values Low to High
  • vmovlpd — Move Low Packed Double-Precision Floating-Point Value
  • vmovlps — Move Low Packed Single-Precision Floating-Point Values
  • vmovmskpd — Extract Packed Double-Precision Floating-Point Sign Mask
  • vmovmskps — Extract Packed Single-Precision Floating-Point Sign Mask
  • vmovntqq
  • vmovq — Move Quadword
  • vmovqqa
  • vmovqqu
  • vmovsd — Move Scalar Double-Precision Floating-Point Value
  • vmovss — Move Scalar Single-Precision Floating-Point Values
  • vmulsd — Multiply Scalar Double-Precision Floating-Point Values
  • vmulss — Multiply Scalar Single-Precision Floating-Point Values
  • vpand — Packed Bitwise Logical AND
  • vpandn — Packed Bitwise Logical AND NOT
  • vpblendvb — Variable Blend Packed Bytes
  • vpblendw — Blend Packed Words
  • vpcmpestri — Packed Compare Explicit Length Strings, Return Index
  • vpcmpestrm — Packed Compare Explicit Length Strings, Return Mask
  • vpcmpistri — Packed Compare Implicit Length Strings, Return Index
  • vpcmpistrm — Packed Compare Implicit Length Strings, Return Mask
  • vperm2f128 — Permute Floating-Point Values
  • vpextrd — Extract Doubleword
  • vpextrq — Extract Quadword
  • vphaddd — Packed Horizontal Add Doubleword Integer
  • vphaddsw — Packed Horizontal Add Signed Word Integers with Signed Saturation
  • vphaddw — Packed Horizontal Add Word Integers
  • vphminposuw — Packed Horizontal Minimum of Unsigned Word Integers
  • vphsubd — Packed Horizontal Subtract Doubleword Integers
  • vphsubsw — Packed Horizontal Subtract Signed Word Integers with Signed Saturation
  • vphsubw — Packed Horizontal Subtract Word Integers
  • vpinsrb — Insert Byte
  • vpinsrd — Insert Doubleword
  • vpinsrq — Insert Quadword
  • vpinsrw — Insert Word
  • vpmovmskb — Move Byte Mask
  • vpor — Packed Bitwise Logical OR
  • vpsignb — Packed Sign of Byte Integers
  • vpsignd — Packed Sign of Doubleword Integers
  • vpsignw — Packed Sign of Word Integers
  • vptest — Packed Logical Compare
  • vpxor — Packed Bitwise Logical Exclusive OR
  • vrcpps — Compute Approximate Reciprocals of Packed Single-Precision Floating-Point Values
  • vrcpss — Compute Approximate Reciprocal of Scalar Single-Precision Floating-Point Values
  • vroundpd — Round Packed Double Precision Floating-Point Values
  • vroundps — Round Packed Single Precision Floating-Point Values
  • vroundsd — Round Scalar Double Precision Floating-Point Values
  • vroundss — Round Scalar Single Precision Floating-Point Values
  • vrsqrtps — Compute Reciprocals of Square Roots of Packed Single-Precision Floating-Point Values
  • vrsqrtss — Compute Reciprocal of Square Root of Scalar Single-Precision Floating-Point Value
  • vsqrtsd — Compute Square Root of Scalar Double-Precision Floating-Point Value
  • vsqrtss — Compute Square Root of Scalar Single-Precision Floating-Point Value
  • vstmxcsr — Store MXCSR Register State
  • vsubsd — Subtract Scalar Double-Precision Floating-Point Values
  • vsubss — Subtract Scalar Single-Precision Floating-Point Values
  • vtestpd — Packed Double-Precision Floating-Point Bit Test
  • vtestps — Packed Single-Precision Floating-Point Bit Test
  • vucomisd — Unordered Compare Scalar Double-Precision Floating-Point Values and Set EFLAGS
  • vucomiss — Unordered Compare Scalar Single-Precision Floating-Point Values and Set EFLAGS
  • vzeroall — Zero All YMM Registers
  • vzeroupper — Zero Upper Bits of YMM Registers

Intel AVX2 instructions

7 instructions
  • vbroadcasti128 — Broadcast 128 Bits of Integer Data
  • vextracti128 — Extract Packed Integer Values
  • vinserti128 — Insert Packed Integer Values
  • vpblendd — Blend Packed Doublewords
  • vperm2i128 — Permute 128-Bit Integer Values
  • vpmaskmovd — Conditional Move Packed Doubleword Integers
  • vpmaskmovq — Conditional Move Packed Quadword Integers

Intel AVX512-FP16 instructions

114 instructions
  • vaddph — Add Packed Half-Precision Floating-Point Values
  • vaddsh — Add Scalar Half-Precision Floating-Point Values
  • vcmpph — Compare Packed Half-Precision Floating-Point Values
  • vcmpsh — Compare Scalar Half-Precision Floating-Point Values
  • vcomish — Compare Scalar Ordered Half-Precision Floating-Point Values and Set EFLAGS
  • vcvtdq2ph — Convert Packed Dword Integers to Packed Half-Precision FP Values
  • vcvtpd2ph — Convert Packed Double-Precision FP Values to Packed Half-Precision FP Values
  • vcvtph2dq — Convert Packed Half-Precision FP Values to Packed Dword Integers
  • vcvtph2pd — Convert Packed Half-Precision FP Values to Packed Double-Precision FP Values
  • vcvtph2ps — Convert Half-Precision FP Values to Single-Precision FP Values
  • vcvtph2psx — Convert Half-Precision FP Values to Single-Precision FP Values
  • vcvtph2qq — Convert Packed Half Precision Floating-Point Values to Packed Singed Quadword Integer Values
  • vcvtph2udq — Convert Packed Half-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Values
  • vcvtph2uqq — Convert Packed Half Precision Floating-Point Values to Packed Unsigned Quadword Integer Values
  • vcvtph2uw — Convert Packed Half-Precision Floating-Point Values to Packed Unsigned Word Integer Values
  • vcvtph2w — Convert Packed Half-Precision Floating-Point Values to Packed Word Integer Values
  • vcvtps2ph — Convert Single-Precision FP value to Half-Precision FP value
  • vcvtps2phx — Convert Single-Precision FP value to Half-Precision FP value
  • vcvtqq2ph — Convert Packed Quadword Integers to Packed Half-Precision Floating-Point Values
  • vcvtsd2sh — Convert Scalar Double-Precision FP Value to Scalar Half-Precision FP Value
  • vcvtsh2sd — Convert Scalar Half-Precision FP Value to Scalar Double-Precision FP Value
  • vcvtsh2si — Convert Scalar Half-Precision FP Value to Dword Integer
  • vcvtsh2ss — Convert Scalar Half-Precision FP Value to Scalar Double-Precision FP Value
  • vcvtsh2usi — Convert Scalar Half-Precision Floating-Point Value to Unsigned Doubleword Integer
  • vcvtsi2sh — Convert Dword Integer to Scalar Half-Precision FP Value
  • vcvtss2sh — Convert Scalar Single-Precision FP Value to Scalar Half-Precision FP Value
  • vcvttph2dq — Convert with Truncation Packed Half-Precision FP Values to Packed Dword Integers
  • vcvttph2qq — Convert with Truncation Packed Half Precision Floating-Point Values to Packed Singed Quadword Integer Values
  • vcvttph2udq — Convert with Truncation Packed Half-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Values
  • vcvttph2uqq — Convert with Truncation Packed Half Precision Floating-Point Values to Packed Unsigned Quadword Integer Values
  • vcvttph2uw — Convert with Truncation Packed Half-Precision Floating-Point Values to Packed Unsigned Word Integer Values
  • vcvttph2w — Convert with Truncation Packed Half-Precision Floating-Point Values to Packed Word Integer Values
  • vcvttsh2si — Convert with Truncation Scalar Half-Precision FP Value to Dword Integer
  • vcvttsh2usi — Convert with Truncation Scalar Half-Precision Floating-Point Value to Unsigned Integer
  • vcvtudq2ph — Convert Packed Unsigned Doubleword Integers to Packed Half-Precision Floating-Point Values
  • vcvtuqq2ph — Convert Packed Unsigned Quadword Integers to Packed Half-Precision Floating-Point Values
  • vcvtusi2sh — Convert Unsigned Integer to Scalar Half-Precision Floating-Point Value
  • vcvtuw2ph — Convert Packed Unsigned Word Integers to Packed Half-Precision Floating-Point Values
  • vcvtw2ph — Convert Packed Word Integers to Packed Half-Precision Floating-Point Values
  • vdivph — Divide Packed Half-Precision Floating-Point Values
  • vdivsh — Divide Scalar Half-Precision Floating-Point Values
  • vendscaleph
  • vendscalesh
  • vfcmaddcph — Fused Conjugate Multiply-Add of Complex Packed Half-Precision Floating-Point Values
  • vfcmaddcsh — Fused Conjugate Multiply-Add of Complex Scalar Half-Precision Floating-Point Values
  • vfcmulcpch
  • vfcmulcsh — Fused Conjugate Multiply of Complex Scalar Half-Precision Floating-Point Values
  • vfmadd132ph — Fused Multiply-Add of Packed Half-Precision Floating-Point Values
  • vfmadd213ph — Fused Multiply-Add of Packed Half-Precision Floating-Point Values
  • vfmadd231ph — Fused Multiply-Add of Packed Half-Precision Floating-Point Values
  • vfmaddcph — Fused Multiply-Add of Complex Packed Half-Precision Floating-Point Values
  • vfmaddcsh — Fused Multiply-Add of Complex Scalar Half-Precision Floating-Point Values
  • vfmaddsub132ph — Fused Multiply-Alternating Add/Subtract of Packed Half-Precision Floating-Point Values
  • vfmaddsub213ph — Fused Multiply-Alternating Add/Subtract of Packed Half-Precision Floating-Point Values
  • vfmaddsub231ph — Fused Multiply-Alternating Add/Subtract of Packed Half-Precision Floating-Point Values
  • vfmsub132ph — Fused Multiply-Subtract of Packed Half-Precision Floating-Point Values
  • vfmsub213ph — Fused Multiply-Subtract of Packed Half-Precision Floating-Point Values
  • vfmsub231ph — Fused Multiply-Subtract of Packed Half-Precision Floating-Point Values
  • vfmsubadd132ph — Fused Multiply-Alternating Subtract/Add of Packed Half-Precision Floating-Point Values
  • vfmsubadd213ph — Fused Multiply-Alternating Subtract/Add of Packed Half-Precision Floating-Point Values
  • vfmsubadd231ph — Fused Multiply-Alternating Subtract/Add of Packed Half-Precision Floating-Point Values
  • vfmulcpch
  • vfmulcsh — Fused Multiply of Complex Scalar Half-Precision Floating-Point Values
  • vfnmadd132ph — Fused Negative Multiply-Add of Packed Half-Precision Floating-Point Values
  • vfnmadd213ph — Fused Negative Multiply-Add of Packed Half-Precision Floating-Point Values
  • vfnmadd231ph — Fused Negative Multiply-Add of Packed Half-Precision Floating-Point Values
  • vfnmsub132ph — Fused Negative Multiply-Subtract of Packed Half-Precision Floating-Point Values
  • vfnmsub213ph — Fused Negative Multiply-Subtract of Packed Half-Precision Floating-Point Values
  • vfnmsub231ph — Fused Negative Multiply-Subtract of Packed Half-Precision Floating-Point Values
  • vfpclassph — Test Class of Packed Half-Precision Floating-Point Values
  • vfpclasssh — Test Class of Scalar Half-Precision Floating-Point Value
  • vgetexpph — Extract Exponents of Packed Half-Precision Floating-Point Values as Half-Precision Floating-Point Values
  • vgetexpsh — Extract Exponent of Scalar Half-Precision Floating-Point Value as Half-Precision Floating-Point Value
  • vgetmantph — Extract Normalized Mantissas from Packed Half-Precision Floating-Point Values
  • vgetmantsh — Extract Normalized Mantissa from Scalar Half-Precision Floating-Point Value
  • vgetmaxph
  • vgetmaxsh
  • vgetminph
  • vgetminsh
  • vmovsh — Move Scalar Half-Precision Floating-Point Values
  • vmovw — Move Word
  • vmulph — Multiply Packed Half-Precision Floating-Point Values
  • vmulsh — Fused Multiply Scalar Half-Precision Floating-Point Values
  • vpmadd132ph
  • vpmadd132sh
  • vpmadd213ph
  • vpmadd213sh
  • vpmadd231ph
  • vpmadd231sh
  • vpmsub132ph
  • vpmsub132sh
  • vpmsub213ph
  • vpmsub213sh
  • vpmsub231ph
  • vpmsub231sh
  • vpnmadd132sh
  • vpnmadd213sh
  • vpnmadd231sh
  • vpnmsub132sh
  • vpnmsub213sh
  • vpnmsub231sh
  • vrcpph — Compute Approximate Reciprocals of Packed Half-Precision Floating-Point Values
  • vrcpsh — Compute Approximate Reciprocal of Scalar Half-Precision Floating-Point Values
  • vreduceph — Perform Reduction Transformation on Packed Half-Precision Floating-Point Values
  • vreducesh — Perform Reduction Transformation on a Scalar Half-Precision Floating-Point Value
  • vrsqrtph — Compute Reciprocals of Square Roots of Packed Half-Precision Floating-Point Values
  • vrsqrtsh — Compute Reciprocal of Square Root of Scalar Half-Precision Floating-Point Value
  • vscalefph — Scale Packed Half-Precision Floating-Point Values With Half-Precision Floating-Point Values
  • vscalefsh — Scale Scalar Half-Precision Floating-Point Value With a Half-Precision Floating-Point Value
  • vsqrtph — Compute Square Roots of Packed Half-Precision Floating-Point Values
  • vsqrtsh — Compute Square Root of Scalar Half-Precision Floating-Point Value
  • vsubph — Subtract Packed Half-Precision Floating-Point Values
  • vsubsh — Subtract Scalar Half-Precision Floating-Point Values
  • vucomish — Unordered Compare Scalar Half-Precision Floating-Point Values and Set EFLAGS

Intel Carry-Less Multiplication instructions (CLMUL)

5 instructions
  • pclmulhqhqdq
  • pclmulhqlqdq
  • pclmullqhqdq
  • pclmullqlqdq
  • pclmulqdq — Carry-Less Quadword Multiplication

Intel Control-Flow Enforcement Technology (CET)

14 instructions
  • clrssbsy
  • endbr32
  • endbr64 — END (terminate) BRanch in 64-bit mode
  • incsspd
  • incsspq
  • rdsspd
  • rdsspq
  • rstorssp
  • saveprevssp
  • setssbsy
  • wrssd
  • wrssq
  • wrussd
  • wrussq

Intel Fused Multiply-Add instructions (FMA)

84 instructions
  • vfmadd123pd
  • vfmadd123ps
  • vfmadd123sd
  • vfmadd123ss
  • vfmadd132sd — Fused Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfmadd132ss — Fused Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfmadd213sd — Fused Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfmadd213ss — Fused Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfmadd231sd — Fused Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfmadd231ss — Fused Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfmadd312pd
  • vfmadd312ps
  • vfmadd312sd
  • vfmadd312ss
  • vfmadd321pd
  • vfmadd321ps
  • vfmadd321sd
  • vfmadd321ss
  • vfmaddsub123pd
  • vfmaddsub123ps
  • vfmaddsub312pd
  • vfmaddsub312ps
  • vfmaddsub321pd
  • vfmaddsub321ps
  • vfmsub123pd
  • vfmsub123ps
  • vfmsub123sd
  • vfmsub123ss
  • vfmsub132sd — Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfmsub132ss — Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfmsub213sd — Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfmsub213ss — Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfmsub231sd — Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfmsub231ss — Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfmsub312pd
  • vfmsub312ps
  • vfmsub312sd
  • vfmsub312ss
  • vfmsub321pd
  • vfmsub321ps
  • vfmsub321sd
  • vfmsub321ss
  • vfmsubadd123pd
  • vfmsubadd123ps
  • vfmsubadd312pd
  • vfmsubadd312ps
  • vfmsubadd321pd
  • vfmsubadd321ps
  • vfnmadd123pd
  • vfnmadd123ps
  • vfnmadd123sd
  • vfnmadd123ss
  • vfnmadd132sd — Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfnmadd132ss — Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfnmadd213sd — Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfnmadd213ss — Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfnmadd231sd — Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Values
  • vfnmadd231ss — Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Values
  • vfnmadd312pd
  • vfnmadd312ps
  • vfnmadd312sd
  • vfnmadd312ss
  • vfnmadd321pd
  • vfnmadd321ps
  • vfnmadd321sd
  • vfnmadd321ss
  • vfnmsub123pd
  • vfnmsub123ps
  • vfnmsub123sd
  • vfnmsub123ss
  • vfnmsub132sd — Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfnmsub132ss — Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfnmsub213sd — Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfnmsub213ss — Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfnmsub231sd — Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Values
  • vfnmsub231ss — Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Values
  • vfnmsub312pd
  • vfnmsub312ps
  • vfnmsub312sd
  • vfnmsub312ss
  • vfnmsub321pd
  • vfnmsub321ps
  • vfnmsub321sd
  • vfnmsub321ss

Intel instruction extension based on pub number 319433-030 dated October 2017

4 instructions
  • vaesdec — Perform One Round of an AES Decryption Flow
  • vaesdeclast — Perform Last Round of an AES Decryption Flow
  • vaesenc — Perform One Round of an AES Encryption Flow
  • vaesenclast — Perform Last Round of an AES Encryption Flow

Intel Memory Protection Extensions (MPX)

7 instructions
  • bndcl
  • bndcn
  • bndcu
  • bndldx
  • bndmk
  • bndmov
  • bndstx

Intel memory protection keys for userspace (PKU aka PKEYs)

2 instructions
  • rdpkru
  • wrpkru

Intel SHA acceleration instructions

10 instructions
  • sha1msg1 — Perform an Intermediate Calculation for the Next Four SHA1 Message Doublewords
  • sha1msg2 — Perform a Final Calculation for the Next Four SHA1 Message Doublewords
  • sha1nexte — Calculate SHA1 State Variable E after Four Rounds
  • sha1rnds4 — Perform Four Rounds of SHA1 Operation
  • sha256msg1 — Perform an Intermediate Calculation for the Next Four SHA256 Message Doublewords
  • sha256msg2 — Perform a Final Calculation for the Next Four SHA256 Message Doublewords
  • sha256rnds2 — Perform Two Rounds of SHA256 Operation
  • vsha512msg1 — Perform an Intermediate Calculation for the Next Four SHA512 Message Quadwords
  • vsha512msg2 — Perform a Final Calculation for the Next Four SHA512 Message Quadwords
  • vsha512rnds2 — Perform Two Rounds of SHA512 Operation

Intel SMX

1 instruction
  • getsec

Intel Software Guard Extensions (SGX)

3 instructions
  • encls
  • enclu
  • enclv

Intel Transactional Synchronization Extensions (TSX)

5 instructions
  • prefetchwt1 — Prefetch Vector Data Into Caches with Intent to Write and T1 Hint
  • xabort
  • xbegin
  • xend
  • xtest

Interleaved flags arithmetic

2 instructions
  • adcx — Unsigned Integer Addition of Two Operands with Carry Flag
  • adox — Unsigned Integer Addition of Two Operands with Overflow Flag

Interrupts, system calls, and returns (extensions)

2 instructions
  • erets (FRED)
  • eretu (FRED)

Introduced in Deschutes but necessary for SSE support

4 instructions
  • fxrstor
  • fxrstor64
  • fxsave
  • fxsave64

Jumps (extensions)

1 instruction
  • jmpabs (APX)

Katmai Streaming SIMD instructions (SSE -- a.k.a. KNI, XMM, MMX2)

62 instructions
  • addps — Add Packed Single-Precision Floating-Point Values
  • addss — Add Scalar Single-Precision Floating-Point Values
  • andnps — Bitwise Logical AND NOT of Packed Single-Precision Floating-Point Values
  • andps — Bitwise Logical AND of Packed Single-Precision Floating-Point Values
  • cmpeqps
  • cmpeqss
  • cmpleps
  • cmpless
  • cmpltps
  • cmpltss
  • cmpneqps
  • cmpneqss
  • cmpnleps
  • cmpnless
  • cmpnltps
  • cmpnltss
  • cmpordps
  • cmpordss
  • cmpps — Compare Packed Single-Precision Floating-Point Values
  • cmpss — Compare Scalar Single-Precision Floating-Point Values
  • cmpunordps
  • cmpunordss
  • comiss — Compare Scalar Ordered Single-Precision Floating-Point Values and Set EFLAGS
  • cvtpi2ps — Convert Packed Dword Integers to Packed Single-Precision FP Values
  • cvtps2pi — Convert Packed Single-Precision FP Values to Packed Dword Integers
  • cvtsi2ss — Convert Dword Integer to Scalar Single-Precision FP Value
  • cvtss2si — Convert Scalar Single-Precision FP Value to Dword Integer
  • cvttps2pi — Convert with Truncation Packed Single-Precision FP Values to Packed Dword Integers
  • cvttss2si — Convert with Truncation Scalar Single-Precision FP Value to Dword Integer
  • divps — Divide Packed Single-Precision Floating-Point Values
  • divss — Divide Scalar Single-Precision Floating-Point Values
  • ldmxcsr — Load MXCSR Register
  • maxps — Return Maximum Packed Single-Precision Floating-Point Values
  • maxss — Return Maximum Scalar Single-Precision Floating-Point Value
  • minps — Return Minimum Packed Single-Precision Floating-Point Values
  • minss — Return Minimum Scalar Single-Precision Floating-Point Value
  • movaps — Move Aligned Packed Single-Precision Floating-Point Values
  • movhlps — Move Packed Single-Precision Floating-Point Values High to Low
  • movhps — Move High Packed Single-Precision Floating-Point Values
  • movlhps — Move Packed Single-Precision Floating-Point Values Low to High
  • movlps — Move Low Packed Single-Precision Floating-Point Values
  • movmskps — Extract Packed Single-Precision Floating-Point Sign Mask
  • movntps — Store Packed Single-Precision Floating-Point Values Using Non-Temporal Hint
  • movss — Move Scalar Single-Precision Floating-Point Values
  • movups — Move Unaligned Packed Single-Precision Floating-Point Values
  • mulps — Multiply Packed Single-Precision Floating-Point Values
  • mulss — Multiply Scalar Single-Precision Floating-Point Values
  • orps — Bitwise Logical OR of Single-Precision Floating-Point Values
  • rcpps — Compute Approximate Reciprocals of Packed Single-Precision Floating-Point Values
  • rcpss — Compute Approximate Reciprocal of Scalar Single-Precision Floating-Point Values
  • rsqrtps — Compute Reciprocals of Square Roots of Packed Single-Precision Floating-Point Values
  • rsqrtss — Compute Reciprocal of Square Root of Scalar Single-Precision Floating-Point Value
  • shufps — Shuffle Packed Single-Precision Floating-Point Values
  • sqrtps — Compute Square Roots of Packed Single-Precision Floating-Point Values
  • sqrtss — Compute Square Root of Scalar Single-Precision Floating-Point Value
  • stmxcsr — Store MXCSR Register State
  • subps — Subtract Packed Single-Precision Floating-Point Values
  • subss — Subtract Scalar Single-Precision Floating-Point Values
  • ucomiss — Unordered Compare Scalar Single-Precision Floating-Point Values and Set EFLAGS
  • unpckhps — Unpack and Interleave High Packed Single-Precision Floating-Point Values
  • unpcklps — Unpack and Interleave Low Packed Single-Precision Floating-Point Values
  • xorps — Bitwise Logical XOR for Single-Precision Floating-Point Values

Machine control and management instructions

20 instructions
  • bb0_reset
  • bb1_reset
  • clts
  • cpu_read
  • cpu_write
  • cpuid — CPU Identification
  • dmint
  • lmsw
  • rdm
  • rdmsr
  • rdmsrlist
  • smint
  • smintold
  • smsw
  • umov
  • urdmsr
  • uwrmsr
  • wrmsr
  • wrmsrlist
  • wrmsrns

Memory management and control

11 instructions
  • clflush — Flush Cache Line
  • clflushopt — Flush Cache Line Optimized
  • clwb — Cache Line Write Back
  • clzero — Zero-out 64-bit Cache Line
  • invd
  • invlpg
  • invlpga
  • invpcid
  • pcommit
  • wbinvd
  • wbnoinvd

MMX (SIMD using the x87 register file)

78 instructions
  • emms — Exit MMX State
  • movd — Move Doubleword
  • packssdw — Pack Doublewords into Words with Signed Saturation
  • packsswb — Pack Words into Bytes with Signed Saturation
  • packuswb — Pack Words into Bytes with Unsigned Saturation
  • paddb — Add Packed Byte Integers
  • paddd — Add Packed Doubleword Integers
  • paddsb — Add Packed Signed Byte Integers with Signed Saturation
  • paddsiw
  • paddsw — Add Packed Signed Word Integers with Signed Saturation
  • paddusb — Add Packed Unsigned Byte Integers with Unsigned Saturation
  • paddusw — Add Packed Unsigned Word Integers with Unsigned Saturation
  • paddw — Add Packed Word Integers
  • pand — Packed Bitwise Logical AND
  • pandn — Packed Bitwise Logical AND NOT
  • paveb
  • pavgusb — Average Packed Byte Integers
  • pcmpeqb — Compare Packed Byte Data for Equality
  • pcmpeqd — Compare Packed Doubleword Data for Equality
  • pcmpeqw — Compare Packed Word Data for Equality
  • pcmpgtb — Compare Packed Signed Byte Integers for Greater Than
  • pcmpgtd — Compare Packed Signed Doubleword Integers for Greater Than
  • pcmpgtw — Compare Packed Signed Word Integers for Greater Than
  • pdistib
  • pf2id — Packed Floating-Point to Integer Doubleword Converson
  • pfacc — Packed Floating-Point Accumulate
  • pfadd — Packed Floating-Point Add
  • pfcmpeq — Packed Floating-Point Compare for Equal
  • pfcmpge — Packed Floating-Point Compare for Greater or Equal
  • pfcmpgt — Packed Floating-Point Compare for Greater Than
  • pfmax — Packed Floating-Point Maximum
  • pfmin — Packed Floating-Point Minimum
  • pfmul — Packed Floating-Point Multiply
  • pfrcp — Packed Floating-Point Reciprocal Approximation
  • pfrcpit1 — Packed Floating-Point Reciprocal Iteration 1
  • pfrcpit2 — Packed Floating-Point Reciprocal Iteration 2
  • pfrsqit1 — Packed Floating-Point Reciprocal Square Root Iteration 1
  • pfrsqrt — Packed Floating-Point Reciprocal Square Root Approximation
  • pfsub — Packed Floating-Point Subtract
  • pfsubr — Packed Floating-Point Subtract Reverse
  • pi2fd — Packed Integer to Floating-Point Doubleword Conversion
  • pmachriw
  • pmaddwd — Multiply and Add Packed Signed Word Integers
  • pmagw
  • pmulhriw
  • pmulhrwa
  • pmulhrwc
  • pmulhw — Multiply Packed Signed Word Integers and Store High Result
  • pmullw — Multiply Packed Signed Word Integers and Store Low Result
  • pmvgezb
  • pmvlzb
  • pmvnzb
  • pmvzb
  • por — Packed Bitwise Logical OR
  • prefetch — Prefetch Data into Caches
  • prefetchw — Prefetch Data into Caches in Anticipation of a Write
  • pslld — Shift Packed Doubleword Data Left Logical
  • psllq — Shift Packed Quadword Data Left Logical
  • psllw — Shift Packed Word Data Left Logical
  • psrad — Shift Packed Doubleword Data Right Arithmetic
  • psraw — Shift Packed Word Data Right Arithmetic
  • psrld — Shift Packed Doubleword Data Right Logical
  • psrlq — Shift Packed Quadword Data Right Logical
  • psrlw — Shift Packed Word Data Right Logical
  • psubb — Subtract Packed Byte Integers
  • psubd — Subtract Packed Doubleword Integers
  • psubsb — Subtract Packed Signed Byte Integers with Signed Saturation
  • psubsiw
  • psubsw — Subtract Packed Signed Word Integers with Signed Saturation
  • psubusb — Subtract Packed Unsigned Byte Integers with Unsigned Saturation
  • psubusw — Subtract Packed Unsigned Word Integers with Unsigned Saturation
  • psubw — Subtract Packed Word Integers
  • punpckhbw — Unpack and Interleave High-Order Bytes into Words
  • punpckhdq — Unpack and Interleave High-Order Doublewords into Quadwords
  • punpckhwd — Unpack and Interleave High-Order Words into Doublewords
  • punpcklbw — Unpack and Interleave Low-Order Bytes into Words
  • punpckldq — Unpack and Interleave Low-Order Doublewords into Quadwords
  • punpcklwd — Unpack and Interleave Low-Order Words into Doublewords

MMX instructions

2 instructions
  • pxor — Packed Bitwise Logical Exclusive OR
  • skinit

Nehalem New Instructions (SSE4.2)

7 instructions
  • crc32 — Accumulate CRC32 Value
  • pcmpestri — Packed Compare Explicit Length Strings, Return Index
  • pcmpestrm — Packed Compare Explicit Length Strings, Return Mask
  • pcmpgtq — Compare Packed Data for Greater Than
  • pcmpistri — Packed Compare Implicit Length Strings, Return Index
  • pcmpistrm — Packed Compare Implicit Length Strings, Return Mask
  • popcnt — Count of Number of Bits Set to 1

New MMX instructions introduced in Katmai

12 instructions
  • maskmovq — Store Selected Bytes of Quadword
  • movntq — Store of Quadword Using Non-Temporal Hint
  • pavgb — Average Packed Byte Integers
  • pavgw — Average Packed Word Integers
  • pmaxsw — Maximum of Packed Signed Word Integers
  • pmaxub — Maximum of Packed Unsigned Byte Integers
  • pminsw — Minimum of Packed Signed Word Integers
  • pminub — Minimum of Packed Unsigned Byte Integers
  • pmovmskb — Move Byte Mask
  • pmulhuw — Multiply Packed Unsigned Word Integers and Store High Result
  • psadbw — Compute Sum of Absolute Differences
  • pshufw — Shuffle Packed Words

Other basic integer arithmetic (extensions)

1 instruction
  • mulx — Unsigned Multiply Without Affecting Flags (BMI2)

Penryn New Instructions (SSE4.1)

49 instructions
  • blendpd — Blend Packed Double Precision Floating-Point Values
  • blendps — Blend Packed Single Precision Floating-Point Values
  • blendvpd — Variable Blend Packed Double Precision Floating-Point Values
  • blendvps — Variable Blend Packed Single Precision Floating-Point Values
  • dppd — Dot Product of Packed Double Precision Floating-Point Values
  • dpps — Dot Product of Packed Single Precision Floating-Point Values
  • extractps — Extract Packed Single Precision Floating-Point Value
  • insertps — Insert Packed Single Precision Floating-Point Value
  • movntdqa — Load Double Quadword Non-Temporal Aligned Hint
  • mpsadbw — Compute Multiple Packed Sums of Absolute Difference
  • packusdw — Pack Doublewords into Words with Unsigned Saturation
  • pblendvb — Variable Blend Packed Bytes
  • pblendw — Blend Packed Words
  • pcmpeqq — Compare Packed Quadword Data for Equality
  • pextrb — Extract Byte
  • pextrd — Extract Doubleword
  • pextrq — Extract Quadword
  • pextrw — Extract Word
  • phminposuw — Packed Horizontal Minimum of Unsigned Word Integers
  • pinsrb — Insert Byte
  • pinsrd — Insert Doubleword
  • pinsrq — Insert Quadword
  • pmaxsb — Maximum of Packed Signed Byte Integers
  • pmaxsd — Maximum of Packed Signed Doubleword Integers
  • pmaxud — Maximum of Packed Unsigned Doubleword Integers
  • pmaxuw — Maximum of Packed Unsigned Word Integers
  • pminsb — Minimum of Packed Signed Byte Integers
  • pminsd — Minimum of Packed Signed Doubleword Integers
  • pminud — Minimum of Packed Unsigned Doubleword Integers
  • pminuw — Minimum of Packed Unsigned Word Integers
  • pmovsxbd — Move Packed Byte Integers to Doubleword Integers with Sign Extension
  • pmovsxbq — Move Packed Byte Integers to Quadword Integers with Sign Extension
  • pmovsxbw — Move Packed Byte Integers to Word Integers with Sign Extension
  • pmovsxdq — Move Packed Doubleword Integers to Quadword Integers with Sign Extension
  • pmovsxwd — Move Packed Word Integers to Doubleword Integers with Sign Extension
  • pmovsxwq — Move Packed Word Integers to Quadword Integers with Sign Extension
  • pmovzxbd — Move Packed Byte Integers to Doubleword Integers with Zero Extension
  • pmovzxbq — Move Packed Byte Integers to Quadword Integers with Zero Extension
  • pmovzxbw — Move Packed Byte Integers to Word Integers with Zero Extension
  • pmovzxdq — Move Packed Doubleword Integers to Quadword Integers with Zero Extension
  • pmovzxwd — Move Packed Word Integers to Doubleword Integers with Zero Extension
  • pmovzxwq — Move Packed Word Integers to Quadword Integers with Zero Extension
  • pmuldq — Multiply Packed Signed Doubleword Integers and Store Quadword Result
  • pmulld — Multiply Packed Signed Doubleword Integers and Store Low Result
  • ptest — Packed Logical Compare
  • roundpd — Round Packed Double Precision Floating-Point Values
  • roundps — Round Packed Single Precision Floating-Point Values
  • roundsd — Round Scalar Double Precision Floating-Point Values
  • roundss — Round Scalar Single Precision Floating-Point Values

Permanently undefined instructions

67 instructions
  • ccmp
  • ccmpa
  • ccmpae
  • ccmpb
  • ccmpbe
  • ccmpc
  • ccmpe
  • ccmpf
  • ccmpg
  • ccmpge
  • ccmpl
  • ccmple
  • ccmpna
  • ccmpnae
  • ccmpnb
  • ccmpnbe
  • ccmpnc
  • ccmpne
  • ccmpng
  • ccmpnge
  • ccmpnl
  • ccmpnle
  • ccmpno
  • ccmpns
  • ccmpnz
  • ccmpo
  • ccmps
  • ccmpt
  • ccmpz
  • ctest
  • ctesta
  • ctestae
  • ctestb
  • ctestbe
  • ctestc
  • cteste
  • ctestf
  • ctestg
  • ctestge
  • ctestl
  • ctestle
  • ctestna
  • ctestnae
  • ctestnb
  • ctestnbe
  • ctestnc
  • ctestne
  • ctestng
  • ctestnge
  • ctestnl
  • ctestnle
  • ctestno
  • ctestns
  • ctestnz
  • ctesto
  • ctests
  • ctestt
  • ctestz
  • fwait
  • ud0
  • ud1
  • ud2 — Undefined Instruction
  • ud2a
  • ud2b
  • udb
  • xlat
  • xlatb — Table Look-up Translation

Power management

12 instructions
  • hlt
  • monitor — Monitor a Linear Address Range
  • monitord
  • monitorq
  • monitorw
  • monitorx — Monitor a Linear Address Range with Timeout
  • mwait — Monitor Wait
  • mwaitx — Monitor Wait with Timeout
  • pause — Spin Loop Hint
  • tpause — Timed PAUSE
  • umonitor — User mode Monitor a Linear Address Range
  • umwait — User mode Monitor Wait

Prescott New Instructions (SSE3)

10 instructions
  • addsubpd — Packed Double-FP Add/Subtract
  • addsubps — Packed Single-FP Add/Subtract
  • haddpd — Packed Double-FP Horizontal Add
  • haddps — Packed Single-FP Horizontal Add
  • hsubpd — Packed Double-FP Horizontal Subtract
  • hsubps — Packed Single-FP Horizontal Subtract
  • lddqu — Load Unaligned Integer 128 Bits
  • movddup — Move One Double-FP and Duplicate
  • movshdup — Move Packed Single-FP High and Duplicate
  • movsldup — Move Packed Single-FP Low and Duplicate

Processor trace write

1 instruction
  • ptwrite

RAO-INT weakly ordered atomic operations

4 instructions
  • aadd — Atomically ADD
  • aand — Atomically AND
  • aor — Atomically OR
  • axor — Atomically XOR

S3M hash instructions

3 instructions
  • vsm3msg1 — Perform Initial Calculation for the Next Four SM3 Message Words
  • vsm3msg2 — Perform Final Calculation for the Next Four SM3 Message Words
  • vsm3rnds2 — Perform Two Rounds of SM3 Operation

Segment handling instructions

26 instructions
  • arpl
  • lar
  • lds
  • les
  • lfs
  • lgdt
  • lgs
  • lidt
  • lkgs
  • lldt
  • loadall
  • loadall286
  • lsl
  • lss
  • ltr
  • rdfsbase — ReaD FS segment BASE
  • rdgsbase — ReaD GS segment BASE
  • sgdt
  • sidt
  • sldt
  • str
  • swapgs
  • verr
  • verw
  • wrfsbase — WRite FS segment BASE
  • wrgsbase — WRite GS segment BASE

SEV-SNP AMD instructions

3 instructions
  • pvalidate
  • rmpadjust
  • vmgexit

SM4 hash instructions

2 instructions
  • vsm4key4 — Perform Four Rounds of SM4 Key Expansion
  • vsm4rnds4 — Performs Four Rounds of SM4 Encryption

Special reads: timestamp, CPU number, performance counters, randomness

6 instructions
  • rdpid — Read Processor ID
  • rdpmc — Read Performance-Monitoring Counter
  • rdrand — Read Random Number
  • rdseed — Read Random SEED
  • rdtsc — Read Time-Stamp Counter
  • rdtscp — Read Time-Stamp Counter and Processor ID

Stack operations (extensions)

6 instructions
  • pop2 (APX)
  • pop2p (APX)
  • popp (APX)
  • push2 (APX)
  • push2p (APX)
  • pushp (APX)

Synchronization and fencing

4 instructions
  • lfence — Load Fence
  • mfence — Memory Fence
  • serialize — Serialize Instruction Execution
  • sfence — Store Fence

System management mode

9 instructions
  • rdshr
  • rsdc
  • rsldt
  • rsm
  • rsts
  • svdc
  • svldt
  • svts
  • wrshr

Systematic names for the hinting nop instructions

65 instructions
  • hint_nop
  • hint_nop0
  • hint_nop1
  • hint_nop10
  • hint_nop11
  • hint_nop12
  • hint_nop13
  • hint_nop14
  • hint_nop15
  • hint_nop16
  • hint_nop17
  • hint_nop18
  • hint_nop19
  • hint_nop2
  • hint_nop20
  • hint_nop21
  • hint_nop22
  • hint_nop23
  • hint_nop24
  • hint_nop25
  • hint_nop26
  • hint_nop27
  • hint_nop28
  • hint_nop29
  • hint_nop3
  • hint_nop30
  • hint_nop31
  • hint_nop32
  • hint_nop33
  • hint_nop34
  • hint_nop35
  • hint_nop36
  • hint_nop37
  • hint_nop38
  • hint_nop39
  • hint_nop4
  • hint_nop40
  • hint_nop41
  • hint_nop42
  • hint_nop43
  • hint_nop44
  • hint_nop45
  • hint_nop46
  • hint_nop47
  • hint_nop48
  • hint_nop49
  • hint_nop5
  • hint_nop50
  • hint_nop51
  • hint_nop52
  • hint_nop53
  • hint_nop54
  • hint_nop55
  • hint_nop56
  • hint_nop57
  • hint_nop58
  • hint_nop59
  • hint_nop6
  • hint_nop60
  • hint_nop61
  • hint_nop62
  • hint_nop63
  • hint_nop7
  • hint_nop8
  • hint_nop9

Tejas New Instructions (SSSE3)

16 instructions
  • pabsb — Packed Absolute Value of Byte Integers
  • pabsd — Packed Absolute Value of Doubleword Integers
  • pabsw — Packed Absolute Value of Word Integers
  • palignr — Packed Align Right
  • phaddd — Packed Horizontal Add Doubleword Integer
  • phaddsw — Packed Horizontal Add Signed Word Integers with Signed Saturation
  • phaddw — Packed Horizontal Add Word Integers
  • phsubd — Packed Horizontal Subtract Doubleword Integers
  • phsubsw — Packed Horizontal Subtract Signed Word Integers with Signed Saturation
  • phsubw — Packed Horizontal Subtract Word Integers
  • pmaddubsw — Multiply and Add Packed Signed and Unsigned Byte Integers
  • pmulhrsw — Packed Multiply Signed Word Integers and Store High Result with Round and Scale
  • pshufb — Packed Shuffle Bytes
  • psignb — Packed Sign of Byte Integers
  • psignd — Packed Sign of Doubleword Integers
  • psignw — Packed Sign of Word Integers

The basic shift and rotate operations (extensions)

6 instructions
  • rolx (BMI2)
  • rorx — Rotate Right Logical Without Affecting Flags (BMI2)
  • salx (BMI2)
  • sarx — Arithmetic Shift Right Without Affecting Flags (BMI2)
  • shlx — Logical Shift Left Without Affecting Flags (BMI2)
  • shrx — Logical Shift Right Without Affecting Flags (BMI2)

User interrupts

5 instructions
  • clui
  • senduipi
  • stui
  • testui
  • uiret

VIA (Centaur) security instructions

9 instructions
  • montmul
  • xcryptcbc
  • xcryptcfb
  • xcryptctr
  • xcryptecb
  • xcryptofb
  • xsha1
  • xsha256
  • xstore

VMX/SVM Instructions

17 instructions
  • clgi
  • stgi
  • vmcall
  • vmclear
  • vmfunc
  • vmlaunch
  • vmload
  • vmmcall
  • vmptrld
  • vmptrst
  • vmread
  • vmresume
  • vmrun
  • vmsave
  • vmwrite
  • vmxoff
  • vmxon

Willamette MMX instructions (SSE2 SIMD Integer Instructions)

16 instructions
  • movdq2q — Move Quadword from XMM to MMX Technology Register
  • movdqa — Move Aligned Double Quadword
  • movdqu — Move Unaligned Double Quadword
  • movq — Move Quadword
  • movq2dq — Move Quadword from MMX Technology to XMM Register
  • paddq — Add Packed Quadword Integers
  • pinsrw — Insert Word
  • pmuludq — Multiply Packed Unsigned Doubleword Integers
  • pshufd — Shuffle Packed Doublewords
  • pshufhw — Shuffle Packed High Words
  • pshuflw — Shuffle Packed Low Words
  • pslldq — Shift Packed Double Quadword Left Logical
  • psrldq — Shift Packed Double Quadword Right Logical
  • psubq — Subtract Packed Quadword Integers
  • punpckhqdq — Unpack and Interleave High-Order Quadwords into Double Quadwords
  • punpcklqdq — Unpack and Interleave Low-Order Quadwords into Double Quadwords

Willamette SSE2 Cacheability Instructions

4 instructions
  • maskmovdqu — Store Selected Bytes of Double Quadword
  • movntdq — Store Double Quadword Using Non-Temporal Hint
  • movnti — Store Doubleword Using Non-Temporal Hint
  • movntpd — Store Packed Double-Precision Floating-Point Values Using Non-Temporal Hint

Willamette Streaming SIMD instructions (SSE2)

62 instructions
  • addpd — Add Packed Double-Precision Floating-Point Values
  • addsd — Add Scalar Double-Precision Floating-Point Values
  • andnpd — Bitwise Logical AND NOT of Packed Double-Precision Floating-Point Values
  • andpd — Bitwise Logical AND of Packed Double-Precision Floating-Point Values
  • cmpeqpd
  • cmpeqsd
  • cmplepd
  • cmplesd
  • cmpltpd
  • cmpltsd
  • cmpneqpd
  • cmpneqsd
  • cmpnlepd
  • cmpnlesd
  • cmpnltpd
  • cmpnltsd
  • cmpordpd
  • cmpordsd
  • cmppd — Compare Packed Double-Precision Floating-Point Values
  • cmpunordpd
  • cmpunordsd
  • comisd — Compare Scalar Ordered Double-Precision Floating-Point Values and Set EFLAGS
  • cvtdq2pd — Convert Packed Dword Integers to Packed Double-Precision FP Values
  • cvtdq2ps — Convert Packed Dword Integers to Packed Single-Precision FP Values
  • cvtpd2dq — Convert Packed Double-Precision FP Values to Packed Dword Integers
  • cvtpd2pi — Convert Packed Double-Precision FP Values to Packed Dword Integers
  • cvtpd2ps — Convert Packed Double-Precision FP Values to Packed Single-Precision FP Values
  • cvtpi2pd — Convert Packed Dword Integers to Packed Double-Precision FP Values
  • cvtps2dq — Convert Packed Single-Precision FP Values to Packed Dword Integers
  • cvtps2pd — Convert Packed Single-Precision FP Values to Packed Double-Precision FP Values
  • cvtsd2si — Convert Scalar Double-Precision FP Value to Integer
  • cvtsd2ss — Convert Scalar Double-Precision FP Value to Scalar Single-Precision FP Value
  • cvtsi2sd — Convert Dword Integer to Scalar Double-Precision FP Value
  • cvtss2sd — Convert Scalar Single-Precision FP Value to Scalar Double-Precision FP Value
  • cvttpd2dq — Convert with Truncation Packed Double-Precision FP Values to Packed Dword Integers
  • cvttpd2pi — Convert with Truncation Packed Double-Precision FP Values to Packed Dword Integers
  • cvttps2dq — Convert with Truncation Packed Single-Precision FP Values to Packed Dword Integers
  • cvttsd2si — Convert with Truncation Scalar Double-Precision FP Value to Signed Integer
  • divpd — Divide Packed Double-Precision Floating-Point Values
  • divsd — Divide Scalar Double-Precision Floating-Point Values
  • maxpd — Return Maximum Packed Double-Precision Floating-Point Values
  • maxsd — Return Maximum Scalar Double-Precision Floating-Point Value
  • minpd — Return Minimum Packed Double-Precision Floating-Point Values
  • minsd — Return Minimum Scalar Double-Precision Floating-Point Value
  • movapd — Move Aligned Packed Double-Precision Floating-Point Values
  • movhpd — Move High Packed Double-Precision Floating-Point Value
  • movlpd — Move Low Packed Double-Precision Floating-Point Value
  • movmskpd — Extract Packed Double-Precision Floating-Point Sign Mask
  • movsd — Move Scalar Double-Precision Floating-Point Value
  • movupd — Move Unaligned Packed Double-Precision Floating-Point Values
  • mulpd — Multiply Packed Double-Precision Floating-Point Values
  • mulsd — Multiply Scalar Double-Precision Floating-Point Values
  • orpd — Bitwise Logical OR of Double-Precision Floating-Point Values
  • shufpd — Shuffle Packed Double-Precision Floating-Point Values
  • sqrtpd — Compute Square Roots of Packed Double-Precision Floating-Point Values
  • sqrtsd — Compute Square Root of Scalar Double-Precision Floating-Point Value
  • subpd — Subtract Packed Double-Precision Floating-Point Values
  • subsd — Subtract Scalar Double-Precision Floating-Point Values
  • ucomisd — Unordered Compare Scalar Double-Precision Floating-Point Values and Set EFLAGS
  • unpckhpd — Unpack and Interleave High Packed Double-Precision Floating-Point Values
  • unpcklpd — Unpack and Interleave Low Packed Double-Precision Floating-Point Values
  • xorpd — Bitwise Logical XOR for Double-Precision Floating-Point Values

x87 floating point

99 instructions
  • f2xm1
  • fabs
  • fadd
  • faddp
  • fbld
  • fbstp
  • fchs
  • fclex
  • fcmovb
  • fcmovbe
  • fcmove
  • fcmovnb
  • fcmovnbe
  • fcmovne
  • fcmovnu
  • fcmovu
  • fcom
  • fcomi
  • fcomip
  • fcomp
  • fcompp
  • fcos
  • fdecstp
  • fdisi
  • fdiv
  • fdivp
  • fdivr
  • fdivrp
  • femms — Fast Exit Multimedia State
  • feni
  • ffree
  • ffreep
  • fiadd
  • ficom
  • ficomp
  • fidiv
  • fidivr
  • fild
  • fimul
  • fincstp
  • finit
  • fist
  • fistp
  • fisttp
  • fisub
  • fisubr
  • fld
  • fld1
  • fldcw
  • fldenv
  • fldl2e
  • fldl2t
  • fldlg2
  • fldln2
  • fldpi
  • fldz
  • fmul
  • fmulp
  • fnclex
  • fndisi
  • fneni
  • fninit
  • fnop
  • fnsave
  • fnstcw
  • fnstenv
  • fnstsw
  • fpatan
  • fprem
  • fprem1
  • fptan
  • frndint
  • frstor
  • fsave
  • fscale
  • fsetpm
  • fsin
  • fsincos
  • fsqrt
  • fst
  • fstcw
  • fstenv
  • fstp
  • fstsw
  • fsub
  • fsubp
  • fsubr
  • fsubrp
  • ftst
  • fucom
  • fucomi
  • fucomip
  • fucomp
  • fucompp
  • fxam
  • fxch
  • fxtract
  • fyl2x
  • fyl2xp1

XSAVE group (AVX and extended state)

14 instructions
  • xgetbv — Get Value of Extended Control Register
  • xrstor
  • xrstor64
  • xrstors
  • xrstors64
  • xsave
  • xsave64
  • xsavec
  • xsavec64
  • xsaveopt
  • xsaveopt64
  • xsaves
  • xsaves64
  • xsetbv