x86-64 extensions
Instructions without a page
Everything else the assembler accepts, under the heading NASM files it under. These have no page of their own: the only description of them we are free to publish is the one line below, and blink implements a part of them.
AMD Enhanced 3DNow! (Athlon) instructions
5 instructionspf2iw— Packed Floating-Point to Integer Word Conversionpfnacc— Packed Floating-Point Negative Accumulatepfpnacc— Packed Floating-Point Positive-Negative Accumulatepi2fw— Packed Integer to Floating-Point Word Conversionpswapd— Packed Swap Doubleword
AMD Lightweight Profiling (LWP) instructions
4 instructionsllwpcblwpinslwpvalslwpcb
AMD SSE4A
4 instructionsextrq— Extract Fieldinsertq— Insert Fieldmovntsd— Store Scalar Double-Precision Floating-Point Values Using Non-Temporal Hintmovntss— Store Scalar Single-Precision Floating-Point Values Using Non-Temporal Hint
AMD XOP and FMA4 instructions (SSE5)
73 instructionsvfmaddpd— Fused Multiply-Add of Packed Double-Precision Floating-Point Valuesvfmaddps— Fused Multiply-Add of Packed Single-Precision Floating-Point Valuesvfmaddsd— Fused Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfmaddss— Fused Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfmaddsubpd— Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Valuesvfmaddsubps— Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Valuesvfmsubaddpd— Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Valuesvfmsubaddps— Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Valuesvfmsubpd— Fused Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfmsubps— Fused Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfmsubsd— Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfmsubss— Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfnmaddpd— Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Valuesvfnmaddps— Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Valuesvfnmaddsd— Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfnmaddss— Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfnmsubpd— Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfnmsubps— Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfnmsubsd— Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfnmsubss— Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfrczpd— Extract Fraction Packed Double-Precision Floating-Pointvfrczps— Extract Fraction Packed Single-Precision Floating-Pointvfrczsd— Extract Fraction Scalar Double-Precision Floating-Pointvfrczss— Extract Fraction Scalar Single-Precision Floating Pointvpcmov— Packed Conditional Movevpcomb— Compare Packed Signed Byte Integersvpcomd— Compare Packed Signed Doubleword Integersvpcomq— Compare Packed Signed Quadword Integersvpcomub— Compare Packed Unsigned Byte Integersvpcomud— Compare Packed Unsigned Doubleword Integersvpcomuq— Compare Packed Unsigned Quadword Integersvpcomuw— Compare Packed Unsigned Word Integersvpcomw— Compare Packed Signed Word Integersvphaddbd— Packed Horizontal Add Signed Byte to Signed Doublewordvphaddbq— Packed Horizontal Add Signed Byte to Signed Quadwordvphaddbw— Packed Horizontal Add Signed Byte to Signed Wordvphadddq— Packed Horizontal Add Signed Doubleword to Signed Quadwordvphaddubd— Packed Horizontal Add Unsigned Byte to Doublewordvphaddubq— Packed Horizontal Add Unsigned Byte to Quadwordvphaddubw— Packed Horizontal Add Unsigned Byte to Wordvphaddudq— Packed Horizontal Add Unsigned Doubleword to Quadwordvphadduwd— Packed Horizontal Add Unsigned Word to Doublewordvphadduwq— Packed Horizontal Add Unsigned Word to Quadwordvphaddwd— Packed Horizontal Add Signed Word to Signed Doublewordvphaddwq— Packed Horizontal Add Signed Word to Signed Quadwordvphsubbw— Packed Horizontal Subtract Signed Byte to Signed Wordvphsubdq— Packed Horizontal Subtract Signed Doubleword to Signed Quadwordvphsubwd— Packed Horizontal Subtract Signed Word to Signed Doublewordvpmacsdd— Packed Multiply Accumulate Signed Doubleword to Signed Doublewordvpmacsdqh— Packed Multiply Accumulate Signed High Doubleword to Signed Quadwordvpmacsdql— Packed Multiply Accumulate Signed Low Doubleword to Signed Quadwordvpmacssdd— Packed Multiply Accumulate with Saturation Signed Doubleword to Signed Doublewordvpmacssdqh— Packed Multiply Accumulate with Saturation Signed High Doubleword to Signed Quadwordvpmacssdql— Packed Multiply Accumulate with Saturation Signed Low Doubleword to Signed Quadwordvpmacsswd— Packed Multiply Accumulate with Saturation Signed Word to Signed Doublewordvpmacssww— Packed Multiply Accumulate with Saturation Signed Word to Signed Wordvpmacswd— Packed Multiply Accumulate Signed Word to Signed Doublewordvpmacsww— Packed Multiply Accumulate Signed Word to Signed Wordvpmadcsswd— Packed Multiply Add Accumulate with Saturation Signed Word to Signed Doublewordvpmadcswd— Packed Multiply Add Accumulate Signed Word to Signed Doublewordvpperm— Packed Permute Bytesvprotb— Packed Rotate Bytesvprotd— Packed Rotate Doublewordsvprotq— Packed Rotate Quadwordsvprotw— Packed Rotate Wordsvpshab— Packed Shift Arithmetic Bytesvpshad— Packed Shift Arithmetic Doublewordsvpshaq— Packed Shift Arithmetic Quadwordsvpshaw— Packed Shift Arithmetic Wordsvpshlb— Packed Shift Logical Bytesvpshld— Packed Shift Logical Doublewordsvpshlq— Packed Shift Logical Quadwordsvpshlw— Packed Shift Logical Words
AMD XOP bit operations
9 instructionsblcfill— Fill From Lowest Clear Bitblci— Isolate Lowest Clear Bitblcic— Isolate Lowest Set Bit and Complementblcmsk— Mask From Lowest Clear Bitblcs— Set Lowest Clear Bitblsfill— Fill From Lowest Set Bitblsic— Isolate Lowest Set Bit and Complementt1mskc— Inverse Mask From Trailing Onestzmsk— Mask From Trailing Zeros
AVX no exception conversions
7 instructionsvbcstnebf162ps— Load BF16 Element and Convert to FP32 Element With Broadcasvbcstnebf16psvbcstnesh2ps— Load FP16 Element and Convert to FP32 Element with Broadcastvcvtneebf162ps— Convert Even Elements of Packed BF16 Values to FP32 Valuesvcvtneeph2ps— Convert Even Elements of Packed FP16 Values to FP32 Valuesvcvtneobf162ps— Convert Odd Elements of Packed BF16 Values to FP32 Valuesvcvtneoph2ps— Convert Odd Elements of Packed FP16 Values to FP32 Values
AVX-512 instructions
605 instructionsvaddpd— Add Packed Double-Precision Floating-Point Valuesvaddps— Add Packed Single-Precision Floating-Point Valuesvalignd— Align Doubleword Vectorsvalignq— Align Quadword Vectorsvandnpd— Bitwise Logical AND NOT of Packed Double-Precision Floating-Point Valuesvandnps— Bitwise Logical AND NOT of Packed Single-Precision Floating-Point Valuesvandpd— Bitwise Logical AND of Packed Double-Precision Floating-Point Valuesvandps— Bitwise Logical AND of Packed Single-Precision Floating-Point Valuesvblendmpd— Blend Packed Double-Precision Floating-Point Vectors Using an OpMask Controlvblendmps— Blend Packed Single-Precision Floating-Point Vectors Using an OpMask Controlvbroadcastf32x2— Broadcast Two Single-Precision Floating-Point Elementsvbroadcastf32x4— Broadcast Four Single-Precision Floating-Point Elementsvbroadcastf32x8— Broadcast Eight Single-Precision Floating-Point Elementsvbroadcastf64x2— Broadcast Two Double-Precision Floating-Point Elementsvbroadcastf64x4— Broadcast Four Double-Precision Floating-Point Elementsvbroadcasti32x2— Broadcast Two Doubleword Elementsvbroadcasti32x4— Broadcast Four Doubleword Elementsvbroadcasti32x8— Broadcast Eight Doubleword Elementsvbroadcasti64x2— Broadcast Two Quadword Elementsvbroadcasti64x4— Broadcast Four Quadword Elementsvbroadcastsd— Broadcast Double-Precision Floating-Point Elementvbroadcastss— Broadcast Single-Precision Floating-Point Elementvcmpeq_oqpdvcmpeq_oqpsvcmpeq_oqsdvcmpeq_oqssvcmpeq_uqpdvcmpeq_uqpsvcmpeq_uspdvcmpeq_uspsvcmpeqpdvcmpeqpsvcmpfalse_oqpdvcmpfalse_oqpsvcmpfalse_ospdvcmpfalse_ospsvcmpfalsepdvcmpfalsepsvcmpge_oqpdvcmpge_oqpsvcmpge_ospdvcmpge_ospsvcmpgepdvcmpgepsvcmpgt_oqpdvcmpgt_oqpsvcmpgt_ospdvcmpgt_ospsvcmpgtpdvcmpgtpsvcmple_oqpdvcmple_oqpsvcmple_ospdvcmple_ospsvcmplepdvcmplepsvcmplt_oqpdvcmplt_oqpsvcmplt_ospdvcmplt_ospsvcmpltpdvcmpltpsvcmpneq_oqpdvcmpneq_oqpsvcmpneq_ospdvcmpneq_ospsvcmpneq_uqpdvcmpneq_uqpsvcmpneq_uspdvcmpneq_uspsvcmpneqpdvcmpneqpsvcmpnge_uqpdvcmpnge_uqpsvcmpnge_uspdvcmpnge_uspsvcmpngepdvcmpngepsvcmpngt_uqpdvcmpngt_uqpsvcmpngt_uspdvcmpngt_uspsvcmpngtpdvcmpngtpsvcmpnle_uqpdvcmpnle_uqpsvcmpnle_uspdvcmpnle_uspsvcmpnlepdvcmpnlepsvcmpnlt_uqpdvcmpnlt_uqpsvcmpnlt_uspdvcmpnlt_uspsvcmpnltpdvcmpnltpsvcmpord_qpdvcmpord_qpsvcmpord_spdvcmpord_spsvcmpordpdvcmpordpsvcmppd— Compare Packed Double-Precision Floating-Point Valuesvcmpps— Compare Packed Single-Precision Floating-Point Valuesvcmptrue_uqpdvcmptrue_uqpsvcmptrue_uspdvcmptrue_uspsvcmptruepdvcmptruepsvcmpunord_qpdvcmpunord_qpsvcmpunord_spdvcmpunord_spsvcmpunordpdvcmpunordpsvcompresspd— Store Sparse Packed Double-Precision Floating-Point Values into Dense Memory/Registervcompressps— Store Sparse Packed Single-Precision Floating-Point Values into Dense Memory/Registervcvtdq2pd— Convert Packed Dword Integers to Packed Double-Precision FP Valuesvcvtdq2ps— Convert Packed Dword Integers to Packed Single-Precision FP Valuesvcvtpd2qq— Convert Packed Double-Precision Floating-Point Values to Packed Quadword Integersvcvtpd2udq— Convert Packed Double-Precision Floating-Point Values to Packed Unsigned Doubleword Integersvcvtpd2uqq— Convert Packed Double-Precision Floating-Point Values to Packed Unsigned Quadword Integersvcvtps2dq— Convert Packed Single-Precision FP Values to Packed Dword Integersvcvtps2pd— Convert Packed Single-Precision FP Values to Packed Double-Precision FP Valuesvcvtps2qq— Convert Packed Single Precision Floating-Point Values to Packed Singed Quadword Integer Valuesvcvtps2udq— Convert Packed Single-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Valuesvcvtps2uqq— Convert Packed Single Precision Floating-Point Values to Packed Unsigned Quadword Integer Valuesvcvtqq2pd— Convert Packed Quadword Integers to Packed Double-Precision Floating-Point Valuesvcvtqq2ps— Convert Packed Quadword Integers to Packed Single-Precision Floating-Point Valuesvcvtsd2usi— Convert Scalar Double-Precision Floating-Point Value to Unsigned Doubleword Integervcvtss2usi— Convert Scalar Single-Precision Floating-Point Value to Unsigned Doubleword Integervcvttpd2qq— Convert with Truncation Packed Double-Precision Floating-Point Values to Packed Quadword Integersvcvttpd2udq— Convert with Truncation Packed Double-Precision Floating-Point Values to Packed Unsigned Doubleword Integersvcvttpd2uqq— Convert with Truncation Packed Double-Precision Floating-Point Values to Packed Unsigned Quadword Integersvcvttps2dq— Convert with Truncation Packed Single-Precision FP Values to Packed Dword Integersvcvttps2qq— Convert with Truncation Packed Single Precision Floating-Point Values to Packed Singed Quadword Integer Valuesvcvttps2udq— Convert with Truncation Packed Single-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Valuesvcvttps2uqq— Convert with Truncation Packed Single Precision Floating-Point Values to Packed Unsigned Quadword Integer Valuesvcvttsd2usi— Convert with Truncation Scalar Double-Precision Floating-Point Value to Unsigned Integervcvttss2usi— Convert with Truncation Scalar Single-Precision Floating-Point Value to Unsigned Integervcvtudq2pd— Convert Packed Unsigned Doubleword Integers to Packed Double-Precision Floating-Point Valuesvcvtudq2ps— Convert Packed Unsigned Doubleword Integers to Packed Single-Precision Floating-Point Valuesvcvtuqq2pd— Convert Packed Unsigned Quadword Integers to Packed Double-Precision Floating-Point Valuesvcvtuqq2ps— Convert Packed Unsigned Quadword Integers to Packed Single-Precision Floating-Point Valuesvcvtusi2sd— Convert Unsigned Integer to Scalar Double-Precision Floating-Point Valuevcvtusi2ss— Convert Unsigned Integer to Scalar Single-Precision Floating-Point Valuevdbpsadbw— Double Block Packed Sum-Absolute-Differences on Unsigned Bytesvdivpd— Divide Packed Double-Precision Floating-Point Valuesvdivps— Divide Packed Single-Precision Floating-Point Valuesvexp2pd— Approximation to the Exponential 2^x of Packed Double-Precision Floating-Point Values with Less Than 2^-23 Relative Errorvexp2ps— Approximation to the Exponential 2^x of Packed Single-Precision Floating-Point Values with Less Than 2^-23 Relative Errorvexpandpd— Load Sparse Packed Double-Precision Floating-Point Values from Dense Memoryvexpandps— Load Sparse Packed Single-Precision Floating-Point Values from Dense Memoryvextractf32x4— Extract 128 Bits of Packed Single-Precision Floating-Point Valuesvextractf32x8— Extract 256 Bits of Packed Single-Precision Floating-Point Valuesvextractf64x2— Extract 128 Bits of Packed Double-Precision Floating-Point Valuesvextractf64x4— Extract 256 Bits of Packed Double-Precision Floating-Point Valuesvextracti32x4— Extract 128 Bits of Packed Doubleword Integer Valuesvextracti32x8— Extract 256 Bits of Packed Doubleword Integer Valuesvextracti64x2— Extract 128 Bits of Packed Quadword Integer Valuesvextracti64x4— Extract 256 Bits of Packed Quadword Integer Valuesvextractps— Extract Packed Single Precision Floating-Point Valuevfixupimmpd— Fix Up Special Packed Double-Precision Floating-Point Valuesvfixupimmps— Fix Up Special Packed Single-Precision Floating-Point Valuesvfixupimmsd— Fix Up Special Scalar Double-Precision Floating-Point Valuevfixupimmss— Fix Up Special Scalar Single-Precision Floating-Point Valuevfmadd132pd— Fused Multiply-Add of Packed Double-Precision Floating-Point Valuesvfmadd132ps— Fused Multiply-Add of Packed Single-Precision Floating-Point Valuesvfmadd213pd— Fused Multiply-Add of Packed Double-Precision Floating-Point Valuesvfmadd213ps— Fused Multiply-Add of Packed Single-Precision Floating-Point Valuesvfmadd231pd— Fused Multiply-Add of Packed Double-Precision Floating-Point Valuesvfmadd231ps— Fused Multiply-Add of Packed Single-Precision Floating-Point Valuesvfmaddsub132pd— Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Valuesvfmaddsub132ps— Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Valuesvfmaddsub213pd— Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Valuesvfmaddsub213ps— Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Valuesvfmaddsub231pd— Fused Multiply-Alternating Add/Subtract of Packed Double-Precision Floating-Point Valuesvfmaddsub231ps— Fused Multiply-Alternating Add/Subtract of Packed Single-Precision Floating-Point Valuesvfmsub132pd— Fused Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfmsub132ps— Fused Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfmsub213pd— Fused Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfmsub213ps— Fused Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfmsub231pd— Fused Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfmsub231ps— Fused Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfmsubadd132pd— Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Valuesvfmsubadd132ps— Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Valuesvfmsubadd213pd— Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Valuesvfmsubadd213ps— Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Valuesvfmsubadd231pd— Fused Multiply-Alternating Subtract/Add of Packed Double-Precision Floating-Point Valuesvfmsubadd231ps— Fused Multiply-Alternating Subtract/Add of Packed Single-Precision Floating-Point Valuesvfnmadd132pd— Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Valuesvfnmadd132ps— Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Valuesvfnmadd213pd— Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Valuesvfnmadd213ps— Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Valuesvfnmadd231pd— Fused Negative Multiply-Add of Packed Double-Precision Floating-Point Valuesvfnmadd231ps— Fused Negative Multiply-Add of Packed Single-Precision Floating-Point Valuesvfnmsub132pd— Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfnmsub132ps— Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfnmsub213pd— Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfnmsub213ps— Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfnmsub231pd— Fused Negative Multiply-Subtract of Packed Double-Precision Floating-Point Valuesvfnmsub231ps— Fused Negative Multiply-Subtract of Packed Single-Precision Floating-Point Valuesvfpclasspd— Test Class of Packed Double-Precision Floating-Point Valuesvfpclassps— Test Class of Packed Single-Precision Floating-Point Valuesvfpclasssd— Test Class of Scalar Double-Precision Floating-Point Valuevfpclassss— Test Class of Scalar Single-Precision Floating-Point Valuevgatherdpd— Gather Packed Double-Precision Floating-Point Values Using Signed Doubleword Indicesvgatherdps— Gather Packed Single-Precision Floating-Point Values Using Signed Doubleword Indicesvgatherpf0dpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hintvgatherpf0dps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hintvgatherpf0qpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hintvgatherpf0qps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hintvgatherpf1dpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hintvgatherpf1dps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hintvgatherpf1qpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hintvgatherpf1qps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hintvgatherqpd— Gather Packed Double-Precision Floating-Point Values Using Signed Quadword Indicesvgatherqps— Gather Packed Single-Precision Floating-Point Values Using Signed Quadword Indicesvgetexppd— Extract Exponents of Packed Double-Precision Floating-Point Values as Double-Precision Floating-Point Valuesvgetexpps— Extract Exponents of Packed Single-Precision Floating-Point Values as Single-Precision Floating-Point Valuesvgetexpsd— Extract Exponent of Scalar Double-Precision Floating-Point Value as Double-Precision Floating-Point Valuevgetexpss— Extract Exponent of Scalar Single-Precision Floating-Point Value as Single-Precision Floating-Point Valuevgetmantpd— Extract Normalized Mantissas from Packed Double-Precision Floating-Point Valuesvgetmantps— Extract Normalized Mantissas from Packed Single-Precision Floating-Point Valuesvgetmantsd— Extract Normalized Mantissa from Scalar Double-Precision Floating-Point Valuevgetmantss— Extract Normalized Mantissa from Scalar Single-Precision Floating-Point Valuevinsertf32x4— Insert 128 Bits of Packed Single-Precision Floating-Point Valuesvinsertf32x8— Insert 256 Bits of Packed Single-Precision Floating-Point Valuesvinsertf64x2— Insert 128 Bits of Packed Double-Precision Floating-Point Valuesvinsertf64x4— Insert 256 Bits of Packed Double-Precision Floating-Point Valuesvinserti32x4— Insert 128 Bits of Packed Doubleword Integer Valuesvinserti32x8— Insert 256 Bits of Packed Doubleword Integer Valuesvinserti64x2— Insert 128 Bits of Packed Quadword Integer Valuesvinserti64x4— Insert 256 Bits of Packed Quadword Integer Valuesvmaxpd— Return Maximum Packed Double-Precision Floating-Point Valuesvmaxph— Return Maximum Packed Half-Precision Floating-Point Valuesvmaxps— Return Maximum Packed Single-Precision Floating-Point Valuesvminpd— Return Minimum Packed Double-Precision Floating-Point Valuesvminph— Return Minimum Packed Half-Precision Floating-Point Valuesvminps— Return Minimum Packed Single-Precision Floating-Point Valuesvmovapd— Move Aligned Packed Double-Precision Floating-Point Valuesvmovaps— Move Aligned Packed Single-Precision Floating-Point Valuesvmovddup— Move One Double-FP and Duplicatevmovdqa32— Move Aligned Doubleword Valuesvmovdqa64— Move Aligned Quadword Valuesvmovdqu16— Move Unaligned Word Valuesvmovdqu32— Move Unaligned Doubleword Valuesvmovdqu64— Move Unaligned Quadword Valuesvmovdqu8— Move Unaligned Byte Valuesvmovntdq— Store Double Quadword Using Non-Temporal Hintvmovntdqa— Load Double Quadword Non-Temporal Aligned Hintvmovntpd— Store Packed Double-Precision Floating-Point Values Using Non-Temporal Hintvmovntps— Store Packed Single-Precision Floating-Point Values Using Non-Temporal Hintvmovshdup— Move Packed Single-FP High and Duplicatevmovsldup— Move Packed Single-FP Low and Duplicatevmovupd— Move Unaligned Packed Double-Precision Floating-Point Valuesvmovups— Move Unaligned Packed Single-Precision Floating-Point Valuesvmulpd— Multiply Packed Double-Precision Floating-Point Valuesvmulps— Multiply Packed Single-Precision Floating-Point Valuesvorpd— Bitwise Logical OR of Double-Precision Floating-Point Valuesvorps— Bitwise Logical OR of Single-Precision Floating-Point Valuesvpabsb— Packed Absolute Value of Byte Integersvpabsd— Packed Absolute Value of Doubleword Integersvpabsq— Packed Absolute Value of Quadword Integersvpabsw— Packed Absolute Value of Word Integersvpackssdw— Pack Doublewords into Words with Signed Saturationvpacksswb— Pack Words into Bytes with Signed Saturationvpackusdw— Pack Doublewords into Words with Unsigned Saturationvpackuswb— Pack Words into Bytes with Unsigned Saturationvpaddb— Add Packed Byte Integersvpaddd— Add Packed Doubleword Integersvpaddq— Add Packed Quadword Integersvpaddsb— Add Packed Signed Byte Integers with Signed Saturationvpaddsw— Add Packed Signed Word Integers with Signed Saturationvpaddusb— Add Packed Unsigned Byte Integers with Unsigned Saturationvpaddusw— Add Packed Unsigned Word Integers with Unsigned Saturationvpaddw— Add Packed Word Integersvpalignr— Packed Align Rightvpandd— Bitwise Logical AND of Packed Doubleword Integersvpandnd— Bitwise Logical AND NOT of Packed Doubleword Integersvpandnq— Bitwise Logical AND NOT of Packed Quadword Integersvpandq— Bitwise Logical AND of Packed Quadword Integersvpavgb— Average Packed Byte Integersvpavgw— Average Packed Word Integersvpblendmb— Blend Byte Vectors Using an OpMask Controlvpblendmd— Blend Doubleword Vectors Using an OpMask Controlvpblendmq— Blend Quadword Vectors Using an OpMask Controlvpblendmw— Blend Word Vectors Using an OpMask Controlvpbroadcastb— Broadcast Byte Integervpbroadcastd— Broadcast Doubleword Integervpbroadcastmb2q— Broadcast Low Byte of Mask Register to Packed Quadword Valuesvpbroadcastmw2d— Broadcast Low Word of Mask Register to Packed Doubleword Valuesvpbroadcastq— Broadcast Quadword Integervpbroadcastw— Broadcast Word Integervpcmpb— Compare Packed Signed Byte Valuesvpcmpd— Compare Packed Signed Doubleword Valuesvpcmpeqb— Compare Packed Byte Data for Equalityvpcmpeqd— Compare Packed Doubleword Data for Equalityvpcmpeqq— Compare Packed Quadword Data for Equalityvpcmpequbvpcmpequdvpcmpequqvpcmpequwvpcmpeqw— Compare Packed Word Data for Equalityvpcmpgebvpcmpgedvpcmpgeqvpcmpgeubvpcmpgeudvpcmpgeuqvpcmpgeuwvpcmpgewvpcmpgtb— Compare Packed Signed Byte Integers for Greater Thanvpcmpgtd— Compare Packed Signed Doubleword Integers for Greater Thanvpcmpgtq— Compare Packed Data for Greater Thanvpcmpgtubvpcmpgtudvpcmpgtuqvpcmpgtuwvpcmpgtw— Compare Packed Signed Word Integers for Greater Thanvpcmplebvpcmpledvpcmpleqvpcmpleubvpcmpleudvpcmpleuqvpcmpleuwvpcmplewvpcmpltbvpcmpltdvpcmpltqvpcmpltubvpcmpltudvpcmpltuqvpcmpltuwvpcmpltwvpcmpneqbvpcmpneqdvpcmpneqqvpcmpnequbvpcmpnequdvpcmpnequqvpcmpnequwvpcmpneqwvpcmpngtbvpcmpngtdvpcmpngtqvpcmpngtubvpcmpngtudvpcmpngtuqvpcmpngtuwvpcmpngtwvpcmpnlebvpcmpnledvpcmpnleqvpcmpnleubvpcmpnleudvpcmpnleuqvpcmpnleuwvpcmpnlewvpcmpnltbvpcmpnltdvpcmpnltqvpcmpnltubvpcmpnltudvpcmpnltuqvpcmpnltuwvpcmpnltwvpcmpq— Compare Packed Signed Quadword Valuesvpcmpub— Compare Packed Unsigned Byte Valuesvpcmpud— Compare Packed Unsigned Doubleword Valuesvpcmpuq— Compare Packed Unsigned Quadword Valuesvpcmpuw— Compare Packed Unsigned Word Valuesvpcmpw— Compare Packed Signed Word Valuesvpcompressd— Store Sparse Packed Doubleword Integer Values into Dense Memory/Registervpcompressq— Store Sparse Packed Quadword Integer Values into Dense Memory/Registervpconflictd— Detect Conflicts Within a Vector of Packed Doubleword Values into Dense Memory/Registervpconflictq— Detect Conflicts Within a Vector of Packed Quadword Values into Dense Memory/Registervpermb— Permute Byte Integersvpermd— Permute Doubleword Integersvpermi2b— Full Permute of Bytes From Two Tables Overwriting the Indexvpermi2d— Full Permute of Doublewords From Two Tables Overwriting the Indexvpermi2pd— Full Permute of Double-Precision Floating-Point Values From Two Tables Overwriting the Indexvpermi2ps— Full Permute of Single-Precision Floating-Point Values From Two Tables Overwriting the Indexvpermi2q— Full Permute of Quadwords From Two Tables Overwriting the Indexvpermi2w— Full Permute of Words From Two Tables Overwriting the Indexvpermilpd— Permute Double-Precision Floating-Point Valuesvpermilps— Permute Single-Precision Floating-Point Valuesvpermpd— Permute Double-Precision Floating-Point Elementsvpermps— Permute Single-Precision Floating-Point Elementsvpermq— Permute Quadword Integersvpermt2b— Full Permute of Bytes From Two Tables Overwriting a Tablevpermt2d— Full Permute of Doublewords From Two Tables Overwriting a Tablevpermt2pd— Full Permute of Double-Precision Floating-Point Values From Two Tables Overwriting a Tablevpermt2ps— Full Permute of Single-Precision Floating-Point Values From Two Tables Overwriting a Tablevpermt2q— Full Permute of Quadwords From Two Tables Overwriting a Tablevpermt2w— Full Permute of Words From Two Tables Overwriting a Tablevpermw— Permute Word Integersvpexpandd— Load Sparse Packed Doubleword Integer Values from Dense Memory/Registervpexpandq— Load Sparse Packed Quadword Integer Values from Dense Memory/Registervpextrb— Extract Bytevpextrw— Extract Wordvpgatherdd— Gather Packed Doubleword Values Using Signed Doubleword Indicesvpgatherdq— Gather Packed Quadword Values Using Signed Doubleword Indicesvpgatherqd— Gather Packed Doubleword Values Using Signed Quadword Indicesvpgatherqq— Gather Packed Quadword Values Using Signed Quadword Indicesvplzcntd— Count the Number of Leading Zero Bits for Packed Doubleword Valuesvplzcntq— Count the Number of Leading Zero Bits for Packed Quadword Valuesvpmadd52huq— Packed Multiply of Unsigned 52-bit Unsigned Integers and Add High 52-bit Products to Quadword Accumulatorsvpmadd52luq— Packed Multiply of Unsigned 52-bit Integers and Add the Low 52-bit Products to Quadword Accumulatorsvpmaddubsw— Multiply and Add Packed Signed and Unsigned Byte Integersvpmaddwd— Multiply and Add Packed Signed Word Integersvpmaxsb— Maximum of Packed Signed Byte Integersvpmaxsd— Maximum of Packed Signed Doubleword Integersvpmaxsq— Maximum of Packed Signed Quadword Integersvpmaxsw— Maximum of Packed Signed Word Integersvpmaxub— Maximum of Packed Unsigned Byte Integersvpmaxud— Maximum of Packed Unsigned Doubleword Integersvpmaxuq— Maximum of Packed Unsigned Quadword Integersvpmaxuw— Maximum of Packed Unsigned Word Integersvpminsb— Minimum of Packed Signed Byte Integersvpminsd— Minimum of Packed Signed Doubleword Integersvpminsq— Minimum of Packed Signed Quadword Integersvpminsw— Minimum of Packed Signed Word Integersvpminub— Minimum of Packed Unsigned Byte Integersvpminud— Minimum of Packed Unsigned Doubleword Integersvpminuq— Minimum of Packed Unsigned Quadword Integersvpminuw— Minimum of Packed Unsigned Word Integersvpmovb2m— Move Signs of Packed Byte Integers to Mask Registervpmovd2m— Move Signs of Packed Doubleword Integers to Mask Registervpmovdb— Down Convert Packed Doubleword Values to Byte Values with Truncationvpmovdw— Down Convert Packed Doubleword Values to Word Values with Truncationvpmovm2b— Expand Bits of Mask Register to Packed Byte Integersvpmovm2d— Expand Bits of Mask Register to Packed Doubleword Integersvpmovm2q— Expand Bits of Mask Register to Packed Quadword Integersvpmovm2w— Expand Bits of Mask Register to Packed Word Integersvpmovq2m— Move Signs of Packed Quadword Integers to Mask Registervpmovqb— Down Convert Packed Quadword Values to Byte Values with Truncationvpmovqd— Down Convert Packed Quadword Values to Doubleword Values with Truncationvpmovqw— Down Convert Packed Quadword Values to Word Values with Truncationvpmovsdb— Down Convert Packed Doubleword Values to Byte Values with Signed Saturationvpmovsdw— Down Convert Packed Doubleword Values to Word Values with Signed Saturationvpmovsqb— Down Convert Packed Quadword Values to Byte Values with Signed Saturationvpmovsqd— Down Convert Packed Quadword Values to Doubleword Values with Signed Saturationvpmovsqw— Down Convert Packed Quadword Values to Word Values with Signed Saturationvpmovswb— Down Convert Packed Word Values to Byte Values with Signed Saturationvpmovsxbd— Move Packed Byte Integers to Doubleword Integers with Sign Extensionvpmovsxbq— Move Packed Byte Integers to Quadword Integers with Sign Extensionvpmovsxbw— Move Packed Byte Integers to Word Integers with Sign Extensionvpmovsxdq— Move Packed Doubleword Integers to Quadword Integers with Sign Extensionvpmovsxwd— Move Packed Word Integers to Doubleword Integers with Sign Extensionvpmovsxwq— Move Packed Word Integers to Quadword Integers with Sign Extensionvpmovusdb— Down Convert Packed Doubleword Values to Byte Values with Unsigned Saturationvpmovusdw— Down Convert Packed Doubleword Values to Word Values with Unsigned Saturationvpmovusqb— Down Convert Packed Quadword Values to Byte Values with Unsigned Saturationvpmovusqd— Down Convert Packed Quadword Values to Doubleword Values with Unsigned Saturationvpmovusqw— Down Convert Packed Quadword Values to Word Values with Unsigned Saturationvpmovuswb— Down Convert Packed Word Values to Byte Values with Unsigned Saturationvpmovw2m— Move Signs of Packed Word Integers to Mask Registervpmovwb— Down Convert Packed Word Values to Byte Values with Truncationvpmovzxbd— Move Packed Byte Integers to Doubleword Integers with Zero Extensionvpmovzxbq— Move Packed Byte Integers to Quadword Integers with Zero Extensionvpmovzxbw— Move Packed Byte Integers to Word Integers with Zero Extensionvpmovzxdq— Move Packed Doubleword Integers to Quadword Integers with Zero Extensionvpmovzxwd— Move Packed Word Integers to Doubleword Integers with Zero Extensionvpmovzxwq— Move Packed Word Integers to Quadword Integers with Zero Extensionvpmuldq— Multiply Packed Signed Doubleword Integers and Store Quadword Resultvpmulhrsw— Packed Multiply Signed Word Integers and Store High Result with Round and Scalevpmulhuw— Multiply Packed Unsigned Word Integers and Store High Resultvpmulhw— Multiply Packed Signed Word Integers and Store High Resultvpmulld— Multiply Packed Signed Doubleword Integers and Store Low Resultvpmullq— Multiply Packed Signed Quadword Integers and Store Low Resultvpmullw— Multiply Packed Signed Word Integers and Store Low Resultvpmultishiftqb— Select Packed Unaligned Bytes from Quadword Sourcesvpmuludq— Multiply Packed Unsigned Doubleword Integersvpord— Bitwise Logical OR of Packed Doubleword Integersvporq— Bitwise Logical OR of Packed Quadword Integersvprold— Rotate Packed Doubleword Leftvprolq— Rotate Packed Quadword Leftvprolvd— Variable Rotate Packed Doubleword Leftvprolvq— Variable Rotate Packed Quadword Leftvprord— Rotate Packed Doubleword Rightvprorq— Rotate Packed Quadword Rightvprorvd— Variable Rotate Packed Doubleword Rightvprorvq— Variable Rotate Packed Quadword Rightvpsadbw— Compute Sum of Absolute Differencesvpscatterdd— Scatter Packed Doubleword Values with Signed Doubleword Indicesvpscatterdq— Scatter Packed Quadword Values with Signed Doubleword Indicesvpscatterqd— Scatter Packed Doubleword Values with Signed Quadword Indicesvpscatterqq— Scatter Packed Quadword Values with Signed Quadword Indicesvpshufb— Packed Shuffle Bytesvpshufd— Shuffle Packed Doublewordsvpshufhw— Shuffle Packed High Wordsvpshuflw— Shuffle Packed Low Wordsvpslld— Shift Packed Doubleword Data Left Logicalvpslldq— Shift Packed Double Quadword Left Logicalvpsllq— Shift Packed Quadword Data Left Logicalvpsllvd— Variable Shift Packed Doubleword Data Left Logicalvpsllvq— Variable Shift Packed Quadword Data Left Logicalvpsllvw— Variable Shift Packed Word Data Left Logicalvpsllw— Shift Packed Word Data Left Logicalvpsrad— Shift Packed Doubleword Data Right Arithmeticvpsraq— Shift Packed Quadword Data Right Arithmeticvpsravd— Variable Shift Packed Doubleword Data Right Arithmeticvpsravq— Variable Shift Packed Quadword Data Right Arithmeticvpsravw— Variable Shift Packed Word Data Right Arithmeticvpsraw— Shift Packed Word Data Right Arithmeticvpsrld— Shift Packed Doubleword Data Right Logicalvpsrldq— Shift Packed Double Quadword Right Logicalvpsrlq— Shift Packed Quadword Data Right Logicalvpsrlvd— Variable Shift Packed Doubleword Data Right Logicalvpsrlvq— Variable Shift Packed Quadword Data Right Logicalvpsrlvw— Variable Shift Packed Word Data Right Logicalvpsrlw— Shift Packed Word Data Right Logicalvpsubb— Subtract Packed Byte Integersvpsubd— Subtract Packed Doubleword Integersvpsubq— Subtract Packed Quadword Integersvpsubsb— Subtract Packed Signed Byte Integers with Signed Saturationvpsubsw— Subtract Packed Signed Word Integers with Signed Saturationvpsubusb— Subtract Packed Unsigned Byte Integers with Unsigned Saturationvpsubusw— Subtract Packed Unsigned Word Integers with Unsigned Saturationvpsubw— Subtract Packed Word Integersvpternlogd— Bitwise Ternary Logical Operation on Doubleword Valuesvpternlogq— Bitwise Ternary Logical Operation on Quadword Valuesvptestmb— Logical AND of Packed Byte Integer Values and Set Maskvptestmd— Logical AND of Packed Doubleword Integer Values and Set Maskvptestmq— Logical AND of Packed Quadword Integer Values and Set Maskvptestmw— Logical AND of Packed Word Integer Values and Set Maskvptestnmb— Logical NAND of Packed Byte Integer Values and Set Maskvptestnmd— Logical NAND of Packed Doubleword Integer Values and Set Maskvptestnmq— Logical NAND of Packed Quadword Integer Values and Set Maskvptestnmw— Logical NAND of Packed Word Integer Values and Set Maskvpunpckhbw— Unpack and Interleave High-Order Bytes into Wordsvpunpckhdq— Unpack and Interleave High-Order Doublewords into Quadwordsvpunpckhqdq— Unpack and Interleave High-Order Quadwords into Double Quadwordsvpunpckhwd— Unpack and Interleave High-Order Words into Doublewordsvpunpcklbw— Unpack and Interleave Low-Order Bytes into Wordsvpunpckldq— Unpack and Interleave Low-Order Doublewords into Quadwordsvpunpcklqdq— Unpack and Interleave Low-Order Quadwords into Double Quadwordsvpunpcklwd— Unpack and Interleave Low-Order Words into Doublewordsvpxord— Bitwise Logical Exclusive OR of Packed Doubleword Integersvpxorq— Bitwise Logical Exclusive OR of Packed Quadword Integersvrangepd— Range Restriction Calculation For Packed Pairs of Double-Precision Floating-Point Valuesvrangeps— Range Restriction Calculation For Packed Pairs of Single-Precision Floating-Point Valuesvrangesd— Range Restriction Calculation For a pair of Scalar Double-Precision Floating-Point Valuesvrangess— Range Restriction Calculation For a pair of Scalar Single-Precision Floating-Point Valuesvrcp14pd— Compute Approximate Reciprocals of Packed Double-Precision Floating-Point Valuesvrcp14ps— Compute Approximate Reciprocals of Packed Single-Precision Floating-Point Valuesvrcp14sd— Compute Approximate Reciprocal of a Scalar Double-Precision Floating-Point Valuevrcp14ss— Compute Approximate Reciprocal of a Scalar Single-Precision Floating-Point Valuevrcp28pd— Approximation to the Reciprocal of Packed Double-Precision Floating-Point Values with Less Than 2^-28 Relative Errorvrcp28ps— Approximation to the Reciprocal of Packed Single-Precision Floating-Point Values with Less Than 2^-28 Relative Errorvrcp28sd— Approximation to the Reciprocal of a Scalar Double-Precision Floating-Point Value with Less Than 2^-28 Relative Errorvrcp28ss— Approximation to the Reciprocal of a Scalar Single-Precision Floating-Point Value with Less Than 2^-28 Relative Errorvreducepd— Perform Reduction Transformation on Packed Double-Precision Floating-Point Valuesvreduceps— Perform Reduction Transformation on Packed Single-Precision Floating-Point Valuesvreducesd— Perform Reduction Transformation on a Scalar Double-Precision Floating-Point Valuevreducess— Perform Reduction Transformation on a Scalar Single-Precision Floating-Point Valuevrndscalepd— Round Packed Double-Precision Floating-Point Values To Include A Given Number Of Fraction Bitsvrndscaleph— Round Packed Half-Precision Floating-Point Values To Include A Given Number Of Fraction Bitsvrndscaleps— Round Packed Single-Precision Floating-Point Values To Include A Given Number Of Fraction Bitsvrndscalesd— Round Scalar Double-Precision Floating-Point Value To Include A Given Number Of Fraction Bitsvrndscalesh— Round Scalar Half-Precision Floating-Point Value To Include A Given Number Of Fraction Bitsvrndscaless— Round Scalar Single-Precision Floating-Point Value To Include A Given Number Of Fraction Bitsvrsqrt14pd— Compute Approximate Reciprocals of Square Roots of Packed Double-Precision Floating-Point Valuesvrsqrt14ps— Compute Approximate Reciprocals of Square Roots of Packed Single-Precision Floating-Point Valuesvrsqrt14sd— Compute Approximate Reciprocal of a Square Root of a Scalar Double-Precision Floating-Point Valuevrsqrt14ss— Compute Approximate Reciprocal of a Square Root of a Scalar Single-Precision Floating-Point Valuevrsqrt28pd— Approximation to the Reciprocal Square Root of Packed Double-Precision Floating-Point Values with Less Than 2^-28 Relative Errorvrsqrt28ps— Approximation to the Reciprocal Square Root of Packed Single-Precision Floating-Point Values with Less Than 2^-28 Relative Errorvrsqrt28sd— Approximation to the Reciprocal Square Root of a Scalar Double-Precision Floating-Point Value with Less Than 2^-28 Relative Errorvrsqrt28ss— Approximation to the Reciprocal Square Root of a Scalar Single-Precision Floating-Point Value with Less Than 2^-28 Relative Errorvscalefpd— Scale Packed Double-Precision Floating-Point Values With Double-Precision Floating-Point Valuesvscalefps— Scale Packed Single-Precision Floating-Point Values With Single-Precision Floating-Point Valuesvscalefsd— Scale Scalar Double-Precision Floating-Point Value With a Double-Precision Floating-Point Valuevscalefss— Scale Scalar Single-Precision Floating-Point Value With a Single-Precision Floating-Point Valuevscatterdpd— Scatter Packed Double-Precision Floating-Point Values with Signed Doubleword Indicesvscatterdps— Scatter Packed Single-Precision Floating-Point Values with Signed Doubleword Indicesvscatterpf0dpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hint with Intent to Writevscatterpf0dps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T0 Hint with Intent to Writevscatterpf0qpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hint with Intent to Writevscatterpf0qps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T0 Hint with Intent to Writevscatterpf1dpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hint with Intent to Writevscatterpf1dps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Doubleword Indices Using T1 Hint with Intent to Writevscatterpf1qpd— Sparse Prefetch Packed Double-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hint with Intent to Writevscatterpf1qps— Sparse Prefetch Packed Single-Precision Floating-Point Data Values with Signed Quadword Indices Using T1 Hint with Intent to Writevscatterqpd— Scatter Packed Double-Precision Floating-Point Values with Signed Quadword Indicesvscatterqps— Scatter Packed Single-Precision Floating-Point Values with Signed Quadword Indicesvshuff32x4— Shuffle 128-Bit Packed Single-Precision Floating-Point Valuesvshuff64x2— Shuffle 128-Bit Packed Double-Precision Floating-Point Valuesvshufi32x4— Shuffle 128-Bit Packed Doubleword Integer Valuesvshufi64x2— Shuffle 128-Bit Packed Quadword Integer Valuesvshufpd— Shuffle Packed Double-Precision Floating-Point Valuesvshufps— Shuffle Packed Single-Precision Floating-Point Valuesvsqrtpd— Compute Square Roots of Packed Double-Precision Floating-Point Valuesvsqrtps— Compute Square Roots of Packed Single-Precision Floating-Point Valuesvsubpd— Subtract Packed Double-Precision Floating-Point Valuesvsubps— Subtract Packed Single-Precision Floating-Point Valuesvunpckhpd— Unpack and Interleave High Packed Double-Precision Floating-Point Valuesvunpckhps— Unpack and Interleave High Packed Single-Precision Floating-Point Valuesvunpcklpd— Unpack and Interleave Low Packed Double-Precision Floating-Point Valuesvunpcklps— Unpack and Interleave Low Packed Single-Precision Floating-Point Valuesvxorpd— Bitwise Logical XOR for Double-Precision Floating-Point Valuesvxorps— Bitwise Logical XOR for Single-Precision Floating-Point Values
AVX-512 mask register instructions
140 instructionsaddbadddaddqaddwandbanddandnbandndandnqandnwandqandwkaddkaddb— ADD Two 8-bit Maskskaddd— ADD Two 32-bit Maskskaddq— ADD Two 64-bit Maskskaddw— ADD Two 16-bit Maskskandkandb— Bitwise Logical AND 8-bit Maskskandd— Bitwise Logical AND 32-bit Maskskandnkandnb— Bitwise Logical AND NOT 8-bit Maskskandnd— Bitwise Logical AND NOT 32-bit Maskskandnq— Bitwise Logical AND NOT 64-bit Maskskandnw— Bitwise Logical AND NOT 16-bit Maskskandq— Bitwise Logical AND 64-bit Maskskandw— Bitwise Logical AND 16-bit Maskskmovkmovb— Move 8-bit Maskkmovd— Move 32-bit Maskkmovq— Move 64-bit Maskkmovw— Move 16-bit Maskknotknotb— NOT 8-bit Mask Registerknotd— NOT 32-bit Mask Registerknotq— NOT 64-bit Mask Registerknotw— NOT 16-bit Mask Registerkorkorb— Bitwise Logical OR 8-bit Maskskord— Bitwise Logical OR 32-bit Maskskorq— Bitwise Logical OR 64-bit Maskskortestkortestb— OR 8-bit Masks and Set Flagskortestd— OR 32-bit Masks and Set Flagskortestq— OR 64-bit Masks and Set Flagskortestw— OR 16-bit Masks and Set Flagskorw— Bitwise Logical OR 16-bit Maskskshiftlkshiftlb— Shift Left 8-bit Maskskshiftld— Shift Left 32-bit Maskskshiftlq— Shift Left 64-bit Maskskshiftlw— Shift Left 16-bit Maskskshiftrkshiftrb— Shift Right 8-bit Maskskshiftrd— Shift Right 32-bit Maskskshiftrq— Shift Right 64-bit Maskskshiftrw— Shift Right 16-bit Maskskshlkshlbkshldkshlqkshlwkshrkshrbkshrdkshrqkshrwktestktestb— Bit Test 8-bit Masks and Set Flagsktestd— Bit Test 32-bit Masks and Set Flagsktestq— Bit Test 64-bit Masks and Set Flagsktestw— Bit Test 16-bit Masks and Set Flagskunpckkunpckbw— Unpack and Interleave 8-bit Maskskunpckdkunpckdq— Unpack and Interleave 32-bit Maskskunpckqkunpckwkunpckwd— Unpack and Interleave 16-bit Maskskxnorkxnorb— Bitwise Logical XNOR 8-bit Maskskxnord— Bitwise Logical XNOR 32-bit Maskskxnorq— Bitwise Logical XNOR 64-bit Maskskxnorw— Bitwise Logical XNOR 16-bit Maskskxorkxorb— Bitwise Logical XOR 8-bit Maskskxord— Bitwise Logical XOR 32-bit Maskskxorq— Bitwise Logical XOR 64-bit Maskskxorw— Bitwise Logical XOR 16-bit Masksmovbmovwnotbnotdnotqnotworbordorqortestortestbortestdortestqortestworwshiftlshiftlbshiftldshiftlqshiftlwshiftrshiftrbshiftrdshiftrqshiftrwshlbshlqshlwshrbshrqshrwtestbtestdtestqtestwunpckunpckbwunpckdunpckdqunpckqunpckwunpckwdxnorxnorbxnordxnorqxnorwxorbxordxorqxorw
AVX10.2 BF16 instructions
29 instructionsvaddbf16vcmpbf16vcomisbf16vdivbf16vfmadd132bf16vfmadd213bf16vfmadd231bf16vfmsub132bf16vfmsub213bf16vfmsub231bf16vfnmadd132bf16vfnmadd213bf16vfnmadd231bf16vfnmsub132bf16vfnmsub213bf16vfnmsub231bf16vfpclassbf16vgetexpbf16vgetmantbf16vmaxbf16vminbf16vmulbf16vrcpbf16vreducebf16vrndscalebf16vrsqrtbf16vscalefbf16vsqrtbf16vsubbf16
AVX10.2 Compare scalar fp with enhanced eflags instructions
6 instructionsvcomxsdvcomxshvcomxssvucomxsdvucomxshvucomxss
AVX10.2 Convert instructions
14 instructionsvcvt2ph2bf8vcvt2ph2bf8svcvt2ph2hf8vcvt2ph2hf8svcvt2ps2phxvcvtbiasph2bf8vcvtbiasph2bf8svcvtbiasph2hf8vcvtbiasph2hf8svcvthf82phvcvtph2bf8vcvtph2bf8svcvtph2hf8vcvtph2hf8s
AVX10.2 Integer and FP16 VNNI, media new instructions
14 instructionsvdpphpsvmpsadbw— Compute Multiple Packed Sums of Absolute Differencevpdpbssd— Packed Dot Product of Signed-by-Singed Byte subvectors into Doublewordvpdpbssds— Packed Dot Product of Signed-by-Singed Byte subvectors into Doubleword with Saturationvpdpbsud— Packed Dot Product of Signed-by-Unsinged Byte subvectors into Doublewordvpdpbsuds— Packed Dot Product of Signed-by-Unsinged Byte subvectors into Doubleword with Saturationvpdpbuud— Packed Dot Product of Unsigned-by-Unsinged Byte subvectors into Doublewordvpdpbuuds— Packed Dot Product of Unsigned-by-Unsinged Byte subvectors into Doubleword with Saturationvpdpwsud— Packed Dot Product of Signed-by-Unsigned Word subvectors into Doublewordvpdpwsuds— Packed Dot Product of Signed-by-Unsigned Word subvectors into Doubleword with Saturationvpdpwusd— Packed Dot Product of Unsigned-by-Signed Word subvectors into Doublewordvpdpwusds— Packed Dot Product of Unsigned-by-Signed Word subvectors into Doubleword with Saturationvpdpwuud— Packed Dot Product of Unsigned-by-Unsigned Word subvectors into Doublewordvpdpwuuds— Packed Dot Product of Unsigned-by-Unsigned Word subvectors into Doubleword with Saturation
AVX10.2 MINMAX instructions
7 instructionsvminmaxbf16vminmaxpdvminmaxphvminmaxpsvminmaxsdvminmaxshvminmaxss
AVX10.2 Saturating convert instructions
24 instructionsvcvtbf162ibsvcvtbf162iubsvcvtph2ibsvcvtph2iubsvcvtps2ibsvcvtps2iubsvcvttbf162ibsvcvttbf162iubsvcvttpd2dqsvcvttpd2qqsvcvttpd2udqsvcvttpd2uqqsvcvttph2ibsvcvttph2iubsvcvttps2dqsvcvttps2ibsvcvttps2iubsvcvttps2qqsvcvttps2udqsvcvttps2uqqsvcvttsd2sisvcvttsd2usisvcvttss2sisvcvttss2usis
AVX512 4-iteration Dot Product
2 instructionsv4dpwssdv4dpwssds
AVX512 4-iteration Multiply-Add
4 instructionsv4fmaddpsv4fmaddssv4fnmaddpsv4fnmaddss
AVX512 Bfloat16 instructions
3 instructionsvcvtne2ps2bf16— Convert with Nearest-Even rounding 2 Single-Precision FP vectors into BFloat16 FP vectorvcvtneps2bf16— Convert with Nearest-Even rounding a Single-Precision FP vector into a BFloat16 FP vectorvdpbf16ps— Packed Dot Product of BFloat16 FP subvectors into Single-Precision FP values
AVX512 Bit Algorithms
5 instructionsvpopcntb— Packed Population Count for Byte Integersvpopcntd— Packed Population Count for Doubleword Integersvpopcntq— Packed Population Count for Quadword Integersvpopcntw— Packed Population Count for Word Integersvpshufbitqmb— Shuffle Bits From Quadword Elements Using Byte Indexes Into Mask
AVX512 mask intersect instructions
2 instructionsvp2intersectdvp2intersectq
AVX512 Vector Bit Manipulation Instructions 2
16 instructionsvpcompressb— Store Sparse Packed Byte Integer Values into Dense Memory/Registervpcompressw— Store Sparse Packed Word Integer Values into Dense Memory/Registervpexpandb— Load Sparse Packed Byte Integer Values from Dense Memory/Registervpexpandw— Load Sparse Packed Word Integer Values from Dense Memory/Registervpshldd— Concatenate and Shift Packed Doubleword Data Left Logicalvpshldq— Concatenate and Shift Packed Quadword Data Left Logicalvpshldvd— Concatenate and Variable Shift Packed Doubleword Data Left Logicalvpshldvq— Concatenate and Variable Shift Packed Quadword Data Left Logicalvpshldvw— Concatenate and Variable Shift Packed Word Data Left Logicalvpshldw— Concatenate and Shift Packed Word Data Left Logicalvpshrdd— Concatenate and Shift Packed Doubleword Data Right Logicalvpshrdq— Concatenate and Shift Packed Quadword Data Right Logicalvpshrdvd— Concatenate and Variable Shift Packed Doubleword Data Right Logicalvpshrdvq— Concatenate and Variable Shift Packed Quadword Data Right Logicalvpshrdvw— Concatenate and Variable Shift Packed Word Data Right Logicalvpshrdw— Concatenate and Shift Packed Word Data Right Logical
AVX512 VNNI
4 instructionsvpdpbusd— Packed Dot Product of Unsigned-by-Singed Byte subvectors into Doublewordvpdpbusds— Packed Dot Product of Unsigned-by-Singed Byte subvectors into Doubleword with Saturationvpdpwssd— Packed Dot Product of Signed-by-Signed Word subvectors into Doublewordvpdpwssds— Packed Dot Product of Signed-by-Signed Word subvectors into Doubleword with Saturation
BMI1 and BMI2 bit operations
10 instructionsandn— Logical AND NOTbextr— Bit Field Extractblsi— Isolate Lowest Set Bitblsmsk— Mask From Lowest Set Bitblsr— Reset Lowest Set Bitbzhi— Zero High Bits Starting with Specified Bit Positionlzcnt— Count the Number of Leading Zero Bitspdep— Parallel Bits Depositpext— Parallel Bits Extracttzcnt— Count the Number of Trailing Zero Bits
Conditional instructions (extensions)
90 instructionscfcmova(APX)cfcmovae(APX)cfcmovb(APX)cfcmovbe(APX)cfcmovc(APX)cfcmove(APX)cfcmovg(APX)cfcmovge(APX)cfcmovl(APX)cfcmovle(APX)cfcmovna(APX)cfcmovnae(APX)cfcmovnb(APX)cfcmovnbe(APX)cfcmovnc(APX)cfcmovne(APX)cfcmovng(APX)cfcmovnge(APX)cfcmovnl(APX)cfcmovnle(APX)cfcmovno(APX)cfcmovnp(APX)cfcmovns(APX)cfcmovnz(APX)cfcmovo(APX)cfcmovp(APX)cfcmovpe(APX)cfcmovpo(APX)cfcmovs(APX)cfcmovz(APX)cmpaexadd(CMPCCXADD)cmpaxadd(CMPCCXADD)cmpbexadd— Compare for Below or Equals and Add (CMPCCXADD)cmpbxadd— Compare for Below and Add (CMPCCXADD)cmpcxadd(CMPCCXADD)cmpexadd(CMPCCXADD)cmpgexadd(CMPCCXADD)cmpgxadd(CMPCCXADD)cmplexadd— Compare for Less or Equals and Add (CMPCCXADD)cmplxadd— Compare for Less and Add (CMPCCXADD)cmpnaexadd(CMPCCXADD)cmpnaxadd(CMPCCXADD)cmpnbexadd— Compare for Not Below or Equals and Add (CMPCCXADD)cmpnbxadd— Compare for Not Below and Add (CMPCCXADD)cmpncxadd(CMPCCXADD)cmpnexadd(CMPCCXADD)cmpngexadd(CMPCCXADD)cmpngxadd(CMPCCXADD)cmpnlexadd— Compare for Not Less or Equals and Add (CMPCCXADD)cmpnlxadd— Compare for Not Less and Add (CMPCCXADD)cmpnoxadd— Compare for Not Overflow and Add (CMPCCXADD)cmpnpxadd— Compare for Not Parity and Add (CMPCCXADD)cmpnsxadd— Compare for Not Sign and Add (CMPCCXADD)cmpnzxadd— Compare for Not Zero and Add (CMPCCXADD)cmpoxadd— Compare for Overflow and Add (CMPCCXADD)cmppexadd(CMPCCXADD)cmppoxadd(CMPCCXADD)cmppxadd— Compare for Parity and Add (CMPCCXADD)cmpsxadd— Compare for Sign and Add (CMPCCXADD)cmpzxadd— Compare for Zero and Add (CMPCCXADD)setaezu(APX)setazu(APX)setbezu(APX)setbzu(APX)setczu(APX)setezu(APX)setgezu(APX)setgzu(APX)setlezu(APX)setlzu(APX)setnaezu(APX)setnazu(APX)setnbezu(APX)setnbzu(APX)setnczu(APX)setnezu(APX)setngezu(APX)setngzu(APX)setnlezu(APX)setnlzu(APX)setnozu(APX)setnpzu(APX)setnszu(APX)setnzzu(APX)setozu(APX)setpezu(APX)setpozu(APX)setpzu(APX)setszu(APX)setzzu(APX)
doc 319433-034 May 2018
4 instructionscldemote— Cache Line Demotemovdir64b— MOVe to DIRect store 64 Bytesmovdiri— MOVe to DIRect store Integerpconfig
doc 319433-058 June 2025
2 instructionspbndkbprefetchrst2
Extended Page Tables VMX instructions
2 instructionsinveptinvvpid
Flag register instructions (extensions)
2 instructionsclac(SMAP)stac(SMAP)
Galois field operations (GFNI)
6 instructionsgf2p8affineinvqb— Galois Field (2^8) Affine Inverse Transformationgf2p8affineqb— Galois Field (2^8) Affine Transformationgf2p8mulb— Galois Field Multiply Bytesvgf2p8affineinvqb— Galois Field (2^8) Affine Inverse Transformationvgf2p8affineqb— Galois Field (2^8) Affine Transformationvgf2p8mulb— Galois Field Multiply Bytes
Generic memory operations
6 instructionsprefetchit0— Prefetch Code Into Instruction Caches using IT0 Hintprefetchit1— Prefetch Code Into Instruction Caches using IT1 Hintprefetchnta— Prefetch Data Into Caches using NTA Hintprefetcht0— Prefetch Data Into Caches using T0 Hintprefetcht1— Prefetch Data Into Caches using T1 Hintprefetcht2— Prefetch Data Into Caches using T2 Hint
Geode (Cyrix) 3DNow! additions
2 instructionspfrcpvpfrsqrtv
History reset
1 instructionhreset
I/O instructions
2 instructionsinout
Instructions from ISE doc 319433-040, June 2020
4 instructionsenqcmdenqcmdsxresldtrkxsusldtrk
Intel Advanced Matrix Extensions (AMX)
44 instructionsldtilecfg— LoaD TILE ConFiGurationsttilecfg— STore TILE ConFiGurationt2rpntlvwz0t2rpntlvwz0rst2rpntlvwz0rst1t2rpntlvwz0t1t2rpntlvwz1t2rpntlvwz1rst2rpntlvwz1rst1t2rpntlvwz1t1tcmmimfp16ps— Tile Complex Matrix Multiply IMaginary part of FP16 tiles with Packed Single-precision accumulationtcmmrlfp16ps— Tile Complex Matrix Multiply ReaL part of FP16 tiles with Packed Single-precision accumulationtconjtcmmimfp16pstconjtfp16tcvtrowd2pstcvtrowps2bf16htcvtrowps2bf16ltcvtrowps2phhtcvtrowps2phltdpbf16ps— Tile Dot Product of BF16 tiles with Packed Single-precision accumulationtdpbf8pstdpbhf8pstdpbssd— Tile Dot Product of Signed bytes by Signed bytes with Doubleword accumulationtdpbsud— Tile Dot Product of Signed bytes by Unsigned bytes with Doubleword accumulationtdpbusd— Tile Dot Product of Unsigned bytes by Signed bytes with Doubleword accumulationtdpbuud— Tile Dot Product of Unsigned bytes by Unsigned bytes with Doubleword accumulationtdpfp16ps— Tile Dot Product of FP16 tiles with Packed Single-precision accumulationtdphbf8pstdphf8pstileloadd— TILE LOAD Datatileloaddrstileloaddrst1tileloaddt1— TILE LOAD Data with T1 caching hinttilemovrowtilerelease— TILE RELEASE register statetilestored— TILE STORE Datatilezero— TILE ZERO datatmmultf32psttcmmimfp16psttcmmrlfp16psttdpbf16psttdpfp16psttmmultf32psttransposed
Intel AES instructions
6 instructionsaesdec— Perform One Round of an AES Decryption Flowaesdeclast— Perform Last Round of an AES Decryption Flowaesenc— Perform One Round of an AES Encryption Flowaesenclast— Perform Last Round of an AES Encryption Flowaesimc— Perform the AES InvMixColumn Transformationaeskeygenassist— AES Round Key Generation Assist
Intel AES Key Locker
11 instructionsaesdec128klaesdec256klaesdecwide128klaesdecwide256klaesenc128klaesenc256klaesencwide128klaesencwide256klencodekey128encodekey256loadiwkey
Intel AVX AES instructions
2 instructionsvaesimc— Perform the AES InvMixColumn Transformationvaeskeygenassist— AES Round Key Generation Assist
Intel AVX Carry-Less Multiplication instructions (CLMUL)
21 instructionsvfcmulcph— Fused Conjugate Multiply of Complex Packed Half-Precision Floating-Point Valuesvfmadd132shvfmadd213shvfmadd231shvfmsub132sh— Fused Multiply-Subtract of Scalar Half-Precision Floating-Point Valuesvfmsub213sh— Fused Multiply-Subtract of Scalar Half-Precision Floating-Point Valuesvfmsub231sh— Fused Multiply-Subtract of Scalar Half-Precision Floating-Point Valuesvfmulcph— Fused Fused Multiply of Complex Packed Half-Precision Floating-Point Valuesvfnmadd132shvfnmadd213shvfnmadd231shvfnmsub132sh— Fused Negative Multiply-Subtract of Scalar Half-Precision Floating-Point Valuesvfnmsub213sh— Fused Negative Multiply-Subtract of Scalar Half-Precision Floating-Point Valuesvfnmsub231sh— Fused Negative Multiply-Subtract of Scalar Half-Precision Floating-Point Valuesvmaxsh— Return Maximum Scalar Half-Precision Floating-Point Valuevminsh— Return Minimum Scalar Half-Precision Floating-Point Valuevpclmulhqhqdqvpclmulhqlqdqvpclmullqhqdqvpclmullqlqdqvpclmulqdq— Carry-Less Quadword Multiplication
Intel AVX instructions
204 instructionsvaddsd— Add Scalar Double-Precision Floating-Point Valuesvaddss— Add Scalar Single-Precision Floating-Point Valuesvaddsubpd— Packed Double-FP Add/Subtractvaddsubps— Packed Single-FP Add/Subtractvblendpd— Blend Packed Double Precision Floating-Point Valuesvblendps— Blend Packed Single Precision Floating-Point Valuesvblendvpd— Variable Blend Packed Double Precision Floating-Point Valuesvblendvps— Variable Blend Packed Single Precision Floating-Point Valuesvbroadcastf128— Broadcast 128 Bit of Floating-Point Datavcmpeq_ospdvcmpeq_ospsvcmpeq_ossdvcmpeq_osssvcmpeq_uqsdvcmpeq_uqssvcmpeq_ussdvcmpeq_usssvcmpeqsdvcmpeqssvcmpfalse_oqsdvcmpfalse_oqssvcmpfalse_ossdvcmpfalse_osssvcmpfalsesdvcmpfalsessvcmpge_oqsdvcmpge_oqssvcmpge_ossdvcmpge_osssvcmpgesdvcmpgessvcmpgt_oqsdvcmpgt_oqssvcmpgt_ossdvcmpgt_osssvcmpgtsdvcmpgtssvcmple_oqsdvcmple_oqssvcmple_ossdvcmple_osssvcmplesdvcmplessvcmplt_oqsdvcmplt_oqssvcmplt_ossdvcmplt_osssvcmpltsdvcmpltssvcmpneq_oqsdvcmpneq_oqssvcmpneq_ossdvcmpneq_osssvcmpneq_uqsdvcmpneq_uqssvcmpneq_ussdvcmpneq_usssvcmpneqsdvcmpneqssvcmpnge_uqsdvcmpnge_uqssvcmpnge_ussdvcmpnge_usssvcmpngesdvcmpngessvcmpngt_uqsdvcmpngt_uqssvcmpngt_ussdvcmpngt_usssvcmpngtsdvcmpngtssvcmpnle_uqsdvcmpnle_uqssvcmpnle_ussdvcmpnle_usssvcmpnlesdvcmpnlessvcmpnlt_uqsdvcmpnlt_uqssvcmpnlt_ussdvcmpnlt_usssvcmpnltsdvcmpnltssvcmpord_qsdvcmpord_qssvcmpord_ssdvcmpord_sssvcmpordsdvcmpordssvcmpsd— Compare Scalar Double-Precision Floating-Point Valuesvcmpss— Compare Scalar Single-Precision Floating-Point Valuesvcmptrue_uqsdvcmptrue_uqssvcmptrue_ussdvcmptrue_usssvcmptruesdvcmptruessvcmpunord_qsdvcmpunord_qssvcmpunord_ssdvcmpunord_sssvcmpunordsdvcmpunordssvcomisd— Compare Scalar Ordered Double-Precision Floating-Point Values and Set EFLAGSvcomiss— Compare Scalar Ordered Single-Precision Floating-Point Values and Set EFLAGSvcvtpd2dq— Convert Packed Double-Precision FP Values to Packed Dword Integersvcvtpd2ps— Convert Packed Double-Precision FP Values to Packed Single-Precision FP Valuesvcvtsd2si— Convert Scalar Double-Precision FP Value to Integervcvtsd2ss— Convert Scalar Double-Precision FP Value to Scalar Single-Precision FP Valuevcvtsi2sd— Convert Dword Integer to Scalar Double-Precision FP Valuevcvtsi2ss— Convert Dword Integer to Scalar Single-Precision FP Valuevcvtss2sd— Convert Scalar Single-Precision FP Value to Scalar Double-Precision FP Valuevcvtss2si— Convert Scalar Single-Precision FP Value to Dword Integervcvttpd2dq— Convert with Truncation Packed Double-Precision FP Values to Packed Dword Integersvcvttsd2si— Convert with Truncation Scalar Double-Precision FP Value to Signed Integervcvttss2si— Convert with Truncation Scalar Single-Precision FP Value to Dword Integervdivsd— Divide Scalar Double-Precision Floating-Point Valuesvdivss— Divide Scalar Single-Precision Floating-Point Valuesvdppd— Dot Product of Packed Double Precision Floating-Point Valuesvdpps— Dot Product of Packed Single Precision Floating-Point Valuesvextractf128— Extract Packed Floating-Point Valuesvhaddpd— Packed Double-FP Horizontal Addvhaddps— Packed Single-FP Horizontal Addvhsubpd— Packed Double-FP Horizontal Subtractvhsubps— Packed Single-FP Horizontal Subtractvinsertf128— Insert Packed Floating-Point Valuesvinsertps— Insert Packed Single Precision Floating-Point Valuevlddqu— Load Unaligned Integer 128 Bitsvldmxcsr— Load MXCSR Registervldqquvmaskmovdqu— Store Selected Bytes of Double Quadwordvmaskmovpd— Conditional Move Packed Double-Precision Floating-Point Valuesvmaskmovps— Conditional Move Packed Single-Precision Floating-Point Valuesvmaxsd— Return Maximum Scalar Double-Precision Floating-Point Valuevmaxss— Return Maximum Scalar Single-Precision Floating-Point Valuevminsd— Return Minimum Scalar Double-Precision Floating-Point Valuevminss— Return Minimum Scalar Single-Precision Floating-Point Valuevmovd— Move Doublewordvmovdqa— Move Aligned Double Quadwordvmovdqu— Move Unaligned Double Quadwordvmovhlps— Move Packed Single-Precision Floating-Point Values High to Lowvmovhpd— Move High Packed Double-Precision Floating-Point Valuevmovhps— Move High Packed Single-Precision Floating-Point Valuesvmovlhps— Move Packed Single-Precision Floating-Point Values Low to Highvmovlpd— Move Low Packed Double-Precision Floating-Point Valuevmovlps— Move Low Packed Single-Precision Floating-Point Valuesvmovmskpd— Extract Packed Double-Precision Floating-Point Sign Maskvmovmskps— Extract Packed Single-Precision Floating-Point Sign Maskvmovntqqvmovq— Move Quadwordvmovqqavmovqquvmovsd— Move Scalar Double-Precision Floating-Point Valuevmovss— Move Scalar Single-Precision Floating-Point Valuesvmulsd— Multiply Scalar Double-Precision Floating-Point Valuesvmulss— Multiply Scalar Single-Precision Floating-Point Valuesvpand— Packed Bitwise Logical ANDvpandn— Packed Bitwise Logical AND NOTvpblendvb— Variable Blend Packed Bytesvpblendw— Blend Packed Wordsvpcmpestri— Packed Compare Explicit Length Strings, Return Indexvpcmpestrm— Packed Compare Explicit Length Strings, Return Maskvpcmpistri— Packed Compare Implicit Length Strings, Return Indexvpcmpistrm— Packed Compare Implicit Length Strings, Return Maskvperm2f128— Permute Floating-Point Valuesvpextrd— Extract Doublewordvpextrq— Extract Quadwordvphaddd— Packed Horizontal Add Doubleword Integervphaddsw— Packed Horizontal Add Signed Word Integers with Signed Saturationvphaddw— Packed Horizontal Add Word Integersvphminposuw— Packed Horizontal Minimum of Unsigned Word Integersvphsubd— Packed Horizontal Subtract Doubleword Integersvphsubsw— Packed Horizontal Subtract Signed Word Integers with Signed Saturationvphsubw— Packed Horizontal Subtract Word Integersvpinsrb— Insert Bytevpinsrd— Insert Doublewordvpinsrq— Insert Quadwordvpinsrw— Insert Wordvpmovmskb— Move Byte Maskvpor— Packed Bitwise Logical ORvpsignb— Packed Sign of Byte Integersvpsignd— Packed Sign of Doubleword Integersvpsignw— Packed Sign of Word Integersvptest— Packed Logical Comparevpxor— Packed Bitwise Logical Exclusive ORvrcpps— Compute Approximate Reciprocals of Packed Single-Precision Floating-Point Valuesvrcpss— Compute Approximate Reciprocal of Scalar Single-Precision Floating-Point Valuesvroundpd— Round Packed Double Precision Floating-Point Valuesvroundps— Round Packed Single Precision Floating-Point Valuesvroundsd— Round Scalar Double Precision Floating-Point Valuesvroundss— Round Scalar Single Precision Floating-Point Valuesvrsqrtps— Compute Reciprocals of Square Roots of Packed Single-Precision Floating-Point Valuesvrsqrtss— Compute Reciprocal of Square Root of Scalar Single-Precision Floating-Point Valuevsqrtsd— Compute Square Root of Scalar Double-Precision Floating-Point Valuevsqrtss— Compute Square Root of Scalar Single-Precision Floating-Point Valuevstmxcsr— Store MXCSR Register Statevsubsd— Subtract Scalar Double-Precision Floating-Point Valuesvsubss— Subtract Scalar Single-Precision Floating-Point Valuesvtestpd— Packed Double-Precision Floating-Point Bit Testvtestps— Packed Single-Precision Floating-Point Bit Testvucomisd— Unordered Compare Scalar Double-Precision Floating-Point Values and Set EFLAGSvucomiss— Unordered Compare Scalar Single-Precision Floating-Point Values and Set EFLAGSvzeroall— Zero All YMM Registersvzeroupper— Zero Upper Bits of YMM Registers
Intel AVX2 instructions
7 instructionsvbroadcasti128— Broadcast 128 Bits of Integer Datavextracti128— Extract Packed Integer Valuesvinserti128— Insert Packed Integer Valuesvpblendd— Blend Packed Doublewordsvperm2i128— Permute 128-Bit Integer Valuesvpmaskmovd— Conditional Move Packed Doubleword Integersvpmaskmovq— Conditional Move Packed Quadword Integers
Intel AVX512-FP16 instructions
114 instructionsvaddph— Add Packed Half-Precision Floating-Point Valuesvaddsh— Add Scalar Half-Precision Floating-Point Valuesvcmpph— Compare Packed Half-Precision Floating-Point Valuesvcmpsh— Compare Scalar Half-Precision Floating-Point Valuesvcomish— Compare Scalar Ordered Half-Precision Floating-Point Values and Set EFLAGSvcvtdq2ph— Convert Packed Dword Integers to Packed Half-Precision FP Valuesvcvtpd2ph— Convert Packed Double-Precision FP Values to Packed Half-Precision FP Valuesvcvtph2dq— Convert Packed Half-Precision FP Values to Packed Dword Integersvcvtph2pd— Convert Packed Half-Precision FP Values to Packed Double-Precision FP Valuesvcvtph2ps— Convert Half-Precision FP Values to Single-Precision FP Valuesvcvtph2psx— Convert Half-Precision FP Values to Single-Precision FP Valuesvcvtph2qq— Convert Packed Half Precision Floating-Point Values to Packed Singed Quadword Integer Valuesvcvtph2udq— Convert Packed Half-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Valuesvcvtph2uqq— Convert Packed Half Precision Floating-Point Values to Packed Unsigned Quadword Integer Valuesvcvtph2uw— Convert Packed Half-Precision Floating-Point Values to Packed Unsigned Word Integer Valuesvcvtph2w— Convert Packed Half-Precision Floating-Point Values to Packed Word Integer Valuesvcvtps2ph— Convert Single-Precision FP value to Half-Precision FP valuevcvtps2phx— Convert Single-Precision FP value to Half-Precision FP valuevcvtqq2ph— Convert Packed Quadword Integers to Packed Half-Precision Floating-Point Valuesvcvtsd2sh— Convert Scalar Double-Precision FP Value to Scalar Half-Precision FP Valuevcvtsh2sd— Convert Scalar Half-Precision FP Value to Scalar Double-Precision FP Valuevcvtsh2si— Convert Scalar Half-Precision FP Value to Dword Integervcvtsh2ss— Convert Scalar Half-Precision FP Value to Scalar Double-Precision FP Valuevcvtsh2usi— Convert Scalar Half-Precision Floating-Point Value to Unsigned Doubleword Integervcvtsi2sh— Convert Dword Integer to Scalar Half-Precision FP Valuevcvtss2sh— Convert Scalar Single-Precision FP Value to Scalar Half-Precision FP Valuevcvttph2dq— Convert with Truncation Packed Half-Precision FP Values to Packed Dword Integersvcvttph2qq— Convert with Truncation Packed Half Precision Floating-Point Values to Packed Singed Quadword Integer Valuesvcvttph2udq— Convert with Truncation Packed Half-Precision Floating-Point Values to Packed Unsigned Doubleword Integer Valuesvcvttph2uqq— Convert with Truncation Packed Half Precision Floating-Point Values to Packed Unsigned Quadword Integer Valuesvcvttph2uw— Convert with Truncation Packed Half-Precision Floating-Point Values to Packed Unsigned Word Integer Valuesvcvttph2w— Convert with Truncation Packed Half-Precision Floating-Point Values to Packed Word Integer Valuesvcvttsh2si— Convert with Truncation Scalar Half-Precision FP Value to Dword Integervcvttsh2usi— Convert with Truncation Scalar Half-Precision Floating-Point Value to Unsigned Integervcvtudq2ph— Convert Packed Unsigned Doubleword Integers to Packed Half-Precision Floating-Point Valuesvcvtuqq2ph— Convert Packed Unsigned Quadword Integers to Packed Half-Precision Floating-Point Valuesvcvtusi2sh— Convert Unsigned Integer to Scalar Half-Precision Floating-Point Valuevcvtuw2ph— Convert Packed Unsigned Word Integers to Packed Half-Precision Floating-Point Valuesvcvtw2ph— Convert Packed Word Integers to Packed Half-Precision Floating-Point Valuesvdivph— Divide Packed Half-Precision Floating-Point Valuesvdivsh— Divide Scalar Half-Precision Floating-Point Valuesvendscalephvendscaleshvfcmaddcph— Fused Conjugate Multiply-Add of Complex Packed Half-Precision Floating-Point Valuesvfcmaddcsh— Fused Conjugate Multiply-Add of Complex Scalar Half-Precision Floating-Point Valuesvfcmulcpchvfcmulcsh— Fused Conjugate Multiply of Complex Scalar Half-Precision Floating-Point Valuesvfmadd132ph— Fused Multiply-Add of Packed Half-Precision Floating-Point Valuesvfmadd213ph— Fused Multiply-Add of Packed Half-Precision Floating-Point Valuesvfmadd231ph— Fused Multiply-Add of Packed Half-Precision Floating-Point Valuesvfmaddcph— Fused Multiply-Add of Complex Packed Half-Precision Floating-Point Valuesvfmaddcsh— Fused Multiply-Add of Complex Scalar Half-Precision Floating-Point Valuesvfmaddsub132ph— Fused Multiply-Alternating Add/Subtract of Packed Half-Precision Floating-Point Valuesvfmaddsub213ph— Fused Multiply-Alternating Add/Subtract of Packed Half-Precision Floating-Point Valuesvfmaddsub231ph— Fused Multiply-Alternating Add/Subtract of Packed Half-Precision Floating-Point Valuesvfmsub132ph— Fused Multiply-Subtract of Packed Half-Precision Floating-Point Valuesvfmsub213ph— Fused Multiply-Subtract of Packed Half-Precision Floating-Point Valuesvfmsub231ph— Fused Multiply-Subtract of Packed Half-Precision Floating-Point Valuesvfmsubadd132ph— Fused Multiply-Alternating Subtract/Add of Packed Half-Precision Floating-Point Valuesvfmsubadd213ph— Fused Multiply-Alternating Subtract/Add of Packed Half-Precision Floating-Point Valuesvfmsubadd231ph— Fused Multiply-Alternating Subtract/Add of Packed Half-Precision Floating-Point Valuesvfmulcpchvfmulcsh— Fused Multiply of Complex Scalar Half-Precision Floating-Point Valuesvfnmadd132ph— Fused Negative Multiply-Add of Packed Half-Precision Floating-Point Valuesvfnmadd213ph— Fused Negative Multiply-Add of Packed Half-Precision Floating-Point Valuesvfnmadd231ph— Fused Negative Multiply-Add of Packed Half-Precision Floating-Point Valuesvfnmsub132ph— Fused Negative Multiply-Subtract of Packed Half-Precision Floating-Point Valuesvfnmsub213ph— Fused Negative Multiply-Subtract of Packed Half-Precision Floating-Point Valuesvfnmsub231ph— Fused Negative Multiply-Subtract of Packed Half-Precision Floating-Point Valuesvfpclassph— Test Class of Packed Half-Precision Floating-Point Valuesvfpclasssh— Test Class of Scalar Half-Precision Floating-Point Valuevgetexpph— Extract Exponents of Packed Half-Precision Floating-Point Values as Half-Precision Floating-Point Valuesvgetexpsh— Extract Exponent of Scalar Half-Precision Floating-Point Value as Half-Precision Floating-Point Valuevgetmantph— Extract Normalized Mantissas from Packed Half-Precision Floating-Point Valuesvgetmantsh— Extract Normalized Mantissa from Scalar Half-Precision Floating-Point Valuevgetmaxphvgetmaxshvgetminphvgetminshvmovsh— Move Scalar Half-Precision Floating-Point Valuesvmovw— Move Wordvmulph— Multiply Packed Half-Precision Floating-Point Valuesvmulsh— Fused Multiply Scalar Half-Precision Floating-Point Valuesvpmadd132phvpmadd132shvpmadd213phvpmadd213shvpmadd231phvpmadd231shvpmsub132phvpmsub132shvpmsub213phvpmsub213shvpmsub231phvpmsub231shvpnmadd132shvpnmadd213shvpnmadd231shvpnmsub132shvpnmsub213shvpnmsub231shvrcpph— Compute Approximate Reciprocals of Packed Half-Precision Floating-Point Valuesvrcpsh— Compute Approximate Reciprocal of Scalar Half-Precision Floating-Point Valuesvreduceph— Perform Reduction Transformation on Packed Half-Precision Floating-Point Valuesvreducesh— Perform Reduction Transformation on a Scalar Half-Precision Floating-Point Valuevrsqrtph— Compute Reciprocals of Square Roots of Packed Half-Precision Floating-Point Valuesvrsqrtsh— Compute Reciprocal of Square Root of Scalar Half-Precision Floating-Point Valuevscalefph— Scale Packed Half-Precision Floating-Point Values With Half-Precision Floating-Point Valuesvscalefsh— Scale Scalar Half-Precision Floating-Point Value With a Half-Precision Floating-Point Valuevsqrtph— Compute Square Roots of Packed Half-Precision Floating-Point Valuesvsqrtsh— Compute Square Root of Scalar Half-Precision Floating-Point Valuevsubph— Subtract Packed Half-Precision Floating-Point Valuesvsubsh— Subtract Scalar Half-Precision Floating-Point Valuesvucomish— Unordered Compare Scalar Half-Precision Floating-Point Values and Set EFLAGS
Intel Carry-Less Multiplication instructions (CLMUL)
5 instructionspclmulhqhqdqpclmulhqlqdqpclmullqhqdqpclmullqlqdqpclmulqdq— Carry-Less Quadword Multiplication
Intel Control-Flow Enforcement Technology (CET)
14 instructionsclrssbsyendbr32endbr64— END (terminate) BRanch in 64-bit modeincsspdincsspqrdsspdrdsspqrstorsspsaveprevsspsetssbsywrssdwrssqwrussdwrussq
Intel Fused Multiply-Add instructions (FMA)
84 instructionsvfmadd123pdvfmadd123psvfmadd123sdvfmadd123ssvfmadd132sd— Fused Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfmadd132ss— Fused Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfmadd213sd— Fused Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfmadd213ss— Fused Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfmadd231sd— Fused Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfmadd231ss— Fused Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfmadd312pdvfmadd312psvfmadd312sdvfmadd312ssvfmadd321pdvfmadd321psvfmadd321sdvfmadd321ssvfmaddsub123pdvfmaddsub123psvfmaddsub312pdvfmaddsub312psvfmaddsub321pdvfmaddsub321psvfmsub123pdvfmsub123psvfmsub123sdvfmsub123ssvfmsub132sd— Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfmsub132ss— Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfmsub213sd— Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfmsub213ss— Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfmsub231sd— Fused Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfmsub231ss— Fused Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfmsub312pdvfmsub312psvfmsub312sdvfmsub312ssvfmsub321pdvfmsub321psvfmsub321sdvfmsub321ssvfmsubadd123pdvfmsubadd123psvfmsubadd312pdvfmsubadd312psvfmsubadd321pdvfmsubadd321psvfnmadd123pdvfnmadd123psvfnmadd123sdvfnmadd123ssvfnmadd132sd— Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfnmadd132ss— Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfnmadd213sd— Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfnmadd213ss— Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfnmadd231sd— Fused Negative Multiply-Add of Scalar Double-Precision Floating-Point Valuesvfnmadd231ss— Fused Negative Multiply-Add of Scalar Single-Precision Floating-Point Valuesvfnmadd312pdvfnmadd312psvfnmadd312sdvfnmadd312ssvfnmadd321pdvfnmadd321psvfnmadd321sdvfnmadd321ssvfnmsub123pdvfnmsub123psvfnmsub123sdvfnmsub123ssvfnmsub132sd— Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfnmsub132ss— Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfnmsub213sd— Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfnmsub213ss— Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfnmsub231sd— Fused Negative Multiply-Subtract of Scalar Double-Precision Floating-Point Valuesvfnmsub231ss— Fused Negative Multiply-Subtract of Scalar Single-Precision Floating-Point Valuesvfnmsub312pdvfnmsub312psvfnmsub312sdvfnmsub312ssvfnmsub321pdvfnmsub321psvfnmsub321sdvfnmsub321ss
Intel instruction extension based on pub number 319433-030 dated October 2017
4 instructionsvaesdec— Perform One Round of an AES Decryption Flowvaesdeclast— Perform Last Round of an AES Decryption Flowvaesenc— Perform One Round of an AES Encryption Flowvaesenclast— Perform Last Round of an AES Encryption Flow
Intel Memory Protection Extensions (MPX)
7 instructionsbndclbndcnbndcubndldxbndmkbndmovbndstx
Intel memory protection keys for userspace (PKU aka PKEYs)
2 instructionsrdpkruwrpkru
Intel SHA acceleration instructions
10 instructionssha1msg1— Perform an Intermediate Calculation for the Next Four SHA1 Message Doublewordssha1msg2— Perform a Final Calculation for the Next Four SHA1 Message Doublewordssha1nexte— Calculate SHA1 State Variable E after Four Roundssha1rnds4— Perform Four Rounds of SHA1 Operationsha256msg1— Perform an Intermediate Calculation for the Next Four SHA256 Message Doublewordssha256msg2— Perform a Final Calculation for the Next Four SHA256 Message Doublewordssha256rnds2— Perform Two Rounds of SHA256 Operationvsha512msg1— Perform an Intermediate Calculation for the Next Four SHA512 Message Quadwordsvsha512msg2— Perform a Final Calculation for the Next Four SHA512 Message Quadwordsvsha512rnds2— Perform Two Rounds of SHA512 Operation
Intel SMX
1 instructiongetsec
Intel Software Guard Extensions (SGX)
3 instructionsenclsencluenclv
Intel Transactional Synchronization Extensions (TSX)
5 instructionsprefetchwt1— Prefetch Vector Data Into Caches with Intent to Write and T1 Hintxabortxbeginxendxtest
Interleaved flags arithmetic
2 instructionsadcx— Unsigned Integer Addition of Two Operands with Carry Flagadox— Unsigned Integer Addition of Two Operands with Overflow Flag
Interrupts, system calls, and returns (extensions)
2 instructionserets(FRED)eretu(FRED)
Introduced in Deschutes but necessary for SSE support
4 instructionsfxrstorfxrstor64fxsavefxsave64
Jumps (extensions)
1 instructionjmpabs(APX)
Katmai Streaming SIMD instructions (SSE -- a.k.a. KNI, XMM, MMX2)
62 instructionsaddps— Add Packed Single-Precision Floating-Point Valuesaddss— Add Scalar Single-Precision Floating-Point Valuesandnps— Bitwise Logical AND NOT of Packed Single-Precision Floating-Point Valuesandps— Bitwise Logical AND of Packed Single-Precision Floating-Point Valuescmpeqpscmpeqsscmplepscmplesscmpltpscmpltsscmpneqpscmpneqsscmpnlepscmpnlesscmpnltpscmpnltsscmpordpscmpordsscmpps— Compare Packed Single-Precision Floating-Point Valuescmpss— Compare Scalar Single-Precision Floating-Point Valuescmpunordpscmpunordsscomiss— Compare Scalar Ordered Single-Precision Floating-Point Values and Set EFLAGScvtpi2ps— Convert Packed Dword Integers to Packed Single-Precision FP Valuescvtps2pi— Convert Packed Single-Precision FP Values to Packed Dword Integerscvtsi2ss— Convert Dword Integer to Scalar Single-Precision FP Valuecvtss2si— Convert Scalar Single-Precision FP Value to Dword Integercvttps2pi— Convert with Truncation Packed Single-Precision FP Values to Packed Dword Integerscvttss2si— Convert with Truncation Scalar Single-Precision FP Value to Dword Integerdivps— Divide Packed Single-Precision Floating-Point Valuesdivss— Divide Scalar Single-Precision Floating-Point Valuesldmxcsr— Load MXCSR Registermaxps— Return Maximum Packed Single-Precision Floating-Point Valuesmaxss— Return Maximum Scalar Single-Precision Floating-Point Valueminps— Return Minimum Packed Single-Precision Floating-Point Valuesminss— Return Minimum Scalar Single-Precision Floating-Point Valuemovaps— Move Aligned Packed Single-Precision Floating-Point Valuesmovhlps— Move Packed Single-Precision Floating-Point Values High to Lowmovhps— Move High Packed Single-Precision Floating-Point Valuesmovlhps— Move Packed Single-Precision Floating-Point Values Low to Highmovlps— Move Low Packed Single-Precision Floating-Point Valuesmovmskps— Extract Packed Single-Precision Floating-Point Sign Maskmovntps— Store Packed Single-Precision Floating-Point Values Using Non-Temporal Hintmovss— Move Scalar Single-Precision Floating-Point Valuesmovups— Move Unaligned Packed Single-Precision Floating-Point Valuesmulps— Multiply Packed Single-Precision Floating-Point Valuesmulss— Multiply Scalar Single-Precision Floating-Point Valuesorps— Bitwise Logical OR of Single-Precision Floating-Point Valuesrcpps— Compute Approximate Reciprocals of Packed Single-Precision Floating-Point Valuesrcpss— Compute Approximate Reciprocal of Scalar Single-Precision Floating-Point Valuesrsqrtps— Compute Reciprocals of Square Roots of Packed Single-Precision Floating-Point Valuesrsqrtss— Compute Reciprocal of Square Root of Scalar Single-Precision Floating-Point Valueshufps— Shuffle Packed Single-Precision Floating-Point Valuessqrtps— Compute Square Roots of Packed Single-Precision Floating-Point Valuessqrtss— Compute Square Root of Scalar Single-Precision Floating-Point Valuestmxcsr— Store MXCSR Register Statesubps— Subtract Packed Single-Precision Floating-Point Valuessubss— Subtract Scalar Single-Precision Floating-Point Valuesucomiss— Unordered Compare Scalar Single-Precision Floating-Point Values and Set EFLAGSunpckhps— Unpack and Interleave High Packed Single-Precision Floating-Point Valuesunpcklps— Unpack and Interleave Low Packed Single-Precision Floating-Point Valuesxorps— Bitwise Logical XOR for Single-Precision Floating-Point Values
Machine control and management instructions
20 instructionsbb0_resetbb1_resetcltscpu_readcpu_writecpuid— CPU Identificationdmintlmswrdmrdmsrrdmsrlistsmintsmintoldsmswumovurdmsruwrmsrwrmsrwrmsrlistwrmsrns
Memory management and control
11 instructionsclflush— Flush Cache Lineclflushopt— Flush Cache Line Optimizedclwb— Cache Line Write Backclzero— Zero-out 64-bit Cache Lineinvdinvlpginvlpgainvpcidpcommitwbinvdwbnoinvd
MMX (SIMD using the x87 register file)
78 instructionsemms— Exit MMX Statemovd— Move Doublewordpackssdw— Pack Doublewords into Words with Signed Saturationpacksswb— Pack Words into Bytes with Signed Saturationpackuswb— Pack Words into Bytes with Unsigned Saturationpaddb— Add Packed Byte Integerspaddd— Add Packed Doubleword Integerspaddsb— Add Packed Signed Byte Integers with Signed Saturationpaddsiwpaddsw— Add Packed Signed Word Integers with Signed Saturationpaddusb— Add Packed Unsigned Byte Integers with Unsigned Saturationpaddusw— Add Packed Unsigned Word Integers with Unsigned Saturationpaddw— Add Packed Word Integerspand— Packed Bitwise Logical ANDpandn— Packed Bitwise Logical AND NOTpavebpavgusb— Average Packed Byte Integerspcmpeqb— Compare Packed Byte Data for Equalitypcmpeqd— Compare Packed Doubleword Data for Equalitypcmpeqw— Compare Packed Word Data for Equalitypcmpgtb— Compare Packed Signed Byte Integers for Greater Thanpcmpgtd— Compare Packed Signed Doubleword Integers for Greater Thanpcmpgtw— Compare Packed Signed Word Integers for Greater Thanpdistibpf2id— Packed Floating-Point to Integer Doubleword Conversonpfacc— Packed Floating-Point Accumulatepfadd— Packed Floating-Point Addpfcmpeq— Packed Floating-Point Compare for Equalpfcmpge— Packed Floating-Point Compare for Greater or Equalpfcmpgt— Packed Floating-Point Compare for Greater Thanpfmax— Packed Floating-Point Maximumpfmin— Packed Floating-Point Minimumpfmul— Packed Floating-Point Multiplypfrcp— Packed Floating-Point Reciprocal Approximationpfrcpit1— Packed Floating-Point Reciprocal Iteration 1pfrcpit2— Packed Floating-Point Reciprocal Iteration 2pfrsqit1— Packed Floating-Point Reciprocal Square Root Iteration 1pfrsqrt— Packed Floating-Point Reciprocal Square Root Approximationpfsub— Packed Floating-Point Subtractpfsubr— Packed Floating-Point Subtract Reversepi2fd— Packed Integer to Floating-Point Doubleword Conversionpmachriwpmaddwd— Multiply and Add Packed Signed Word Integerspmagwpmulhriwpmulhrwapmulhrwcpmulhw— Multiply Packed Signed Word Integers and Store High Resultpmullw— Multiply Packed Signed Word Integers and Store Low Resultpmvgezbpmvlzbpmvnzbpmvzbpor— Packed Bitwise Logical ORprefetch— Prefetch Data into Cachesprefetchw— Prefetch Data into Caches in Anticipation of a Writepslld— Shift Packed Doubleword Data Left Logicalpsllq— Shift Packed Quadword Data Left Logicalpsllw— Shift Packed Word Data Left Logicalpsrad— Shift Packed Doubleword Data Right Arithmeticpsraw— Shift Packed Word Data Right Arithmeticpsrld— Shift Packed Doubleword Data Right Logicalpsrlq— Shift Packed Quadword Data Right Logicalpsrlw— Shift Packed Word Data Right Logicalpsubb— Subtract Packed Byte Integerspsubd— Subtract Packed Doubleword Integerspsubsb— Subtract Packed Signed Byte Integers with Signed Saturationpsubsiwpsubsw— Subtract Packed Signed Word Integers with Signed Saturationpsubusb— Subtract Packed Unsigned Byte Integers with Unsigned Saturationpsubusw— Subtract Packed Unsigned Word Integers with Unsigned Saturationpsubw— Subtract Packed Word Integerspunpckhbw— Unpack and Interleave High-Order Bytes into Wordspunpckhdq— Unpack and Interleave High-Order Doublewords into Quadwordspunpckhwd— Unpack and Interleave High-Order Words into Doublewordspunpcklbw— Unpack and Interleave Low-Order Bytes into Wordspunpckldq— Unpack and Interleave Low-Order Doublewords into Quadwordspunpcklwd— Unpack and Interleave Low-Order Words into Doublewords
MMX instructions
2 instructionspxor— Packed Bitwise Logical Exclusive ORskinit
Nehalem New Instructions (SSE4.2)
7 instructionscrc32— Accumulate CRC32 Valuepcmpestri— Packed Compare Explicit Length Strings, Return Indexpcmpestrm— Packed Compare Explicit Length Strings, Return Maskpcmpgtq— Compare Packed Data for Greater Thanpcmpistri— Packed Compare Implicit Length Strings, Return Indexpcmpistrm— Packed Compare Implicit Length Strings, Return Maskpopcnt— Count of Number of Bits Set to 1
New MMX instructions introduced in Katmai
12 instructionsmaskmovq— Store Selected Bytes of Quadwordmovntq— Store of Quadword Using Non-Temporal Hintpavgb— Average Packed Byte Integerspavgw— Average Packed Word Integerspmaxsw— Maximum of Packed Signed Word Integerspmaxub— Maximum of Packed Unsigned Byte Integerspminsw— Minimum of Packed Signed Word Integerspminub— Minimum of Packed Unsigned Byte Integerspmovmskb— Move Byte Maskpmulhuw— Multiply Packed Unsigned Word Integers and Store High Resultpsadbw— Compute Sum of Absolute Differencespshufw— Shuffle Packed Words
Other basic integer arithmetic (extensions)
1 instructionmulx— Unsigned Multiply Without Affecting Flags (BMI2)
Penryn New Instructions (SSE4.1)
49 instructionsblendpd— Blend Packed Double Precision Floating-Point Valuesblendps— Blend Packed Single Precision Floating-Point Valuesblendvpd— Variable Blend Packed Double Precision Floating-Point Valuesblendvps— Variable Blend Packed Single Precision Floating-Point Valuesdppd— Dot Product of Packed Double Precision Floating-Point Valuesdpps— Dot Product of Packed Single Precision Floating-Point Valuesextractps— Extract Packed Single Precision Floating-Point Valueinsertps— Insert Packed Single Precision Floating-Point Valuemovntdqa— Load Double Quadword Non-Temporal Aligned Hintmpsadbw— Compute Multiple Packed Sums of Absolute Differencepackusdw— Pack Doublewords into Words with Unsigned Saturationpblendvb— Variable Blend Packed Bytespblendw— Blend Packed Wordspcmpeqq— Compare Packed Quadword Data for Equalitypextrb— Extract Bytepextrd— Extract Doublewordpextrq— Extract Quadwordpextrw— Extract Wordphminposuw— Packed Horizontal Minimum of Unsigned Word Integerspinsrb— Insert Bytepinsrd— Insert Doublewordpinsrq— Insert Quadwordpmaxsb— Maximum of Packed Signed Byte Integerspmaxsd— Maximum of Packed Signed Doubleword Integerspmaxud— Maximum of Packed Unsigned Doubleword Integerspmaxuw— Maximum of Packed Unsigned Word Integerspminsb— Minimum of Packed Signed Byte Integerspminsd— Minimum of Packed Signed Doubleword Integerspminud— Minimum of Packed Unsigned Doubleword Integerspminuw— Minimum of Packed Unsigned Word Integerspmovsxbd— Move Packed Byte Integers to Doubleword Integers with Sign Extensionpmovsxbq— Move Packed Byte Integers to Quadword Integers with Sign Extensionpmovsxbw— Move Packed Byte Integers to Word Integers with Sign Extensionpmovsxdq— Move Packed Doubleword Integers to Quadword Integers with Sign Extensionpmovsxwd— Move Packed Word Integers to Doubleword Integers with Sign Extensionpmovsxwq— Move Packed Word Integers to Quadword Integers with Sign Extensionpmovzxbd— Move Packed Byte Integers to Doubleword Integers with Zero Extensionpmovzxbq— Move Packed Byte Integers to Quadword Integers with Zero Extensionpmovzxbw— Move Packed Byte Integers to Word Integers with Zero Extensionpmovzxdq— Move Packed Doubleword Integers to Quadword Integers with Zero Extensionpmovzxwd— Move Packed Word Integers to Doubleword Integers with Zero Extensionpmovzxwq— Move Packed Word Integers to Quadword Integers with Zero Extensionpmuldq— Multiply Packed Signed Doubleword Integers and Store Quadword Resultpmulld— Multiply Packed Signed Doubleword Integers and Store Low Resultptest— Packed Logical Compareroundpd— Round Packed Double Precision Floating-Point Valuesroundps— Round Packed Single Precision Floating-Point Valuesroundsd— Round Scalar Double Precision Floating-Point Valuesroundss— Round Scalar Single Precision Floating-Point Values
Permanently undefined instructions
67 instructionsccmpccmpaccmpaeccmpbccmpbeccmpcccmpeccmpfccmpgccmpgeccmplccmpleccmpnaccmpnaeccmpnbccmpnbeccmpncccmpneccmpngccmpngeccmpnlccmpnleccmpnoccmpnsccmpnzccmpoccmpsccmptccmpzctestctestactestaectestbctestbectestcctestectestfctestgctestgectestlctestlectestnactestnaectestnbctestnbectestncctestnectestngctestngectestnlctestnlectestnoctestnsctestnzctestoctestsctesttctestzfwaitud0ud1ud2— Undefined Instructionud2aud2budbxlatxlatb— Table Look-up Translation
Power management
12 instructionshltmonitor— Monitor a Linear Address Rangemonitordmonitorqmonitorwmonitorx— Monitor a Linear Address Range with Timeoutmwait— Monitor Waitmwaitx— Monitor Wait with Timeoutpause— Spin Loop Hinttpause— Timed PAUSEumonitor— User mode Monitor a Linear Address Rangeumwait— User mode Monitor Wait
Prescott New Instructions (SSE3)
10 instructionsaddsubpd— Packed Double-FP Add/Subtractaddsubps— Packed Single-FP Add/Subtracthaddpd— Packed Double-FP Horizontal Addhaddps— Packed Single-FP Horizontal Addhsubpd— Packed Double-FP Horizontal Subtracthsubps— Packed Single-FP Horizontal Subtractlddqu— Load Unaligned Integer 128 Bitsmovddup— Move One Double-FP and Duplicatemovshdup— Move Packed Single-FP High and Duplicatemovsldup— Move Packed Single-FP Low and Duplicate
Processor trace write
1 instructionptwrite
RAO-INT weakly ordered atomic operations
4 instructionsaadd— Atomically ADDaand— Atomically ANDaor— Atomically ORaxor— Atomically XOR
S3M hash instructions
3 instructionsvsm3msg1— Perform Initial Calculation for the Next Four SM3 Message Wordsvsm3msg2— Perform Final Calculation for the Next Four SM3 Message Wordsvsm3rnds2— Perform Two Rounds of SM3 Operation
Segment handling instructions
26 instructionsarpllarldsleslfslgdtlgslidtlkgslldtloadallloadall286lsllssltrrdfsbase— ReaD FS segment BASErdgsbase— ReaD GS segment BASEsgdtsidtsldtstrswapgsverrverwwrfsbase— WRite FS segment BASEwrgsbase— WRite GS segment BASE
SEV-SNP AMD instructions
3 instructionspvalidatermpadjustvmgexit
SM4 hash instructions
2 instructionsvsm4key4— Perform Four Rounds of SM4 Key Expansionvsm4rnds4— Performs Four Rounds of SM4 Encryption
Special reads: timestamp, CPU number, performance counters, randomness
6 instructionsrdpid— Read Processor IDrdpmc— Read Performance-Monitoring Counterrdrand— Read Random Numberrdseed— Read Random SEEDrdtsc— Read Time-Stamp Counterrdtscp— Read Time-Stamp Counter and Processor ID
Stack operations (extensions)
6 instructionspop2(APX)pop2p(APX)popp(APX)push2(APX)push2p(APX)pushp(APX)
Synchronization and fencing
4 instructionslfence— Load Fencemfence— Memory Fenceserialize— Serialize Instruction Executionsfence— Store Fence
System management mode
9 instructionsrdshrrsdcrsldtrsmrstssvdcsvldtsvtswrshr
Systematic names for the hinting nop instructions
65 instructionshint_nophint_nop0hint_nop1hint_nop10hint_nop11hint_nop12hint_nop13hint_nop14hint_nop15hint_nop16hint_nop17hint_nop18hint_nop19hint_nop2hint_nop20hint_nop21hint_nop22hint_nop23hint_nop24hint_nop25hint_nop26hint_nop27hint_nop28hint_nop29hint_nop3hint_nop30hint_nop31hint_nop32hint_nop33hint_nop34hint_nop35hint_nop36hint_nop37hint_nop38hint_nop39hint_nop4hint_nop40hint_nop41hint_nop42hint_nop43hint_nop44hint_nop45hint_nop46hint_nop47hint_nop48hint_nop49hint_nop5hint_nop50hint_nop51hint_nop52hint_nop53hint_nop54hint_nop55hint_nop56hint_nop57hint_nop58hint_nop59hint_nop6hint_nop60hint_nop61hint_nop62hint_nop63hint_nop7hint_nop8hint_nop9
Tejas New Instructions (SSSE3)
16 instructionspabsb— Packed Absolute Value of Byte Integerspabsd— Packed Absolute Value of Doubleword Integerspabsw— Packed Absolute Value of Word Integerspalignr— Packed Align Rightphaddd— Packed Horizontal Add Doubleword Integerphaddsw— Packed Horizontal Add Signed Word Integers with Signed Saturationphaddw— Packed Horizontal Add Word Integersphsubd— Packed Horizontal Subtract Doubleword Integersphsubsw— Packed Horizontal Subtract Signed Word Integers with Signed Saturationphsubw— Packed Horizontal Subtract Word Integerspmaddubsw— Multiply and Add Packed Signed and Unsigned Byte Integerspmulhrsw— Packed Multiply Signed Word Integers and Store High Result with Round and Scalepshufb— Packed Shuffle Bytespsignb— Packed Sign of Byte Integerspsignd— Packed Sign of Doubleword Integerspsignw— Packed Sign of Word Integers
The basic shift and rotate operations (extensions)
6 instructionsrolx(BMI2)rorx— Rotate Right Logical Without Affecting Flags (BMI2)salx(BMI2)sarx— Arithmetic Shift Right Without Affecting Flags (BMI2)shlx— Logical Shift Left Without Affecting Flags (BMI2)shrx— Logical Shift Right Without Affecting Flags (BMI2)
User interrupts
5 instructionscluisenduipistuitestuiuiret
VIA (Centaur) security instructions
9 instructionsmontmulxcryptcbcxcryptcfbxcryptctrxcryptecbxcryptofbxsha1xsha256xstore
VMX/SVM Instructions
17 instructionsclgistgivmcallvmclearvmfuncvmlaunchvmloadvmmcallvmptrldvmptrstvmreadvmresumevmrunvmsavevmwritevmxoffvmxon
Willamette MMX instructions (SSE2 SIMD Integer Instructions)
16 instructionsmovdq2q— Move Quadword from XMM to MMX Technology Registermovdqa— Move Aligned Double Quadwordmovdqu— Move Unaligned Double Quadwordmovq— Move Quadwordmovq2dq— Move Quadword from MMX Technology to XMM Registerpaddq— Add Packed Quadword Integerspinsrw— Insert Wordpmuludq— Multiply Packed Unsigned Doubleword Integerspshufd— Shuffle Packed Doublewordspshufhw— Shuffle Packed High Wordspshuflw— Shuffle Packed Low Wordspslldq— Shift Packed Double Quadword Left Logicalpsrldq— Shift Packed Double Quadword Right Logicalpsubq— Subtract Packed Quadword Integerspunpckhqdq— Unpack and Interleave High-Order Quadwords into Double Quadwordspunpcklqdq— Unpack and Interleave Low-Order Quadwords into Double Quadwords
Willamette SSE2 Cacheability Instructions
4 instructionsmaskmovdqu— Store Selected Bytes of Double Quadwordmovntdq— Store Double Quadword Using Non-Temporal Hintmovnti— Store Doubleword Using Non-Temporal Hintmovntpd— Store Packed Double-Precision Floating-Point Values Using Non-Temporal Hint
Willamette Streaming SIMD instructions (SSE2)
62 instructionsaddpd— Add Packed Double-Precision Floating-Point Valuesaddsd— Add Scalar Double-Precision Floating-Point Valuesandnpd— Bitwise Logical AND NOT of Packed Double-Precision Floating-Point Valuesandpd— Bitwise Logical AND of Packed Double-Precision Floating-Point Valuescmpeqpdcmpeqsdcmplepdcmplesdcmpltpdcmpltsdcmpneqpdcmpneqsdcmpnlepdcmpnlesdcmpnltpdcmpnltsdcmpordpdcmpordsdcmppd— Compare Packed Double-Precision Floating-Point Valuescmpunordpdcmpunordsdcomisd— Compare Scalar Ordered Double-Precision Floating-Point Values and Set EFLAGScvtdq2pd— Convert Packed Dword Integers to Packed Double-Precision FP Valuescvtdq2ps— Convert Packed Dword Integers to Packed Single-Precision FP Valuescvtpd2dq— Convert Packed Double-Precision FP Values to Packed Dword Integerscvtpd2pi— Convert Packed Double-Precision FP Values to Packed Dword Integerscvtpd2ps— Convert Packed Double-Precision FP Values to Packed Single-Precision FP Valuescvtpi2pd— Convert Packed Dword Integers to Packed Double-Precision FP Valuescvtps2dq— Convert Packed Single-Precision FP Values to Packed Dword Integerscvtps2pd— Convert Packed Single-Precision FP Values to Packed Double-Precision FP Valuescvtsd2si— Convert Scalar Double-Precision FP Value to Integercvtsd2ss— Convert Scalar Double-Precision FP Value to Scalar Single-Precision FP Valuecvtsi2sd— Convert Dword Integer to Scalar Double-Precision FP Valuecvtss2sd— Convert Scalar Single-Precision FP Value to Scalar Double-Precision FP Valuecvttpd2dq— Convert with Truncation Packed Double-Precision FP Values to Packed Dword Integerscvttpd2pi— Convert with Truncation Packed Double-Precision FP Values to Packed Dword Integerscvttps2dq— Convert with Truncation Packed Single-Precision FP Values to Packed Dword Integerscvttsd2si— Convert with Truncation Scalar Double-Precision FP Value to Signed Integerdivpd— Divide Packed Double-Precision Floating-Point Valuesdivsd— Divide Scalar Double-Precision Floating-Point Valuesmaxpd— Return Maximum Packed Double-Precision Floating-Point Valuesmaxsd— Return Maximum Scalar Double-Precision Floating-Point Valueminpd— Return Minimum Packed Double-Precision Floating-Point Valuesminsd— Return Minimum Scalar Double-Precision Floating-Point Valuemovapd— Move Aligned Packed Double-Precision Floating-Point Valuesmovhpd— Move High Packed Double-Precision Floating-Point Valuemovlpd— Move Low Packed Double-Precision Floating-Point Valuemovmskpd— Extract Packed Double-Precision Floating-Point Sign Maskmovsd— Move Scalar Double-Precision Floating-Point Valuemovupd— Move Unaligned Packed Double-Precision Floating-Point Valuesmulpd— Multiply Packed Double-Precision Floating-Point Valuesmulsd— Multiply Scalar Double-Precision Floating-Point Valuesorpd— Bitwise Logical OR of Double-Precision Floating-Point Valuesshufpd— Shuffle Packed Double-Precision Floating-Point Valuessqrtpd— Compute Square Roots of Packed Double-Precision Floating-Point Valuessqrtsd— Compute Square Root of Scalar Double-Precision Floating-Point Valuesubpd— Subtract Packed Double-Precision Floating-Point Valuessubsd— Subtract Scalar Double-Precision Floating-Point Valuesucomisd— Unordered Compare Scalar Double-Precision Floating-Point Values and Set EFLAGSunpckhpd— Unpack and Interleave High Packed Double-Precision Floating-Point Valuesunpcklpd— Unpack and Interleave Low Packed Double-Precision Floating-Point Valuesxorpd— Bitwise Logical XOR for Double-Precision Floating-Point Values
x87 floating point
99 instructionsf2xm1fabsfaddfaddpfbldfbstpfchsfclexfcmovbfcmovbefcmovefcmovnbfcmovnbefcmovnefcmovnufcmovufcomfcomifcomipfcompfcomppfcosfdecstpfdisifdivfdivpfdivrfdivrpfemms— Fast Exit Multimedia Statefeniffreeffreepfiaddficomficompfidivfidivrfildfimulfincstpfinitfistfistpfisttpfisubfisubrfldfld1fldcwfldenvfldl2efldl2tfldlg2fldln2fldpifldzfmulfmulpfnclexfndisifnenifninitfnopfnsavefnstcwfnstenvfnstswfpatanfpremfprem1fptanfrndintfrstorfsavefscalefsetpmfsinfsincosfsqrtfstfstcwfstenvfstpfstswfsubfsubpfsubrfsubrpftstfucomfucomifucomipfucompfucomppfxamfxchfxtractfyl2xfyl2xp1
XSAVE group (AVX and extended state)
14 instructionsxgetbv— Get Value of Extended Control Registerxrstorxrstor64xrstorsxrstors64xsavexsave64xsavecxsavec64xsaveoptxsaveopt64xsavesxsaves64xsetbv