Introduction to Intel AES-NI
Contents
Translation
This post is also available in Simplified Chinese.
Notice
This article was originally published on Anquanke: original article
AES-NI is Intel’s x86-64 SIMD extension for accelerating AES. Anyone familiar with SIMD probably knows it exists, but AES’s asymmetric structure and AES-NI’s unusual design require considerable detail and theory to use correctly. Using easyRE from N1CTF 2021 as an example, this article summarizes my understanding; corrections are welcome.
AES Structure
AES-128 is a ten-round 4×4 substitution-permutation network whose final round omits MixColumns.

Although AES has ten rounds, an AddRoundKey occurs before round one, so there are eleven round keys. Encryption begins and ends with AddRoundKey, a design called key whitening. The reason is straightforward: the other three operations are fixed and key-independent. At either boundary anyone could invert them, so they would add no security.
AESENC and AESENCLAST
These are AES-NI’s encryption instructions and the easiest to understand. The Intel Intrinsics Guide shows that AESENC applies ShiftRows, SubBytes, MixColumns, and AddRoundKey. Since SubBytes operates independently on bytes, it commutes with ShiftRows; AESENC therefore implements one ordinary round in the diagram.
AESENCLAST applies ShiftRows, SubBytes, and AddRoundKey, implementing the final round.
The initial round-key XOR uses PXOR, giving the following complete encryption (pt is plaintext, k[x] a round key, and ct ciphertext):
| |
Nine AESENC rounds plus one AESENCLAST are easy to remember; the easily overlooked detail is the direct PXOR with round key zero.
AES Decryption and Equivalent Inverse Cipher
AES’s asymmetric design is deceptive. The decryption side of the diagram also consists of whitening, nine ordinary rounds, and one final round.
Naively reversing encryption suggests decrypting the final round first and ordinary rounds afterward, yet the diagram clearly does not do that.
Ignoring round boundaries, decryption performs the four inverse operations in reverse order. There are many ways to group that sequence into rounds, however. Under the conventional grouping shown above, a decryption round is not the inverse of an encryption round—AES’s first counterintuitive feature.
Here a decryption round contains InvShiftRows, InvSubBytes, AddRoundKey, and InvMixColumns; the final round again omits the MixColumns operation.
AES was originally Rijndael. Its original proposal also defines an “equivalent inverse cipher” (section 5.3.3), swapping AddRoundKey and InvMixColumns so decryption mirrors encryption with AddRoundKey last. InvSubBytes and InvShiftRows themselves commute.
This swap is not inherently equivalent. InvMixColumns multiplies each four-byte column by a 4×4 matrix over GF(2^8), while AddRoundKey XORs each byte. XOR is addition in GF(2^8); distributivity therefore requires multiplying the round key by the same matrix when moving AddRoundKey after InvMixColumns.
Round key zero is XORed directly and the last key belongs to the final decryption round, so neither participates in the swap. Thus equivalent decryption reverses the encryption keys and applies InvMixColumns to keys 1 through n−1, producing decryption keys.
Encryption and equivalent-decryption rounds have an elegant symmetry but use different round keys—AES’s second counterintuitive feature.
AESDEC, AESDECLAST, and AESIMC
Intel also uses equivalent decryption, as described in the AES-NI white paper. AESDEC is not the inverse of AESENC, nor is AESDECLAST the inverse of AESENCLAST. Complete decryption is:
| |
Here k[0] and k[n] are unchanged, while k'[1] through k'[n-1] are encryption keys transformed by InvMixColumns. Intel provides AESIMC specifically to perform this single operation.
AESKEYGENASSIST and PCLMULQDQ
AESKEYGENASSIST supports key expansion; see page 19 of the white paper.
PCLMULQDQ, Carry-Less Multiplication Quadword, multiplies two polynomials over GF(2^128). It is not formally part of AES-NI, but besides accelerating CRC32 it computes GCM’s GMAC and therefore often appears in SIMD cryptography. Libsodium’s AES-256-GCM implementation is an excellent example.
Advanced AES-NI Uses
Isolating AES Operations
I initially wondered why Intel chose equivalent decryption, which requires an extra AESIMC step to generate decryption keys. The white paper revealed the elegance of the design.
Page 34 shows how AES-NI can isolate the individual AES operations:
| |
ShiftRows is directly expressible with SSSE3 PSHUFB. SubBytes reverses the shuffle and performs a final round with a zero key, canceling the other two operations. MixColumns combines encryption and decryption so the final-round behavior leaves only MixColumns. The composition is remarkably clever.
AESIMC converts an encryption key to an equivalent-decryption key. To reverse this without a direct MixColumns instruction, combine AESDECLAST and AESENC as shown above.
The Intrinsics Guide shows that on Skylake, AESIMC has twice the latency and reciprocal throughput of AESENC. I suspect it internally composes AESENCLAST and AESDEC similarly.
Accelerating Other Algorithms
AES-NI’s flexibility supports larger substitution-permutation networks. The Rijndael proposal defines 128-, 192-, and 256-bit block sizes—not key sizes—but only Rijndael-128 became AES. Intel’s white paper implements other variants, including Rijndael-256:
| |
Rijndael-256 is an 8×4 network. Byte-level SubBytes and AddRoundKey work normally, as does per-column MixColumns; only ShiftRows needs SSE4.1 PBLENDVB and SSSE3 PSHUFB to adjust offsets. Since 8×4 is twice 4×4, every round uses two AESENC instructions and ends with two AESENCLASTs. This orderly, elegant code is part of what makes computers fascinating.
SM4’s nonlinear transform τ is also an S-box over GF(2^8), differing from AES only in its generating polynomial. The two fields are isomorphic, so algebraic transformations map their elements. Markku-Juhani O. Saarinen used this to accelerate SM4 with AES-NI; see sm4ni.
N1CTF 2021 easyRe
The challenge is available here. Its encryption function repeatedly encrypts and shuffles the plaintext in xmm0:
Although the program uses V-prefixed AVX2 instructions, it touches only XMM registers, so decryption can use SSE alone. Parse the function with Capstone into an expression tree whose input is a leaf and ciphertext is the root. Rotate the tree left and right until the input becomes the root; the resulting expression is the decryption formula.
During rotation, inverses of VPXOR, VPADDQ, and VPSUBQ are straightforward. VPSHUFD permutes four 32-bit XMM elements and can permute them back. For VAESENC, extract the entire VAESENC/VAESENCLAST block, invert the intermediate round keys with VAESIMC, and generate the reverse decryption tree. Since AES’s key zero is applied by VPXOR, a VAESENC not preceded by VPXOR can be treated as XOR with an all-zero key, making the final VAESDECLAST key zero. Handle VAESDEC blocks similarly, using the VAESDECLAST plus VAESENC composition above to obtain MixColumns and transform round keys.
I wrote a JIT from the expression tree; compiling and running its generated code produces the flag:
| |
Compile with -maes to enable AES-NI and produce SSE code. Adding -march=native enables more extensions and can automatically produce AVX2 plus VAES code; modern compilers are remarkably capable.
Flag: n1ctf{Easy_AVX!}—not easy at all.