When writing shellcode or analyzing crash dumps, you frequently encounter raw hex bytes like \x31\xc0, \x66\x50, or \xcd\x80. Understanding what these bytes mean — and how to find them — is a fundamental skill in binary exploitation.

This post explains how x86 instructions are encoded into machine code.

Instruction Format

Every x86 instruction is encoded as a sequence of 1–15 bytes with the following structure:

[Prefixes] [Opcode] [ModR/M] [SIB] [Displacement] [Immediate]

Not every field is present in every instruction. Simple instructions like nop are a single byte. Complex memory operations may use all six fields.

Field Size Purpose
Prefix 0–4 bytes Modifies the instruction’s behavior (operand size, segment, etc.)
Opcode 1–3 bytes The actual operation (push, mov, xor, etc.)
ModR/M 0–1 bytes Specifies registers and addressing modes
SIB 0–1 bytes Scale-Index-Base for complex memory addressing
Displacement 0–4 bytes Memory offset
Immediate 0–4 bytes Constant value embedded in the instruction

Common Instructions and Their Encodings

Here are the instructions most frequently used in shellcode, with their raw byte encodings:

Assembly Hex Size Notes
nop 90 1 byte No operation
push eax 50 1 byte Opcode 50+r where EAX = register 0
push ecx 51 1 byte ECX = register 1
pop eax 58 1 byte Opcode 58+r
push ax 66 50 2 bytes 66 = operand size override (16-bit)
xor eax, eax 31 c0 2 bytes Common way to zero a register
xor ecx, ecx 31 c9 2 bytes  
mov al, 0x0b b0 0b 2 bytes Opcode b0+r with 8-bit immediate
mov ah, 0x73 b4 73 2 bytes b4 = b0 + 4 (AH is register 4 in encoding)
mov ebx, esp 89 e3 2 bytes Uses ModR/M byte
int 0x80 cd 80 2 bytes Software interrupt (Linux x86 syscall)
ret c3 1 byte Return from function
mov eax, 0x0b b8 0b 00 00 00 5 bytes Full 32-bit immediate — too long for constrained shellcode

The Operand Size Override Prefix (0x66)

In 32-bit mode, the CPU defaults to 32-bit operands. The byte 0x66 tells the CPU to temporarily switch to 16-bit operands for the next instruction.

Without the prefix:

  • push eax = 50 (pushes 4 bytes, ESP decreases by 4)

With the prefix:

  • push ax = 66 50 (pushes 2 bytes, ESP decreases by 2)

This prefix is essential when writing size-constrained shellcode, where you need to build strings on the stack two bytes at a time.

The Register Encoding Scheme

x86 assigns each general-purpose register a number (0–7). Many opcodes use the formula base_opcode + register_number:

Register Number push (50+r) pop (58+r) mov r8, imm8 (b0+r)
EAX/AL 0 50 58 b0
ECX/CL 1 51 59 b1
EDX/DL 2 52 5a b2
EBX/BL 3 53 5b b3
ESP/AH 4 54 5c b4
EBP/CH 5 55 5d b5
ESI/DH 6 56 5e b6
EDI/BH 7 57 5f b7

Note that b4 encodes mov ah, imm8 — AH is register 4 in the 8-bit encoding scheme, which is separate from ESP (register 4 in the 32-bit scheme).

How to Look Up Encodings

Using pwntools:

from pwn import *
context.arch = 'i386'
print(asm('push ax').hex())   # → 6650
print(asm('xor eax, eax').hex())  # → 31c0

Using rasm2 (Radare2):

rasm2 -a x86 -b 32 "push ax"        # → 6650
rasm2 -a x86 -b 32 -d "6650"        # → push ax

Reference tables:

Why This Matters

When writing shellcode under constraints (byte filters, length limits, or alignment requirements), you cannot rely on an assembler to make good choices for you. You need to know exactly how many bytes each instruction compiles to, what bytes it contains, and whether alternative encodings exist.

Understanding instruction encoding transforms shellcode from guesswork into engineering.

References