x86 Instruction Encoding: How Assembly Becomes Machine Code
When writing shellcode or analyzing crash dumps, you frequently encounter raw hex bytes like \x31\xc0, \x66\x50, or \xcd\x80. Understanding what these bytes mean — and how to find them — is a fundamental skill in binary exploitation.
This post explains how x86 instructions are encoded into machine code.
Instruction Format
Every x86 instruction is encoded as a sequence of 1–15 bytes with the following structure:
[Prefixes] [Opcode] [ModR/M] [SIB] [Displacement] [Immediate]
Not every field is present in every instruction. Simple instructions like nop are a single byte. Complex memory operations may use all six fields.
| Field | Size | Purpose |
|---|---|---|
| Prefix | 0–4 bytes | Modifies the instruction’s behavior (operand size, segment, etc.) |
| Opcode | 1–3 bytes | The actual operation (push, mov, xor, etc.) |
| ModR/M | 0–1 bytes | Specifies registers and addressing modes |
| SIB | 0–1 bytes | Scale-Index-Base for complex memory addressing |
| Displacement | 0–4 bytes | Memory offset |
| Immediate | 0–4 bytes | Constant value embedded in the instruction |
Common Instructions and Their Encodings
Here are the instructions most frequently used in shellcode, with their raw byte encodings:
| Assembly | Hex | Size | Notes |
|---|---|---|---|
nop |
90 |
1 byte | No operation |
push eax |
50 |
1 byte | Opcode 50+r where EAX = register 0 |
push ecx |
51 |
1 byte | ECX = register 1 |
pop eax |
58 |
1 byte | Opcode 58+r |
push ax |
66 50 |
2 bytes | 66 = operand size override (16-bit) |
xor eax, eax |
31 c0 |
2 bytes | Common way to zero a register |
xor ecx, ecx |
31 c9 |
2 bytes | |
mov al, 0x0b |
b0 0b |
2 bytes | Opcode b0+r with 8-bit immediate |
mov ah, 0x73 |
b4 73 |
2 bytes | b4 = b0 + 4 (AH is register 4 in encoding) |
mov ebx, esp |
89 e3 |
2 bytes | Uses ModR/M byte |
int 0x80 |
cd 80 |
2 bytes | Software interrupt (Linux x86 syscall) |
ret |
c3 |
1 byte | Return from function |
mov eax, 0x0b |
b8 0b 00 00 00 |
5 bytes | Full 32-bit immediate — too long for constrained shellcode |
The Operand Size Override Prefix (0x66)
In 32-bit mode, the CPU defaults to 32-bit operands. The byte 0x66 tells the CPU to temporarily switch to 16-bit operands for the next instruction.
Without the prefix:
push eax=50(pushes 4 bytes, ESP decreases by 4)
With the prefix:
push ax=66 50(pushes 2 bytes, ESP decreases by 2)
This prefix is essential when writing size-constrained shellcode, where you need to build strings on the stack two bytes at a time.
The Register Encoding Scheme
x86 assigns each general-purpose register a number (0–7). Many opcodes use the formula base_opcode + register_number:
| Register | Number | push (50+r) | pop (58+r) | mov r8, imm8 (b0+r) |
|---|---|---|---|---|
| EAX/AL | 0 | 50 |
58 |
b0 |
| ECX/CL | 1 | 51 |
59 |
b1 |
| EDX/DL | 2 | 52 |
5a |
b2 |
| EBX/BL | 3 | 53 |
5b |
b3 |
| ESP/AH | 4 | 54 |
5c |
b4 |
| EBP/CH | 5 | 55 |
5d |
b5 |
| ESI/DH | 6 | 56 |
5e |
b6 |
| EDI/BH | 7 | 57 |
5f |
b7 |
Note that b4 encodes mov ah, imm8 — AH is register 4 in the 8-bit encoding scheme, which is separate from ESP (register 4 in the 32-bit scheme).
How to Look Up Encodings
Using pwntools:
from pwn import *
context.arch = 'i386'
print(asm('push ax').hex()) # → 6650
print(asm('xor eax, eax').hex()) # → 31c0
Using rasm2 (Radare2):
rasm2 -a x86 -b 32 "push ax" # → 6650
rasm2 -a x86 -b 32 -d "6650" # → push ax
Reference tables:
- x86 Opcode Reference
- Intel Software Developer Manual, Volume 2 (Instruction Set Reference)
Why This Matters
When writing shellcode under constraints (byte filters, length limits, or alignment requirements), you cannot rely on an assembler to make good choices for you. You need to know exactly how many bytes each instruction compiles to, what bytes it contains, and whether alternative encodings exist.
Understanding instruction encoding transforms shellcode from guesswork into engineering.