x86-64 · The Canonical Cheatsheet

Intel syntax. One page, every load-bearing table for reading and reversing 64-bit code — registers, flags, addressing, the full instruction set by family, both calling conventions, the condition-code matrix, syscalls, SIMD, encoding, and idioms. Syscall numbers and opcode bytes below were pulled from a live Linux/gcc toolchain.

System V AMD64 + Microsoft x64 · verified on gcc 10.5 / binutils 2.42
1 Registers 2 RFLAGS 3 Addressing 4 Data movement 5 Arithmetic 6 Logic & bits 7 Control flow 8 Condition codes 9 String ops 10 Calling conventions 11 Syscalls 12 SIMD / FPU 13 Encoding 14 Idioms & syntax

1Registers

volatile / caller-saved (scratch) preserved / callee-saved writing the 32-bit name zero-extends to 64

General-purpose (16) 4 name-widths, one physical register

6432168SysV roleSave
raxeaxaxalreturn / syscall# / al=#vec argsvol
rbxebxbxbl—pres
rcxecxcxclarg4 · shift count · rep countervol
rdxedxdxdlarg3 · hi half rdx:rax · remaindervol
rsiesisisilarg2 · rep sourcevol
rdiedididilarg1 · rep destvol
rbpebpbpbplframe ptr (−O0) / sparepres
rspespspsplstack pointerpres
r8r8dr8wr8barg5vol
r9r9dr9wr9barg6vol
r10r10dr10wr10bscratch · syscall arg4vol
r11r11dr11wr11bscratch · syscall clobbersvol
r12–r15r12d…r12w…r12b…long-lived localspres

Special & hidden

RegMeaning
ripInstruction pointer. Not directly R/W; changed only by jmp/call/ret. Read only via [rip+disp] (PIE data).
rflagsStatus + control flags (§2).
cs ss ds esSegment selectors — flat/ignored in long mode (base 0).
fsfs:0x28=stack canary; fs:0x0=TLS self-ptr (glibc). Any fs:=thread-local.
gsTLS base on Windows/kernel (per-CPU data in kernel).
cr0/2/3/4/8Control regs (paging, CR3=page-table base, CR2=fault addr). Ring 0 only.
xmm0–15128-bit SSE (float/double + int vectors). §12.
ymm0–15256-bit AVX (lower half = xmm).
zmm0–31512-bit AVX-512 + mask regs k0–k7.
Partial-register rule: writing a 32-bit reg zero-extends to 64 (mov eax,edi clears bits 63–32). Writing 16/8-bit merges (upper bits survive). Legacy high-byte ah bh ch dh are unencodable alongside any REX-prefixed register.
63 31 15 7 0 rax eax 32-bit write → these bits become 0 ah al 16/8-bit write → upper bits preserved (merge)
The aliasing map, valid for every GP register (r8–r15 use the …d/…w/…b names).

2RFLAGS

Status flags set by COMPUTE, read by JUMP

BitFlagSet when…Tested by
0CF carryunsigned overflow / borrow; last bit shifted outjb/jae, jc/jnc, adc/sbb
2PF paritylow byte has even # of 1-bitsjp/jnp (NaN check after float cmp)
4AF auxcarry out of bit 3 (BCD)daa/das only
6ZF zeroresult == 0 (or operands equal after cmp)je/jne, jz/jnz
7SF signresult's top bit = 1 (negative)js/jns; jl/jge combine w/ OF
10DF dircontrol: string ops go downwardset/clear w/ std/cld
11OF overflowsigned overflow (sign wrong)jo/jno; jg/jl combine w/ SF
The four that matter: ZF, CF, SF, OF. cmp a,b computes a−b, discards result, keeps flags. test a,b computes a&b the same way — test r,r is the canonical "is it zero/negative?".

Control / system flags

BitFlagMeaning
8TFtrap — single-step (debuggers set this)
9IFinterrupt enable (usually 1 in user code)
12–13IOPLI/O privilege level
14NTnested task
16RFresume (skip one debug fault)
17VMvirtual-8086 mode
18ACalignment check
21IDCPUID available (togglable)
Never touch flags: mov, lea, push/pop. That's why the compiler slips them between a cmp and its jcc — read past them to find which compare owns a jump. Do touch flags: add/sub/and/or/xor/inc/dec/shifts/cmp/test.

3Addressing modes — one formula

[ base + index × scale + disp ] baseany reg / rip; "which object" indexany reg but rsp; "which elem" scale1·2·4·8 only; "elem size" dispsigned 8/32-bit const; "which field"
Every memory operand is one instance of this. In lea the brackets mean pure arithmetic — no memory touched.

Legal forms & what they encode

FormReads as
[rax]*rax — dereference
[rax+8]*(rax+8) — struct field / stack local
[rax+rcx]base+index
[rax+rcx*4]int array: rax[rcx]
[rax+rcx*8+16]arr[i] inside a struct at +16
[rip+0x1234]PIE global/string (objdump prints target)
[0x404050]absolute (non-PIE / abs32)
fs:[0x28]segment-relative (TLS / canary)

Size keywords & data

Ptr keywordBitsC typeDef dir
BYTE PTR8chardb / .byte
WORD PTR16shortdw / .word
DWORD PTR32int/floatdd / .long
QWORD PTR64long/ptr/doubledq / .quad
XMMWORD PTR128vector—
YMMWORD PTR256vector—
Size keyword only needed when ambiguous — e.g. mov QWORD PTR [rax], 1 (no register fixes the width). Register operands imply the size.

4Data movement

Move, extend, address

InstrEffect
mov d, scopy (no flags)
movzx d, szero-extend 8/16 → wider
movsx d, ssign-extend 8/16 → wider
movsxd r64, r32sign-extend 32 → 64 (signed widen)
mov r32, r32zero-extend 32 → 64 (unsigned widen — free)
lea d, [expr]d = address expr (mul+add, no memory, no flags)
xchg d, sswap (auto-lock if memory!)
bswap rreverse byte order (endian flip)
movbe d, sload/store byte-swapped
cmovcc d, sconditional move (branchless, §8)
setcc r8set byte to 0/1 by condition (§8)

Stack

InstrEffect (stack grows down)
push srsp −= 8; [rsp] = s
pop dd = [rsp]; rsp += 8
pushf / popfpush / pop RFLAGS
enter n,0push rbp; mov rbp,rsp; sub rsp,n (rare)
leavemov rsp,rbp; pop rbp (epilogue)
call tpush rip; jmp t
retpop rip
Sizes: stack ops are always 8-byte in 64-bit mode (no push eax). Alignment: rsp must be 16-byte aligned at every call; since call pushes 8, a function body sees rsp ≡ 8 mod 16 on entry — hence the tell-tale sub rsp,8 / extra push to re-align.

5Arithmetic

Add / subtract / negate

InstrEffect · flags
add d, sd += s
sub d, sd −= s
adc d, sd += s + CF (multi-word add)
sbb d, sd −= s + CF (multi-word sub)
inc d / dec d±1 — does not touch CF
neg dd = −d (0 − d)
xadd d, sswap then add (atomic w/ lock)
adcx / adoxwide-mul carry chains (crypto)

Multiply / divide use rdx:rax

InstrEffect
imul r, ssigned r *= s (2-op, truncated)
imul r, s, immr = s * imm (3-op)
imul s / mul srdx:rax = rax * s (full 128-bit)
mulx / (BMI2)flag-free full multiply
div sunsigned rdx:rax / s → rax, rem rdx
idiv ssigned divide, same regs
Division ritual — the type leak: before idiv the dividend's sign must fill rdx via cdq/cqo; before div, rdx is zeroed (xor edx,edx). So cqo;idiv=signed, xor edx,edx;div=unsigned. Constant divisors never emit div — they become imul by a magic reciprocal + shift.

Sign-extend accumulator (the div/pointer-math helpers)

8→1616→3232→64Fill rdx from rax signEncoding note
cbwcwdecdqe (=48 98)cwd / cdq / cqo (48 99)cqo heralds signed idiv; cdqe widens an int index before [base+idx*s]

6Logic, shifts & bit manipulation

Boolean

InstrEffect
and d, sd &= s
or d, sd |= s
xor d, sd ^= s
not dd = ~d (no flags)
test a, ba & b → flags only
andn (BMI1)d = ~a & b

Shift / rotate count = imm or cl

InstrEffect
shl / sal<< (×2ⁿ); 0-fill
shr>> unsigned; 0-fill
sar>> signed; sign-fill
rol / rorrotate left / right
rcl / rcrrotate through CF
shld / shrddouble-precision shift
shlx/shrx/sarxflag-free shift (BMI2)

Bit scan / count / test

InstrEffect
bt / bts / btr / btctest / set / reset / flip bit → CF
bsf / bsrindex of lowest / highest 1-bit
tzcnt / lzcnttrailing / leading zero count
popcntcount 1-bits
blsi/blsr/blsmsklowest-bit tricks (BMI1)
pext / pdepparallel bit gather/scatter (BMI2)
bzhi / bextrzero-high / field extract
Shift = multiply/divide by 2ⁿ. shl rax,4 = ×16; sar (not shr) for signed ÷. sar reg,31/63 smears the sign bit into an all-0 or all-1 mask — the seed of branchless abs/min/max. and reg, 2ⁿ−1 = unsigned % 2ⁿ.

7Control flow

Jumps & calls

InstrEffect
jmp tunconditional (t = label / rax / [mem])
jcc tconditional (§8) — the only readers of flags
call tpush return addr; jmp t
ret / ret npop rip (+ discard n bytes of args)
loop / loope / loopnedec rcx; jump if rcx≠0 (& ZF)
jrcxz / jecxzjump if rcx/ecx == 0
int3 (0xCC)breakpoint trap (debuggers patch this)
int 0x80legacy 32-bit syscall gate
syscallfast kernel entry (§11)
ud2 / hltguaranteed #UD / halt (dead-end markers)
endbr64 (f3 0f 1e fa)CET landing pad — semantic no-op, skip it
nop / multi-byte noppadding / alignment — skip it

Reading structure from jump direction

loop head body cmp / jcc ↑ backward = LOOP fall-through
Jump shapeSource construct
backward jccloop (target = loop head)
forward jccif / guard (skips the "then")
forward jmp at block endelse / break
cmp/ja + jmp regswitch jump-table (§ atlas)
call [rip+x] / jmp [rip+x]PLT stub → dynamic import

8Condition codes — the master matrix

The single most valuable table in reverse engineering. After cmp a,b, the g/l family reads the signed story, the a/b family reads the unsigned story. The compiler picked the mnemonic from the C declaration — so the jump you read tells you the variable's type. Each cc exists in three forms: jcc (branch), setcc (→ 0/1 byte), cmovcc (select). Opcodes below are from your assembler.

Signed & equality

ccjmp ifflagsj-opset/cmov
e / za == bZF=1740f 94 / 0f 44
ne / nza != bZF=0750f 95 / 0f 45
g / nlea > bZF=0·SF=OF7f0f 9f / 0f 4f
ge / nla ≥ bSF=OF7d0f 9d / 0f 4d
l / ngea < bSF≠OF7c0f 9c / 0f 4c
le / nga ≤ bZF=1|SF≠OF7e0f 9e / 0f 4e

Unsigned & single-flag

ccjmp ifflagsj-opset/cmov
a / nbea > bCF=0·ZF=0770f 97 / 0f 47
ae / nb / nca ≥ bCF=0730f 93 / 0f 43
b / nae / ca < bCF=1720f 92 / 0f 42
be / naa ≤ bCF=1|ZF=1760f 96 / 0f 46
s / nssign 1 / 0SF78/790f 98 / 0f 48
o / nosigned ovfOF70/710f 90 / 0f 40
p / npparity / NaNPF7a/7b0f 9a / 0f 4a
Why "SF≠OF" = less-than: a negative result (SF=1) usually means a<b, but if the subtraction overflowed (OF=1) the sign is a lie — so truth is SF XOR OF. Short jcc opcodes are 7x (±127 bytes); near form is 0f 8x (32-bit displacement).
Three heuristics: ① ja/jae/jb/jbe ⇒ the compared value was unsigned/pointer/size_t. ② jp/jnp right after ucomisd/comiss ⇒ a NaN check. ③ float compares always use the a/b family (they set flags like unsigned) — never jg/jl.

9String / block instructions

Primitives rsi=src, rdi=dst, direction = DF

InstrEffect (b/w/d/q suffix = width)
movs[rdi] = [rsi]; advance both
stos[rdi] = al/ax/eax/rax; advance rdi
lodsal/…/rax = [rsi]; advance rsi
scascmp rax-family, [rdi]; advance rdi
cmpscmp [rsi], [rdi]; advance both

Repeat prefixes & direction

PrefixMeaning
reprepeat rcx times → memcpy(movs) / memset(stos)
repe / repzrepeat while equal & rcx≠0 (cmps/scas)
repne / repnzrepeat while not-equal → strlen/memchr
cld / stdDF=0 forward (normal) / DF=1 backward
lockatomic RMW prefix (see idioms)
rep movsb with rdi=dst rsi=src rcx=n is memcpy; rep stosq zero-filling = memset/bss init. Seeing these = the compiler inlined a libc call.

10Calling conventions — the two you'll meet

System V AMD64 · Linux/macOS/BSD Microsoft x64 · Windows rdi rsi rdx rcx r8 r9 rcx rdx r8 r9 only 4 in regs int/ptr args 1–6 · rest on stack right-to-left int args 1–4 · rest on stack xmm0–7 for floats xmm0–3 (paired w/ int slot) return: rax (rdx:rax if 128) / xmm0 preserved: rbx rbp r12–r15 128-byte red zone below rsp · no shadow space return: rax / xmm0 preserved: rbx rbp rdi rsi r12–r15 + xmm6–15 32-byte shadow space caller-reserved · no red zone
Same ISA, different etiquette. The killer differences: arg registers, rdi/rsi preserved on Windows, shadow space vs red zone.

System V — full contract

ItemDetail
int argsrdi rsi rdx rcx r8 r9, then stack
fp argsxmm0–xmm7 (counted separately)
varargsal = # of xmm regs used (before printf-style call)
returnrax / rdx:rax / xmm0(:xmm1)
volatilerax rcx rdx rsi rdi r8–r11, all xmm
preservedrbx rbp rsp r12–r15
stack align16-byte at each call
red zone128 B below rsp usable by leaf fns
struct return>16 B: hidden ptr in rdi (shifts args)

Prologue / epilogue silhouettes

−O0 (frame ptr)−O2 (optimized)
push rbppush rbp/rbx (if needed)
mov rbp, rspsub rsp, N
sub rsp, N… locals via [rsp+x]
locals [rbp−x]add rsp, N
leave ; retpop ; ret
Stack canary: mov rax,fs:0x28 → [rbp−8] in prologue; epilogue sub rax,fs:0x28; jne __stack_chk_fail. Low byte is always 00 (NUL-terminates to stop string leaks). Sits between locals and the return address.

11Linux syscalls

The syscall ABI differs from function calls!

rax rdi rsi rdx r10 r8 r9 number not rcx! → result in rax · rcx & r11 destroyed by the CPU
Two changes from the function ABI: arg 4 is r10 (not rcx), and syscall hardware-clobbers rcx (saved rip) and r11 (saved flags). Number goes in rax. A bare syscall not routed through the PLT ⇒ static binary or deliberate libc bypass (shellcode / anti-analysis).

Common numbers from your unistd_64.h

#name#name
0read21access
1write22pipe
2open33dup2
3close39getpid
9mmap41socket
10mprotect42connect
11munmap56clone
12brk57fork
16ioctl59execve
101ptrace60exit
257openat231exit_group

12SIMD & floating point

vAVX/VEX (absent=SSE) mulop: add sub mul div sqrt min max pp=packed s=scalar ss=single(f32) d=double(f64) vmulps AVX · packed · single mulsd = scalar double · addss = scalar float · integer vectors: paddd/paddq/pcmpeqb (p + element width)
Four glyphs decode most vector mnemonics.

Move / convert

InstrEffect
movss / movsdscalar f32 / f64
movaps / movupsaligned / unaligned 128b
movdqa / movdquint vector aligned/un
cvtsi2sd/ssint → double/float
cvttsd2sidouble → int (truncate)
cvtss2sdfloat → double
pxor x,xzero (float 0.0 / vector)
vzeroupperclear ymm tops (SSE↔AVX boundary)

Compute & compare

InstrEffect
add/sub/mul/div ss/sd/ps/pdarithmetic
sqrt·min·max·(s/p)(s/d)math
ucomisd / comissscalar cmp → ZF/CF/PF
cmpps / cmppdpacked cmp → mask
andps / xorpsbitwise (sign tricks)
shufps / pshufdlane permute
unpck* / insertf128interleave / build wide
vfmadd… (FMA)a*b+c fused

Reading vector code

You seeMeaning
3 operands (v…)AVX; 2 operands = SSE
ja/jb after ucomisdfloat compare (unsigned-style)
jp after cmpNaN check
shuffle stormhorizontal reduce (sum/dot)
xmm on stack copy16-byte memcpy (not float!)
3 versions of a loopvector body + scalar tail + guard
x87 FPU (st0–st7 stack, fld/fmul/fstp) appears only for long double or very old code. Normal float/double = SSE.

13Instruction encoding

prefixes REX opcode 1–3B ModRM SIB disp 0/1/4B imm 0/1/2/4/8B all optional except opcode · total length 1–15 bytes
The universal template. ModRM's reg/rm and SIB's scale/index/base are the §3 address formula, serialized.

REX prefix (0x40–0x4F)

BitNameMeaning
0.4W1 = 64-bit operand
0.2Rhi bit of ModRM.reg (reach r8–r15)
0.1Xhi bit of SIB.index
0.0Bhi bit of ModRM.rm / SIB.base
48 = REX.W → half of all 64-bit code starts here. 48 89/48 8b = 64-bit mov.

Byte signatures worth memorizing

BytesInstruction
c3ret
55push rbp
55 48 89 e5push rbp; mov rbp,rsp (prologue)
48 89 e5mov rbp,rsp
c9leave
90nop
ccint3 (breakpoint / guard fill)
f3 0f 1e faendbr64
0f 05syscall
31 c0xor eax,eax (=0)
e8 / e9call / jmp rel32
0f 8xjcc near (rel32)
Variable length ⇒ disassembler desync. Start one byte off and you get plausible garbage until it re-syncs. Anti-disassembly abuses this with jumps into mid-instruction. Trust function boundaries; use recursive-descent tools (Ghidra) over linear sweep when output looks wrong.

14Idioms & the AT&T ↔ Intel bridge

Compiler idioms — the Rosetta stone

You seeIt means
xor r,r / pxor x,x= 0 (shortest, breaks dep chain)
test r,r ; jeif (!x) / NULL check
test al,al after callcheck returned bool
lea r,[r+r*4]×5 (mul via address unit)
imul + shr by constdivision by constant (magic reciprocal)
sar r,31/63sign mask 0 or −1
cqo;xor;subbranchless abs()
cmovcc / setccternary / bool (branchless)
mov r,fs:0x28stack canary
mov edi,edizero upper 32 bits (not a nop)
and r,2ⁿ−1unsigned % 2ⁿ
shl r,n× 2ⁿ
lock xadd / cmpxchgatomic ++ / compare-and-swap
big round hex constbitmask / INT_MAX / sign bit
endbr64 · nop… · xchg ax,axpadding — skip

AT&T vs Intel same bytes, mirrored

Intel (this sheet)AT&T (objdump default)
mov rax, rdimovq %rdi, %rax
mov eax, 5movl $5, %eax
mov rax, [rbx+rcx*4+8]movq 8(%rbx,%rcx,4), %rax
lea rax, [rip+0x10]leaq 0x10(%rip), %rax
add QWORD PTR [rax], 1addq $1, (%rax)
Rules: AT&T reverses operands (src, dst), prefixes regs with % and immediates with $, and encodes size in a mnemonic suffix (b/w/l/q). Force Intel: objdump -M intel; in GDB set disassembly-flavor intel.

Atomics & barriers

InstrEffect
lock add/xadd/or…atomic read-modify-write
cmpxchg / cmpxchg16bcompare-and-swap (lock-free)
mfence/lfence/sfencememory ordering barriers
pausespin-loop hint
rdtsc / rdrand / cpuidtimestamp / RNG / feature probe