Intel syntax. One page, every load-bearing table for reading and reversing 64-bit code — registers, flags, addressing, the full instruction set by family, both calling conventions, the condition-code matrix, syscalls, SIMD, encoding, and idioms. Syscall numbers and opcode bytes below were pulled from a live Linux/gcc toolchain.
System V AMD64 + Microsoft x64 · verified on gcc 10.5 / binutils 2.42
volatile / caller-saved (scratch)preserved / callee-savedwriting the 32-bit name zero-extends to 64
General-purpose (16) 4 name-widths, one physical register
64
32
16
8
SysV role
Save
rax
eax
ax
al
return / syscall# / al=#vec args
vol
rbx
ebx
bx
bl
—
pres
rcx
ecx
cx
cl
arg4 · shift count · rep counter
vol
rdx
edx
dx
dl
arg3 · hi half rdx:rax · remainder
vol
rsi
esi
si
sil
arg2 · rep source
vol
rdi
edi
di
dil
arg1 · rep dest
vol
rbp
ebp
bp
bpl
frame ptr (−O0) / spare
pres
rsp
esp
sp
spl
stack pointer
pres
r8
r8d
r8w
r8b
arg5
vol
r9
r9d
r9w
r9b
arg6
vol
r10
r10d
r10w
r10b
scratch · syscall arg4
vol
r11
r11d
r11w
r11b
scratch · syscall clobbers
vol
r12–r15
r12d…
r12w…
r12b…
long-lived locals
pres
Special & hidden
Reg
Meaning
rip
Instruction pointer. Not directly R/W; changed only by jmp/call/ret. Read only via [rip+disp] (PIE data).
rflags
Status + control flags (§2).
cs ss ds es
Segment selectors — flat/ignored in long mode (base 0).
fs
fs:0x28=stack canary; fs:0x0=TLS self-ptr (glibc). Any fs:=thread-local.
gs
TLS base on Windows/kernel (per-CPU data in kernel).
cr0/2/3/4/8
Control regs (paging, CR3=page-table base, CR2=fault addr). Ring 0 only.
xmm0–15
128-bit SSE (float/double + int vectors). §12.
ymm0–15
256-bit AVX (lower half = xmm).
zmm0–31
512-bit AVX-512 + mask regs k0–k7.
Partial-register rule: writing a 32-bit reg zero-extends to 64 (mov eax,edi clears bits 63–32). Writing 16/8-bit merges (upper bits survive). Legacy high-byte ah bh ch dh are unencodable alongside any REX-prefixed register.
The aliasing map, valid for every GP register (r8–r15 use the …d/…w/…b names).
2RFLAGS
Status flags set by COMPUTE, read by JUMP
Bit
Flag
Set when…
Tested by
0
CF carry
unsigned overflow / borrow; last bit shifted out
jb/jae, jc/jnc, adc/sbb
2
PF parity
low byte has even # of 1-bits
jp/jnp (NaN check after float cmp)
4
AF aux
carry out of bit 3 (BCD)
daa/das only
6
ZF zero
result == 0 (or operands equal after cmp)
je/jne, jz/jnz
7
SF sign
result's top bit = 1 (negative)
js/jns; jl/jge combine w/ OF
10
DF dir
control: string ops go downward
set/clear w/ std/cld
11
OF overflow
signed overflow (sign wrong)
jo/jno; jg/jl combine w/ SF
The four that matter: ZF, CF, SF, OF. cmp a,b computes a−b, discards result, keeps flags. test a,b computes a&b the same way — test r,r is the canonical "is it zero/negative?".
Control / system flags
Bit
Flag
Meaning
8
TF
trap — single-step (debuggers set this)
9
IF
interrupt enable (usually 1 in user code)
12–13
IOPL
I/O privilege level
14
NT
nested task
16
RF
resume (skip one debug fault)
17
VM
virtual-8086 mode
18
AC
alignment check
21
ID
CPUID available (togglable)
Never touch flags:mov, lea, push/pop. That's why the compiler slips them between a cmp and its jcc — read past them to find which compare owns a jump. Do touch flags: add/sub/and/or/xor/inc/dec/shifts/cmp/test.
3Addressing modes — one formula
Every memory operand is one instance of this. In lea the brackets mean pure arithmetic — no memory touched.
Legal forms & what they encode
Form
Reads as
[rax]
*rax — dereference
[rax+8]
*(rax+8) — struct field / stack local
[rax+rcx]
base+index
[rax+rcx*4]
int array: rax[rcx]
[rax+rcx*8+16]
arr[i] inside a struct at +16
[rip+0x1234]
PIE global/string (objdump prints target)
[0x404050]
absolute (non-PIE / abs32)
fs:[0x28]
segment-relative (TLS / canary)
Size keywords & data
Ptr keyword
Bits
C type
Def dir
BYTE PTR
8
char
db / .byte
WORD PTR
16
short
dw / .word
DWORD PTR
32
int/float
dd / .long
QWORD PTR
64
long/ptr/double
dq / .quad
XMMWORD PTR
128
vector
—
YMMWORD PTR
256
vector
—
Size keyword only needed when ambiguous — e.g. mov QWORD PTR [rax], 1 (no register fixes the width). Register operands imply the size.
4Data movement
Move, extend, address
Instr
Effect
mov d, s
copy (no flags)
movzx d, s
zero-extend 8/16 → wider
movsx d, s
sign-extend 8/16 → wider
movsxd r64, r32
sign-extend 32 → 64 (signed widen)
mov r32, r32
zero-extend 32 → 64 (unsigned widen — free)
lea d, [expr]
d = address expr (mul+add, no memory, no flags)
xchg d, s
swap (auto-lock if memory!)
bswap r
reverse byte order (endian flip)
movbe d, s
load/store byte-swapped
cmovcc d, s
conditional move (branchless, §8)
setcc r8
set byte to 0/1 by condition (§8)
Stack
Instr
Effect (stack grows down)
push s
rsp −= 8; [rsp] = s
pop d
d = [rsp]; rsp += 8
pushf / popf
push / pop RFLAGS
enter n,0
push rbp; mov rbp,rsp; sub rsp,n (rare)
leave
mov rsp,rbp; pop rbp (epilogue)
call t
push rip; jmp t
ret
pop rip
Sizes: stack ops are always 8-byte in 64-bit mode (no push eax). Alignment: rsp must be 16-byte aligned at everycall; since call pushes 8, a function body sees rsp ≡ 8 mod 16 on entry — hence the tell-tale sub rsp,8 / extra push to re-align.
5Arithmetic
Add / subtract / negate
Instr
Effect · flags
add d, s
d += s
sub d, s
d −= s
adc d, s
d += s + CF (multi-word add)
sbb d, s
d −= s + CF (multi-word sub)
inc d / dec d
±1 — does not touch CF
neg d
d = −d (0 − d)
xadd d, s
swap then add (atomic w/ lock)
adcx / adox
wide-mul carry chains (crypto)
Multiply / divide use rdx:rax
Instr
Effect
imul r, s
signed r *= s (2-op, truncated)
imul r, s, imm
r = s * imm (3-op)
imul s / mul s
rdx:rax = rax * s (full 128-bit)
mulx / (BMI2)
flag-free full multiply
div s
unsigned rdx:rax / s → rax, rem rdx
idiv s
signed divide, same regs
Division ritual — the type leak: before idiv the dividend's sign must fill rdx via cdq/cqo; before div, rdx is zeroed (xor edx,edx). So cqo;idiv=signed, xor edx,edx;div=unsigned. Constant divisors never emit div — they become imul by a magic reciprocal + shift.
Sign-extend accumulator (the div/pointer-math helpers)
8→16
16→32
32→64
Fill rdx from rax sign
Encoding note
cbw
cwde
cdqe (=48 98)
cwd / cdq / cqo(48 99)
cqo heralds signed idiv; cdqe widens an int index before [base+idx*s]
6Logic, shifts & bit manipulation
Boolean
Instr
Effect
and d, s
d &= s
or d, s
d |= s
xor d, s
d ^= s
not d
d = ~d (no flags)
test a, b
a & b → flags only
andn (BMI1)
d = ~a & b
Shift / rotate count = imm or cl
Instr
Effect
shl / sal
<< (×2ⁿ); 0-fill
shr
>> unsigned; 0-fill
sar
>> signed; sign-fill
rol / ror
rotate left / right
rcl / rcr
rotate through CF
shld / shrd
double-precision shift
shlx/shrx/sarx
flag-free shift (BMI2)
Bit scan / count / test
Instr
Effect
bt / bts / btr / btc
test / set / reset / flip bit → CF
bsf / bsr
index of lowest / highest 1-bit
tzcnt / lzcnt
trailing / leading zero count
popcnt
count 1-bits
blsi/blsr/blsmsk
lowest-bit tricks (BMI1)
pext / pdep
parallel bit gather/scatter (BMI2)
bzhi / bextr
zero-high / field extract
Shift = multiply/divide by 2ⁿ.shl rax,4 = ×16; sar (not shr) for signed ÷. sar reg,31/63 smears the sign bit into an all-0 or all-1 mask — the seed of branchless abs/min/max. and reg, 2ⁿ−1 = unsigned % 2ⁿ.
7Control flow
Jumps & calls
Instr
Effect
jmp t
unconditional (t = label / rax / [mem])
jcc t
conditional (§8) — the only readers of flags
call t
push return addr; jmp t
ret / ret n
pop rip (+ discard n bytes of args)
loop / loope / loopne
dec rcx; jump if rcx≠0 (& ZF)
jrcxz / jecxz
jump if rcx/ecx == 0
int3 (0xCC)
breakpoint trap (debuggers patch this)
int 0x80
legacy 32-bit syscall gate
syscall
fast kernel entry (§11)
ud2 / hlt
guaranteed #UD / halt (dead-end markers)
endbr64 (f3 0f 1e fa)
CET landing pad — semantic no-op, skip it
nop / multi-byte nop
padding / alignment — skip it
Reading structure from jump direction
Jump shape
Source construct
backward jcc
loop (target = loop head)
forward jcc
if / guard (skips the "then")
forward jmp at block end
else / break
cmp/ja + jmp reg
switch jump-table (§ atlas)
call [rip+x] / jmp [rip+x]
PLT stub → dynamic import
8Condition codes — the master matrix
The single most valuable table in reverse engineering. After cmp a,b, the g/l family reads the signed story, the a/b family reads the unsigned story. The compiler picked the mnemonic from the C declaration — so the jump you read tells you the variable's type. Each cc exists in three forms: jcc (branch), setcc (→ 0/1 byte), cmovcc (select). Opcodes below are from your assembler.
Signed & equality
cc
jmp if
flags
j-op
set/cmov
e / z
a == b
ZF=1
74
0f 94 / 0f 44
ne / nz
a != b
ZF=0
75
0f 95 / 0f 45
g / nle
a > b
ZF=0·SF=OF
7f
0f 9f / 0f 4f
ge / nl
a ≥ b
SF=OF
7d
0f 9d / 0f 4d
l / nge
a < b
SF≠OF
7c
0f 9c / 0f 4c
le / ng
a ≤ b
ZF=1|SF≠OF
7e
0f 9e / 0f 4e
Unsigned & single-flag
cc
jmp if
flags
j-op
set/cmov
a / nbe
a > b
CF=0·ZF=0
77
0f 97 / 0f 47
ae / nb / nc
a ≥ b
CF=0
73
0f 93 / 0f 43
b / nae / c
a < b
CF=1
72
0f 92 / 0f 42
be / na
a ≤ b
CF=1|ZF=1
76
0f 96 / 0f 46
s / ns
sign 1 / 0
SF
78/79
0f 98 / 0f 48
o / no
signed ovf
OF
70/71
0f 90 / 0f 40
p / np
parity / NaN
PF
7a/7b
0f 9a / 0f 4a
Why "SF≠OF" = less-than: a negative result (SF=1) usually means a<b, but if the subtraction overflowed (OF=1) the sign is a lie — so truth is SF XOR OF. Short jcc opcodes are 7x (±127 bytes); near form is 0f 8x (32-bit displacement).
Three heuristics: ① ja/jae/jb/jbe ⇒ the compared value was unsigned/pointer/size_t. ② jp/jnp right after ucomisd/comiss ⇒ a NaN check. ③ float compares always use the a/b family (they set flags like unsigned) — never jg/jl.
9String / block instructions
Primitives rsi=src, rdi=dst, direction = DF
Instr
Effect (b/w/d/q suffix = width)
movs
[rdi] = [rsi]; advance both
stos
[rdi] = al/ax/eax/rax; advance rdi
lods
al/…/rax = [rsi]; advance rsi
scas
cmp rax-family, [rdi]; advance rdi
cmps
cmp [rsi], [rdi]; advance both
Repeat prefixes & direction
Prefix
Meaning
rep
repeat rcx times → memcpy(movs) / memset(stos)
repe / repz
repeat while equal & rcx≠0 (cmps/scas)
repne / repnz
repeat while not-equal → strlen/memchr
cld / std
DF=0 forward (normal) / DF=1 backward
lock
atomic RMW prefix (see idioms)
rep movsb with rdi=dst rsi=src rcx=n is memcpy; rep stosq zero-filling = memset/bss init. Seeing these = the compiler inlined a libc call.
10Calling conventions — the two you'll meet
Same ISA, different etiquette. The killer differences: arg registers, rdi/rsi preserved on Windows, shadow space vs red zone.
System V — full contract
Item
Detail
int args
rdi rsi rdx rcx r8 r9, then stack
fp args
xmm0–xmm7 (counted separately)
varargs
al = # of xmm regs used (before printf-style call)
return
rax / rdx:rax / xmm0(:xmm1)
volatile
rax rcx rdx rsi rdi r8–r11, all xmm
preserved
rbx rbp rsp r12–r15
stack align
16-byte at each call
red zone
128 B below rsp usable by leaf fns
struct return
>16 B: hidden ptr in rdi (shifts args)
Prologue / epilogue silhouettes
−O0 (frame ptr)
−O2 (optimized)
push rbp
push rbp/rbx (if needed)
mov rbp, rsp
sub rsp, N
sub rsp, N
… locals via [rsp+x]
locals [rbp−x]
add rsp, N
leave ; ret
pop ; ret
Stack canary:mov rax,fs:0x28 → [rbp−8] in prologue; epilogue sub rax,fs:0x28; jne __stack_chk_fail. Low byte is always 00 (NUL-terminates to stop string leaks). Sits between locals and the return address.
11Linux syscalls
The syscall ABI differs from function calls!
Two changes from the function ABI: arg 4 is r10 (not rcx), and syscall hardware-clobbers rcx (saved rip) and r11 (saved flags). Number goes in rax. A bare syscall not routed through the PLT ⇒ static binary or deliberate libc bypass (shellcode / anti-analysis).
Common numbers from your unistd_64.h
#
name
#
name
0
read
21
access
1
write
22
pipe
2
open
33
dup2
3
close
39
getpid
9
mmap
41
socket
10
mprotect
42
connect
11
munmap
56
clone
12
brk
57
fork
16
ioctl
59
execve
101
ptrace
60
exit
257
openat
231
exit_group
12SIMD & floating point
Four glyphs decode most vector mnemonics.
Move / convert
Instr
Effect
movss / movsd
scalar f32 / f64
movaps / movups
aligned / unaligned 128b
movdqa / movdqu
int vector aligned/un
cvtsi2sd/ss
int → double/float
cvttsd2si
double → int (truncate)
cvtss2sd
float → double
pxor x,x
zero (float 0.0 / vector)
vzeroupper
clear ymm tops (SSE↔AVX boundary)
Compute & compare
Instr
Effect
add/sub/mul/div ss/sd/ps/pd
arithmetic
sqrt·min·max·(s/p)(s/d)
math
ucomisd / comiss
scalar cmp → ZF/CF/PF
cmpps / cmppd
packed cmp → mask
andps / xorps
bitwise (sign tricks)
shufps / pshufd
lane permute
unpck* / insertf128
interleave / build wide
vfmadd… (FMA)
a*b+c fused
Reading vector code
You see
Meaning
3 operands (v…)
AVX; 2 operands = SSE
ja/jb after ucomisd
float compare (unsigned-style)
jp after cmp
NaN check
shuffle storm
horizontal reduce (sum/dot)
xmm on stack copy
16-byte memcpy (not float!)
3 versions of a loop
vector body + scalar tail + guard
x87 FPU (st0–st7 stack, fld/fmul/fstp) appears only for long double or very old code. Normal float/double = SSE.
13Instruction encoding
The universal template. ModRM's reg/rm and SIB's scale/index/base are the §3 address formula, serialized.
REX prefix (0x40–0x4F)
Bit
Name
Meaning
0.4
W
1 = 64-bit operand
0.2
R
hi bit of ModRM.reg (reach r8–r15)
0.1
X
hi bit of SIB.index
0.0
B
hi bit of ModRM.rm / SIB.base
48 = REX.W → half of all 64-bit code starts here. 48 89/48 8b = 64-bit mov.
Byte signatures worth memorizing
Bytes
Instruction
c3
ret
55
push rbp
55 48 89 e5
push rbp; mov rbp,rsp (prologue)
48 89 e5
mov rbp,rsp
c9
leave
90
nop
cc
int3 (breakpoint / guard fill)
f3 0f 1e fa
endbr64
0f 05
syscall
31 c0
xor eax,eax (=0)
e8 / e9
call / jmp rel32
0f 8x
jcc near (rel32)
Variable length ⇒ disassembler desync. Start one byte off and you get plausible garbage until it re-syncs. Anti-disassembly abuses this with jumps into mid-instruction. Trust function boundaries; use recursive-descent tools (Ghidra) over linear sweep when output looks wrong.
14Idioms & the AT&T ↔ Intel bridge
Compiler idioms — the Rosetta stone
You see
It means
xor r,r / pxor x,x
= 0 (shortest, breaks dep chain)
test r,r ; je
if (!x) / NULL check
test al,al after call
check returned bool
lea r,[r+r*4]
×5 (mul via address unit)
imul + shr by const
division by constant (magic reciprocal)
sar r,31/63
sign mask 0 or −1
cqo;xor;sub
branchless abs()
cmovcc / setcc
ternary / bool (branchless)
mov r,fs:0x28
stack canary
mov edi,edi
zero upper 32 bits (not a nop)
and r,2ⁿ−1
unsigned % 2ⁿ
shl r,n
× 2ⁿ
lock xadd / cmpxchg
atomic ++ / compare-and-swap
big round hex const
bitmask / INT_MAX / sign bit
endbr64 · nop… · xchg ax,ax
padding — skip
AT&T vs Intel same bytes, mirrored
Intel (this sheet)
AT&T (objdump default)
mov rax, rdi
movq %rdi, %rax
mov eax, 5
movl $5, %eax
mov rax, [rbx+rcx*4+8]
movq 8(%rbx,%rcx,4), %rax
lea rax, [rip+0x10]
leaq 0x10(%rip), %rax
add QWORD PTR [rax], 1
addq $1, (%rax)
Rules: AT&T reverses operands (src, dst), prefixes regs with % and immediates with $, and encodes size in a mnemonic suffix (b/w/l/q). Force Intel: objdump -M intel; in GDB set disassembly-flavor intel.