Intel syntax, all the way down. Every byte, register value, flag bit, and address in this atlas was harvested from specimen.c compiled and executed on your machine — rerun any command and watch it agree.
The specimen is a small C file where every function exists to trigger exactly one assembly pattern you will meet in real reverse engineering: addressing modes, magic-number division, jump tables, canaries, callee-saved shuffles, SIMD lanes. Rebuild it any time:
gcc -O2 -fno-inline -o specimen specimen.c # the RE-realistic build
gcc -O0 -o specimen-O0 specimen.c # the teaching build (frame pointers)
objdump -d -M intel specimen # everything below comes from this
x86-64 gives you sixteen 64-bit general-purpose registers. Each one answers to four names depending on how much of it you're touching. Here is the anatomy of rax — all four names refer to the same physical register:
That asymmetry is not trivia — it's the single most common source of confusion when reading 64-bit disassembly, so we proved it live on your CPU. In maxi, right before mov eax, edi executes, we poisoned rax and rdi from GDB, then single-stepped the real instruction:
(gdb) set $rax = 0xdeadbeefcafebabe
(gdb) set $rdi = 0x11112222deadbeef
(gdb) x/i $rip
=> 0x5555555553f6 <maxi+6>: mov eax, edi
(gdb) stepi
after real mov eax,edi: rax = 0xdeadbeef ← upper 32 bits wiped to zero
Only the low 32 bits of rdi were read, and the whole upper half of rax became zero. Two consequences you'll use constantly:
mov eax, edi is a free zero-extension to 64 bits and one byte shorter than mov rax, rdi. When you see 32-bit registers in 64-bit code, nothing is wrong — it's the default.mov edi, edi is not a no-op. Your own day_name contains one at 1430 — it zeroes bits 63–32 of rdi so the next instruction can safely use rdi*4 as a table index. (Careful: a GDB set $eax=… does not emulate this — only real instructions zero-extend. We hit that gotcha while making this atlas.)The full atlas — learn the roles, because the roles are how you decompile in your head:
| 64 | 32 | 16 | 8 | System V role | Survives a call? | What it means when you see it |
|---|---|---|---|---|---|---|
| rax | eax | ax | al | return value | no — volatile | after a call: the result. Before syscall: the syscall number. al before varargs call: count of vector registers used. |
| rdi | edi | di | dil | argument 1 | no | loaded right before a call = first argument. Dest for rep movs/stos. |
| rsi | esi | si | sil | argument 2 | no | second argument. Source for rep movs. |
| rdx | edx | dx | dl | argument 3 | no | third argument; high half of 128-bit results (rdx:rax in mul/div); remainder after div. |
| rcx | ecx | cx | cl | argument 4 | no | fourth argument; cl = variable shift count; rep counter; destroyed by syscall (hardware stores return RIP there). |
| r8 | r8d | r8w | r8b | argument 5 | no | fifth argument. |
| r9 | r9d | r9w | r9b | argument 6 | no | sixth argument. Argument 7+ goes on the stack. |
| r10 | r10d | r10w | r10b | scratch | no | replaces rcx as argument 4 in the kernel syscall convention; nested-function static chain. |
| r11 | r11d | r11w | r11b | scratch | no | destroyed by syscall (hardware stores RFLAGS there); favorite PLT/veneer scratch. |
| rbx | ebx | bx | bl | callee-saved | yes — preserved | a value the function wants to keep across calls — usually a pointer or loop-carried variable. Watch the push rbx/pop rbx pair. |
| rbp | ebp | bp | bpl | callee-saved | yes | frame pointer at -O0 (locals are [rbp-X]); just another preserved register at -O2. |
| r12–r15 | r12d… | r12w… | r12b… | callee-saved | yes | more long-lived locals. The number of push r1X in a prologue ≈ how many variables outlive calls. |
| rsp | esp | sp | spl | stack pointer | yes (by contract) | always points at the last pushed qword. Must be 16-byte aligned at every call. |
| rip | — | instruction ptr | — | unreadable/unwritable directly; only call/jmp/ret change it, only [rip+disp] reads it. | ||
| fs | — | TLS base | — | fs:0x28 = the stack canary. Any fs: access = thread-local storage. | ||
sph/dih: the byte registers ah cH dh bh (bits 15–8) are leftovers from the 8086. Any instruction carrying a REX prefix (needed for r8–r15, spl, dil…) reinterprets those four encodings as spl bpl sil dil instead. So mov ah, r8b is physically unencodable — the two register families can't meet in one instruction. When you see ah in modern code it's almost always lahf-style flag juggling or byte-swapping tricks.
[base + index·scale + disp]Every memory operand in x86-64 — every array access, struct field, local variable, jump table — is one instance of a single hardware formula. Your specimen's index_scale function compiles to exactly one instance of it:
long index_scale(long *arr, long i) { return arr[i]; }
1344: mov rax, QWORD PTR [rdi + rsi*8] ; arr in rdi, i in rsi
The three specimens, decoded with that lens:
; int struct_field(struct player *p) { return p->mp; } — disp alone picks a field
1354: mov eax, DWORD PTR [rdi+0x4] ; mp is 4 bytes into the struct
; int matrix(int m[][10], long r, long c) { return m[r][c]; } — r*10 doesn't fit scale∈{1,2,4,8},
; so the compiler builds ×10 out of two lea's: (r + r*4)*8 = r*40 = r * 10 ints
1364: lea rax, [rsi+rsi*4] ; rax = r*5
1368: lea rax, [rdi+rax*8] ; rax = m + r*40 (= &m[r][0])
136c: mov eax, DWORD PTR [rax+rdx*4] ; load m[r][c]
lea — the address formula hijacked for arithmeticlea (load effective address) computes the bracket without touching memory — it's a 3-operand add-and-multiply that happens to use address syntax. Your lea_math:
long lea_math(long x) { return x * 5 + 7; }
1374: lea rax, [rdi + rdi*4 + 0x7] ; rax = x + x*4 + 7 — no RAM involved
mov rax, [rdi] dereferences — it reads 8 bytes of memory. lea rax, [rdi+8] only does the math — it's rax = rdi + 8, memory untouched. If a bracket appears in lea, mentally erase the brackets. Compilers love it because it does mul+add+move in one instruction and — unlike add — doesn't clobber RFLAGS, so it can sit between a cmp and its jcc. You'll also constantly see lea rdi, [rip+0xXXX] — that's "take the address of this string/global", the position-independent way.
You don't need to hand-assemble, but knowing the byte anatomy pays off in RE: it's how you recognize function boundaries in raw bytes, spot misaligned disassembly, and understand why the same mnemonic has many lengths. Here is your own mov rax, [rdi+rsi*8] — four bytes, fully decoded:
index_scale's only real instruction. mod/reg/rm and scale/index/base are just the formula from Step 2, serialized.The general template every instruction follows (most parts optional): [prefixes] [REX] opcode [ModRM] [SIB] [disp] [imm] — from 1 byte (ret = c3, push rbp = 55) to the architectural maximum of 15 bytes. Practical payoffs:
48 is everywhere. Half of 64-bit code starts with 48 (REX.W). Seeing 48 8b/48 89 in a hexdump = 64-bit mov. 55 48 89 e5 = push rbp; mov rbp,rsp — a function prologue signature you can grep raw bytes for.f3 0f 1e fa = endbr64 — a CET landing pad stamped at every indirect-jump target. Pure armor, zero semantics: skip it while reading. Same for the padding zoo between functions: nop WORD PTR [rax+rax*1+0x0], xchg ax,ax, data16 cs nop — multi-byte NOPs aligning the next function to 16 bytes.0f — your cmovge is 0f 4d c6, all SSE is 0f-something. AVX replaces the prefix soup with c4/c5 (VEX).There is no "if" instruction. There is cmp (a subtraction that throws away the result but keeps the side effects) and, later, a conditional jump that inspects those side effects. The side effects live in RFLAGS. We captured a real one on your CPU — inside maxi, we forced edi=5, esi=3 and stepped over cmp esi, edi (computing 3−5):
(gdb) set $edi=5 ; set $esi=3 ; stepi # executes: cmp esi, edi → 3 - 5
eflags 0x293 [ CF AF SF IF ] # ZF=0 OF=0
cmp esi, edi with 3 and 5. One subtraction feeds every possible question about a vs b.cmp a, b, "greater/less" (jg jge jl jle) test the signed story (SF, OF), while "above/below" (ja jae jb jbe) test the unsigned story (CF). The compiler chose the mnemonic based on the C declaration — so when you read ja, you just learned the variable was unsigned (or a pointer/size_t); when you read jg, it was signed. Disassembly leaks the type system. Nothing tells you more about the original source for less effort.
After cmp a, b, jump if… | Signed (g/l family) | Flags tested | Unsigned (a/b family) | Flags tested |
|---|---|---|---|---|
| equal / not equal | je (jz) · jne (jnz) — ZF, shared by both worlds | |||
| a > b | jg | ZF=0 and SF=OF | ja | CF=0 and ZF=0 |
| a ≥ b | jge | SF=OF | jae (jnc) | CF=0 |
| a < b | jl | SF≠OF | jb (jc) | CF=1 |
| a ≤ b | jle | ZF=1 or SF≠OF | jbe | CF=1 or ZF=1 |
| sign / overflow / parity | js · jo · jp — direct single-flag tests (jp after float compare = NaN!) | |||
Why "SF≠OF" means "less than", in one sentence: normally a negative result (SF=1) means a<b, but if the subtraction overflowed (OF=1) the sign is a lie, so the truth is the XOR of the two. Your capture shows the honest case: SF=1, OF=0 → 3<5 signed; and CF=1 → 3<5 unsigned too.
Every condition code cc exists in three instruction forms, and your specimen has all three:
; 1. jcc — branch on it (day_name's bounds check)
1424: cmp edi, 0x6
1427: ja 14a8 ; unsigned trick: one ja handles d<0 AND d>6 (see Step 7)
; 2. cmovcc — select on it, no branch (maxi: return a > b ? a : b)
13f4: cmp esi, edi
13f6: mov eax, edi
13f8: cmovge eax, esi ; if (b >= a) result = b — a ternary with zero branches
; 3. setcc — materialize it as 0/1 (is_zero: return x == 0)
1404: xor eax, eax ; clear first — sete only writes 1 byte (al)!
1406: test rdi, rdi
1409: sete al ; al = (x == 0)
test vs cmp, the two flag-setters you'll see most: cmp a,b is a−b discarded; test a,b is a&b discarded. test reg, reg (same register twice) is the idiomatic "is it zero / is it negative?" — after it, je means "was zero" (NULL check after loading a pointer, error check after a call), js means "was negative". test al, al after a call = checking a returned bool/char. Also remember both xor reg,reg and sub/add/and/or update flags as a side effect, but mov, lea, push/pop never touch flags — that's why compilers can schedule them between a cmp and its jcc, and you must read past them to find which comparison a jump belongs to.
Nothing in the hardware defines "arguments". The System V AMD64 ABI (Linux, macOS, BSD — Windows differs!) says: first six integer/pointer args ride in rdi rsi rdx rcx r8 r9, first eight float/double args in xmm0–xmm7 (counted separately), the rest go on the stack right-to-left, and the return comes back in rax (or xmm0). We froze your many_args(1,2,3,4,5,6,7,8) at its first instruction and photographed the contract in action:
(gdb) break *many_args ; run
reg args: rdi=1 rsi=2 rdx=3 rcx=4 r8=5 r9=6
0x7fffffffd6d8: 0x0000555555555186 ← [rsp] return address (pushed by call)
0x7fffffffd6e0: 0x0000000000000007 ← [rsp+0x08] argument 7
0x7fffffffd6e8: 0x0000000000000008 ← [rsp+0x10] argument 8
many_args got 1–6 in registers, 7 and 8 at [rsp+8] and [rsp+0x10], and returned 36 in rax.The volatile/preserved split is where reverse engineering gets real leverage. Your callee_saved_demo needs a and b to survive the call to greet — watch what the compiler does:
long callee_saved_demo(long a, long b) { greet("demo"); return a*3 + b; }
1584: push rbp ; save caller's rbp — I'm about to use it
1585: mov rbp, rsi ; b moves into a PRESERVED register
1588: push rbx ; save caller's rbx too
1589: mov rbx, rdi ; a moves into a preserved register
158c: lea rdi, [rip+0xa9b] ; arg1 = "demo"
1593: sub rsp, 0x8 ; realign: 2 pushes + retaddr = 24 bytes → +8 = 32 ✔
1597: call 1520 <greet> ; may destroy rax rcx rdx rsi rdi r8-r11…
159c: lea rax, [rbx+rbx*2] ; …but rbx still holds a: rax = a*3
15a4: add rax, rbp ; + b, still alive in rbp
15a7: pop rbx ; restore in reverse order
15a8: pop rbp
15a9: ret
rdi/rsi/rdx… before a call are the arguments — that's how you recover prototypes of unnamed functions. A value moved into rbx/r12–r15 at the top is a long-lived local. eax examined right after a call is the return value being checked. And at -O2, note what's absent: no frame pointer, no wasted moves — rbp here is just a spare preserved register holding b, nothing to do with frames.
The stack grows downward (push = rsp -= 8 then store), and a function's frame is just the slice of stack between the return address and wherever rsp stops. Compare the two builds of greet — same C, two dialects:
;— -O0: the textbook frame —————————————; ;— -O2: frame pointer optimized out ——————
push rbp push rbp ; (just preserving it)
mov rbp, rsp ; anchor the frame ; args staged into r9, r8, rcx, rdx, rsi…
sub rsp, 0x60 ; room for locals sub rsp, 0x50
mov [rbp-0x58], rdi mov rax, fs:0x28 ; canary in…
mov rax, fs:0x28 mov [rsp+0x48], rax
mov [rbp-0x8], rax ; canary ; locals addressed from RSP, not rbp
; locals at [rbp-X] — easy to read mov rbp, rsp ; rbp = &buf (reused!)
… …
leave ; = mov rsp,rbp; pop rbp add rsp, 0x50
ret pop rbp
ret
We froze the -O2 frame live, right after the canary store, and mapped every byte (all values real):
greet frame, photographed live. The canary's low byte is always 00 — it terminates C strings so string functions can't leak or write past it quietly.And the widest view — where that frame sits in the whole process. This is info proc mappings from your actual run:
info proc mappings. Code low, heap above it, libraries high, stack highest — with a canyon between.Compilers emit a small number of recognizable shapes. Once you know the silhouettes, you read structure instead of instructions.
sum_array 14b4: test rsi, rsi ; ── guard: n <= 0?
14b7: jle 14d0 ; skip loop entirely (early-out block below)
14b9: lea rcx, [rdi+rsi*4] ; ── setup: rcx = one-past-the-end pointer!
14bd: xor eax, eax ; s = 0
14c0: movsxd rdx, DWORD PTR [rdi] ; ─┐ body: load a[i] (sign-extend int→long)
14c3: add rdi, 0x4 ; │ pointer walks — the compiler killed i!
14c7: add rax, rdx ; │ s += a[i]
14ca: cmp rdi, rcx ; │ pointer == end?
14cd: jne 14c0 ; ─┘ BACKWARD jump = "this is a loop"
14cf: ret
14d0: xor eax, eax ; the n<=0 path: return 0
14d2: ret
Two habits of optimized loops to recognize: the compiler replaced the index with a walking pointer (no i exists — rdi strides by 4 toward a precomputed end pointer), and the loop condition moved to the bottom (do-while shape) with a guard on top — cheaper than jumping to a test each iteration.
day_name's jump table 1424: cmp edi, 0x6
1427: ja 14a8 ; unsigned trick: d<0 wraps to huge → one test kills both sides
1429: lea rdx, [rip+0xc1c] # 204c — table base, in .rodata
1430: mov edi, edi ; zero upper 32 (Step 1's "useless" mov!)
1432: movsxd rax, DWORD PTR [rdx+rdi*4] ; fetch 32-bit OFFSET, not a pointer
1436: add rax, rdx ; target = base + offset (PIE-friendly)
1439: notrack jmp rax ; indirect jump into the case blocks
.rodata:0x204c, decoded entry by entry. cmp/ja + lea rip-rel + movsxd + jmp rax = "this was a switch".ja bounds trick is a mini signed/unsigned masterclass: d is a signed int, but the compiler tests cmp edi,6; ja default — an unsigned compare. A negative d reinterpreted as unsigned becomes huge (≥ 0x80000000), so one unsigned test rejects both d > 6 and d < 0. When you see ja right before an indirect jump, it's a switch bounds check, and the case count is the immediate + 1.
The if/else silhouette barely needs a diagram once you've internalized Step 4: cmp/test → jcc over the "then" block, optional jmp over the "else". At -O2 short ifs usually vanish into cmov/setcc (branchless), so a real branch surviving in optimized code hints the compiler thought it unpredictable or the sides too big.
Mixing widths is where silent bugs and RE misreadings live. The machine has exactly three widening strategies — and your specimen used all of them:
movzx exists for 8/16-bit sources; 32→64 zero-extension is just a plain 32-bit mov.The in-place cousins act on the accumulator family only, and the division ritual depends on them:
| Instruction | Effect | Where you'll meet it |
|---|---|---|
| cbw / cwde / cdqe | al→ax → eax → rax (sign, in place) | widening a local before pointer math |
| cwd / cdq / cqo | sign of ax/eax/rax fills dx/edx/rdx | the herald of signed division — your absolute uses cqo |
| cqo ; idiv rcx | signed divide rdx:rax → quot rax, rem rdx | cqo/cdq before idiv = signed / or % |
| xor edx,edx ; div rcx | unsigned divide | zeroed rdx before div = unsigned — type leak #3 |
Optimizing compilers speak in idioms — fixed phrases that mean something quite different from their literal instructions. This table is the heart of "reading assembly like C". Every row marked ● was pulled from your own specimen's disassembly:
| You see | It means | Why / notes |
|---|---|---|
| ● xor eax, eax | eax = 0 | 2 bytes vs 5 for mov eax,0; also breaks CPU dependency chains. Ubiquitous. |
| ● test rdi, rdi ; je … | if (!ptr) / if (x == 0) | AND with itself only sets flags. js variant = "if negative". |
| ● lea rax, [rdi+rdi*4] | x * 5 | lea = mul/add without flags. ×3, ×5, ×9 direct; ×10 = two leas (your matrix). |
| ● imul rax, rdi ; shr rax, 35 (after mov edi, 0xcccccccd) | x / 10 | Magic-number division — walkthrough below. Constant divisors never emit div. |
| ● sar edx, 31 | sign mask: 0 or −1 | Arithmetic shift smears the sign bit; the all-ones mask feeds branchless tricks. |
| ● cqo ; xor rax,rdx ; sub rax,rdx | labs(x) | Your absolute: (x XOR mask) − mask negates iff mask=−1. No branch. |
| ● (sar,shr,lea) ; and ; sub | x % 8 (signed) | Signed modulo must round toward zero, so a bias built from the sign mask wraps the simple and eax,7. Unsigned % 2ⁿ is just the and. |
| ● cmovge / sete / ja-vs-jg | ternary / bool / type leak | Steps 4 & 7. |
| ● mov rax, fs:0x28 | stack canary | Any fs: access = thread-local storage; offset 0x28 is glibc's canary slot. |
| ● mov edi, edi | zero bits 63–32 of rdi | Not a no-op! (Step 1.) |
| ● endbr64 / nop WORD PTR […] / xchg ax,ax | nothing | CET landing pads and alignment padding. Train your eyes to skip them. |
| shl rax, 4 ; add/or rax, … | x*16 + y | Index into array of 16-byte structs; shifts are multiplies by 2ⁿ. |
| rep movsb / rep stosq | memcpy / memset | rdi=dest rsi=src rcx=count; DF decides direction (cld = forward). |
| pxor xmm0, xmm0 | 0.0 (or zero vector) | The float twin of xor eax,eax — your mixed does it before cvtsi2sd. |
| add rsp, 8 … or sub rsp, 8 around a call | alignment, not data | Keeping rsp ≡ 16 at the call (your callee_saved_demo). |
| mov reg, 0x7fffffff… / 0x80000000… | INT_MAX / sign-bit mask | Big round hex constants are almost always masks or limits — decode them. |
x / 10 is a multiply. Division is ~20–40 cycles; multiplication ~3. So for any constant divisor the compiler precomputes a fixed-point reciprocal. Your div10:
1384: mov eax, edi ; zero-extend x to 64 bits (Step 1!)
1386: mov edi, 0xcccccccd ; ⌈2³⁵ / 10⌉ — "0.1" in fixed-point
138b: imul rax, rdi ; 64-bit product = x · (2³⁵/10)
138f: shr rax, 0x23 ; ÷ 2³⁵ ⇒ floor(x/10), exact for all u32
The math: 0xcccccccd = 3435973837 = ⌈2³⁵/10⌉, so (x·0xcccccccd) >> 35 = x·(2³⁵/10)/2³⁵ ≈ x/10, and the ceiling guarantees exactness for every 32-bit x. Verified live: your program printed div10(1234) = 123. RE recipe: when you meet imul with an ugly constant followed by a shift, the divisor is round(2^(32+shift) / magic) — for 0xcccccccd, shift 3 (35−32): 2³⁵/0xcccccccd ≈ 9.99999… → 10. Repeating-pattern constants are the tell: 0xcccccccd→10, 0xaaaaaaab→3 or 6, 0x92492493→7, 0x38e38e39→9 or 18.
All floating point in x86-64 goes through the vector registers — there is no separate "float unit" in normal compiled code (the ancient x87 stack only appears for long double). So even scalar double math lives in xmm0…, and the ABI returns floats in xmm0. The registers nest like the GPRs do:
Scalar reality first — your mixed(double x, long n, double y) is pure lane-0 code:
1504: movapd xmm2, xmm0 ; save x (xmm0 is needed as the result reg)
1508: pxor xmm0, xmm0 ; zero idiom — breaks false dependency
150c: cvtsi2sd xmm0, rdi ; ConVerT Signed Int 2 Scalar Double: n → double
1511: mulsd xmm0, xmm2 ; x * n (scalar double)
1515: addsd xmm0, xmm1 ; + y → returned in xmm0 = 6.25 ✔ (printed by your run)
Packed reality — dot() from simd.c built with -O3 -mavx2. Eight multiplications happen in one instruction; the rest is folding the lanes down to one scalar:
154: vmovups xmm4, [rsi] ; load b[0..3]
160: vinsertf128 ymm0, ymm5, [rdi+0x10], 1 ; build a[0..7] in one ymm
16e: vmulps ymm1, ymm1, ymm0 ; a[i]*b[i] — ×8 AT ONCE
172: vaddss xmm3, xmm3, xmm1 ; ── horizontal reduction begins:
176: vshufps xmm2, xmm1, xmm1, 0x55 ; shuffle lane k to lane 0,
… (vshufps / vunpckhps / vextractf128 …) ; add, repeat for all 8 lanes
1b4: vzeroupper ; drop ymm halves before returning
Field notes for reading vector code:
vmulps ymm1, ymm1, ymm0 = dst, src1, src2) means AVX; two operands means SSE, where dst is also src1.shufps/unpck/extractf128/hadd after packed math = reducing lanes to a scalar (sum, max, dot). Recognize the shape; don't trace every lane.vzeroupper marks the SSE/AVX boundary (avoids a big transition stall) — harmless, and a handy "vector region ends here" landmark.movaps vs movups: aligned vs unaligned load/store. movaps xmm, [rsp+…] is also how compilers copy any 16-byte blob — seeing xmm registers does not imply float data (memcpy of structs uses them constantly).ucomisd/comiss) set ZF/CF like an unsigned integer compare — so float branches use ja/jb/jae, never jg/jl. PF=1 signals NaN: a jp right after a float compare is a NaN check, one more type leak.seta/test maze at the top of add_arrays is an overlap check between the arrays), the wide body, and a scalar tail for leftover elements. One C loop, three assembly loops — don't mistake them for three source loops.Your main is not the entry point. The ELF header's e_entry points at _start; main's address is merely handed to __libc_start_main as an argument. Straight from your _start:
1254: xor ebp, ebp ; mark the outermost frame (rbp=0 ends backtraces)
1256: mov r9, rdx ; rtld_fini
1259: pop rsi ; argc (kernel put it on the stack)
125a: mov rdx, rsp ; argv
125d: and rsp, 0xfffffffffffffff0 ; 16-align the stack (the ABI mandate)
1268: lea rdi, [rip+0xfffffffffffffe51] # 10c0 <main> — main is just an ARGUMENT
126f: call [rip+0x2d63] # __libc_start_main — never returns
1275: hlt
__attribute__((constructor)), C++ static init) run in __libc_start_main — malware hides here to execute before main.file specimen # arch, PIE?, stripped?, static/dynamic
objdump -d -M intel specimen # the static disassembly (Intel syntax!)
readelf -h -l -d specimen # entry point, segments, libraries needed
strings -a -t x specimen # literals + their offsets — the fastest clues
nm -C specimen / readelf --dyn-syms # symbols, imports (UND = comes from libc)
ltrace ./specimen self # library-call trace: the API storyline
strace ./specimen self # syscall trace: what it does to the OS
gdb -q ./specimen # dynamic: break main, stepi, x/i $pc, watch
gef> break main ; run ; context # GEF/pwndbg give regs+stack+code at a glance
Static tells you what could happen; dynamic tells you what did. Alternate between them. For serious work, a disassembler with a decompiler (Ghidra is free, IDA/Binary Ninja commercial) turns all of Steps 4–10 into pseudo-C automatically — but only someone who can read the raw form knows when the decompiler is lying.
dst, src (mov rax, rdi); AT&T is src, dst with %/$ sigils and (%rdi,%rsi,8) for memory. objdump defaults to AT&T — always pass -M intel. GDB: set disassembly-flavor intel.rsp without adjusting it — so signal handlers and naive stack scanners can corrupt data that looks "free". Windows has no red zone.call pushes the return address; that's the only reason the stack and the "return-oriented" attack surface exist. ret pops it back into RIP — overwrite that saved qword (past the canary) and you redirect control. This is why the canary sits between buffers and the return address (Step 6).[rip+disp]; objdump helpfully prints the resolved target as a # comment. Follow the comment, not the disp.00, cc (int3), or multi-byte NOPs between functions are filler. cc specifically is a breakpoint trap — a sea of cc = uninitialized/guard space.jmp into the middle of an instruction desyncs linear disassemblers (Step 3). If objdump output looks like nonsense mid-function, try disassembling from a different start offset, or use a recursive-descent tool (Ghidra) that follows control flow.syscall enters the kernel with the number in rax and args in rdi rsi rdx r10 r8 r9 (note r10, not rcx — the CPU clobbers rcx/r11). A bare syscall in .text without going through the PLT = static binary or deliberate libc-bypass (common in shellcode and anti-analysis).mov eax, 0x1234 stores 34 12 00 00 in the instruction stream and in memory; registers show the logical value. You only meet little-endian when reading hex dumps and building constants by hand.gdb -q ./specimen -ex 'break *maxi+6' -ex run -ex 'set $rax=0xffffffffffffffff' -ex stepi -ex 'p/x $rax'. Watch mov eax,edi blow away the top half. Then try the same idea after an ax-writing instruction and confirm it does not.objdump -d -M intel specimen | less, jump to many_args, and reconstruct its signature purely from which registers/stack slots feed the adds. Check against the source.*maxi+4, set edi/esi to different signed and unsigned extremes (e.g. 0x80000000 vs 1), step the cmp, and predict eflags before revealing it. Feel why jg and ja disagree on the same bytes.unsigned d7(unsigned x){return x/7;} to specimen, recompile, and reverse the emitted imul constant + shift back to "7" using the recipe in Step 9. Try a signed version and spot the extra sign-fix instructions.objdump -s -j .rodata specimen, find the 7 little-endian dwords at 0x204c, add each (as a signed offset) to 0x204c, and confirm the targets match the lea rax,[rip+…] case blocks in day_name. You'll see the cases are physically out of order.greet after the canary store, set a byte inside buf's canary slot, continue, and watch __stack_chk_fail abort the program instead of returning — the mitigation firing in real time.objdump -d -M intel --start-address=0x1345 specimen (one byte into a real instruction) and watch plausible garbage appear — the variable-length desync from Step 3, first-hand.simd.c at -O3 -mavx2 vs -O3 -mno-avx vs -O0 and diff dot. One math loop becomes eight-wide, four-wide, or scalar — the compiler's vectorization decision made visible.