x86-64 — The Reverse Engineer's Atlas

Intel syntax, all the way down. Every byte, register value, flag bit, and address in this atlas was harvested from specimen.c compiled and executed on your machine — rerun any command and watch it agree.

/src/as/specimen  ·  gcc 10.5 -O2  ·  objdump -M intel  ·  gdb + GEF

The specimen is a small C file where every function exists to trigger exactly one assembly pattern you will meet in real reverse engineering: addressing modes, magic-number division, jump tables, canaries, callee-saved shuffles, SIMD lanes. Rebuild it any time:

gcc -O2 -fno-inline -o specimen specimen.c   # the RE-realistic build
gcc -O0 -o specimen-O0 specimen.c            # the teaching build (frame pointers)
objdump -d -M intel specimen                 # everything below comes from this
The one mental model to hold onto: the CPU is an infinite loop — fetch the instruction at RIP, execute it, repeat — and every instruction belongs to one of three families: MOVE bytes between registers and memory, COMPUTE on them (which silently drops breadcrumbs into RFLAGS), or JUMP based on those breadcrumbs. That's the entire machine. Everything else — functions, arguments, stack frames, "local variables" — is not hardware; it's social convention (the System V ABI) layered on those three verbs. When you read disassembly you are reading two languages at once: the machine's three verbs, and the compiler's etiquette. This atlas teaches both.
fetch @ RIP 1–15 bytes, variable execute RIP += len MOVE mov movzx movsx lea push pop COMPUTE add sub imul and shr cmp test → RFLAGS JUMP jmp jcc call ret ← reads RFLAGS is one of
The whole machine: a fetch-execute loop over three instruction families. RFLAGS is the only channel between COMPUTE and JUMP.

1The register atlas: 16 boxes, 4 name widths, 1 trap

x86-64 gives you sixteen 64-bit general-purpose registers. Each one answers to four names depending on how much of it you're touching. Here is the anatomy of rax — all four names refer to the same physical register:

63 31 15 7 0 rax full 64 bits eax writing eax ZEROES all of this ax writing ax/al MERGES — upper bits survive (partial-register write) ah al ah = bits 15–8 · exists only for a/b/c/d; dies if a REX prefix is present
One physical register, four names. The red dashed zone is the #1 trap: 32-bit writes clear bits 63–32; 16- and 8-bit writes don't.

That asymmetry is not trivia — it's the single most common source of confusion when reading 64-bit disassembly, so we proved it live on your CPU. In maxi, right before mov eax, edi executes, we poisoned rax and rdi from GDB, then single-stepped the real instruction:

(gdb) set $rax = 0xdeadbeefcafebabe
(gdb) set $rdi = 0x11112222deadbeef
(gdb) x/i $rip
=> 0x5555555553f6 <maxi+6>:  mov  eax, edi
(gdb) stepi
after real mov eax,edi: rax = 0xdeadbeef        ← upper 32 bits wiped to zero

Only the low 32 bits of rdi were read, and the whole upper half of rax became zero. Two consequences you'll use constantly:

The full atlas — learn the roles, because the roles are how you decompile in your head:

6432168System V roleSurvives a call?What it means when you see it
raxeaxaxalreturn valueno — volatileafter a call: the result. Before syscall: the syscall number. al before varargs call: count of vector registers used.
rdiedididilargument 1noloaded right before a call = first argument. Dest for rep movs/stos.
rsiesisisilargument 2nosecond argument. Source for rep movs.
rdxedxdxdlargument 3nothird argument; high half of 128-bit results (rdx:rax in mul/div); remainder after div.
rcxecxcxclargument 4nofourth argument; cl = variable shift count; rep counter; destroyed by syscall (hardware stores return RIP there).
r8r8dr8wr8bargument 5nofifth argument.
r9r9dr9wr9bargument 6nosixth argument. Argument 7+ goes on the stack.
r10r10dr10wr10bscratchnoreplaces rcx as argument 4 in the kernel syscall convention; nested-function static chain.
r11r11dr11wr11bscratchnodestroyed by syscall (hardware stores RFLAGS there); favorite PLT/veneer scratch.
rbxebxbxblcallee-savedyes — preserveda value the function wants to keep across calls — usually a pointer or loop-carried variable. Watch the push rbx/pop rbx pair.
rbpebpbpbplcallee-savedyesframe pointer at -O0 (locals are [rbp-X]); just another preserved register at -O2.
r12–r15r12d…r12w…r12b…callee-savedyesmore long-lived locals. The number of push r1X in a prologue ≈ how many variables outlive calls.
rspespspsplstack pointeryes (by contract)always points at the last pushed qword. Must be 16-byte aligned at every call.
rip—instruction ptr—unreadable/unwritable directly; only call/jmp/ret change it, only [rip+disp] reads it.
fs—TLS base—fs:0x28 = the stack canary. Any fs: access = thread-local storage.
Why there's no sph/dih: the byte registers ah cH dh bh (bits 15–8) are leftovers from the 8086. Any instruction carrying a REX prefix (needed for r8–r15, spl, dil…) reinterprets those four encodings as spl bpl sil dil instead. So mov ah, r8b is physically unencodable — the two register families can't meet in one instruction. When you see ah in modern code it's almost always lahf-style flag juggling or byte-swapping tricks.

2One formula rules all memory: [base + index·scale + disp]

Every memory operand in x86-64 — every array access, struct field, local variable, jump table — is one instance of a single hardware formula. Your specimen's index_scale function compiles to exactly one instance of it:

long index_scale(long *arr, long i) { return arr[i]; }
    1344:  mov  rax, QWORD PTR [rdi + rsi*8]      ; arr in rdi, i in rsi
[ base  +  index × scale  +  disp ] base any register (or rip) yours: rdi = &arr[0] index any register except rsp yours: rsi = i scale 1, 2, 4 or 8 — nothing else yours: 8 = sizeof(long) disp constant, 8- or 32-bit yours: 0 (omitted) "which object" "which element" "element size" "which field" pointer + i·sizeof + offsetof   ⇒   arr[i] · p->field · matrix[r][c] Read every bracket as C: base names the object, index·scale walks elements, disp picks the field.
The universal address formula. Decompiling in your head = mapping each part back to C.

The three specimens, decoded with that lens:

; int struct_field(struct player *p) { return p->mp; }   — disp alone picks a field
    1354:  mov  eax, DWORD PTR [rdi+0x4]        ; mp is 4 bytes into the struct

; int matrix(int m[][10], long r, long c) { return m[r][c]; }  — r*10 doesn't fit scale∈{1,2,4,8},
; so the compiler builds ×10 out of two lea's: (r + r*4)*8 = r*40 = r * 10 ints
    1364:  lea  rax, [rsi+rsi*4]                ; rax = r*5
    1368:  lea  rax, [rdi+rax*8]                ; rax = m + r*40  (= &m[r][0])
    136c:  mov  eax, DWORD PTR [rax+rdx*4]      ; load m[r][c]

lea — the address formula hijacked for arithmetic

lea (load effective address) computes the bracket without touching memory — it's a 3-operand add-and-multiply that happens to use address syntax. Your lea_math:

long lea_math(long x) { return x * 5 + 7; }
    1374:  lea  rax, [rdi + rdi*4 + 0x7]        ; rax = x + x*4 + 7 — no RAM involved
The disambiguation rule you'll use every day: mov rax, [rdi] dereferences — it reads 8 bytes of memory. lea rax, [rdi+8] only does the math — it's rax = rdi + 8, memory untouched. If a bracket appears in lea, mentally erase the brackets. Compilers love it because it does mul+add+move in one instruction and — unlike add — doesn't clobber RFLAGS, so it can sit between a cmp and its jcc. You'll also constantly see lea rdi, [rip+0xXXX] — that's "take the address of this string/global", the position-independent way.

3How instructions are actually spelled: REX · opcode · ModRM · SIB

You don't need to hand-assemble, but knowing the byte anatomy pays off in RE: it's how you recognize function boundaries in raw bytes, spot misaligned disassembly, and understand why the same mnemonic has many lengths. Here is your own mov rax, [rdi+rsi*8] — four bytes, fully decoded:

48 8b 04 f7 REX prefix opcode: MOV r64, r/m64 ModRM SIB 0100 W=1 R=0 X=0 B=0 W: 64-bit operand  ·  R/X/B: high bit of reg / index / base (how r8–r15 are reached) mod=00 reg=000 rm=100 no disp dest: rax "SIB byte follows" s=11 i=110 b=111 ×8 rsi rdi 48 8b 04 f7  ⇒  mov rax, QWORD PTR [rdi + rsi*8]
Byte anatomy of index_scale's only real instruction. mod/reg/rm and scale/index/base are just the formula from Step 2, serialized.

The general template every instruction follows (most parts optional): [prefixes] [REX] opcode [ModRM] [SIB] [disp] [imm] — from 1 byte (ret = c3, push rbp = 55) to the architectural maximum of 15 bytes. Practical payoffs:

4RFLAGS & condition codes: the machine's only memory of a comparison

There is no "if" instruction. There is cmp (a subtraction that throws away the result but keeps the side effects) and, later, a conditional jump that inspects those side effects. The side effects live in RFLAGS. We captured a real one on your CPU — inside maxi, we forced edi=5, esi=3 and stepped over cmp esi, edi (computing 3−5):

(gdb) set $edi=5 ; set $esi=3 ; stepi       # executes: cmp esi, edi   → 3 - 5
eflags  0x293  [ CF AF SF IF ]              # ZF=0 OF=0
OF DF IF TF SF ZF AF PF CF 11 10 9 8 7 6 4 2 0 0 0 1 0 1 0 1 0 1 signed overflow string direction interrupts on single-step result negative result zero BCD carry low-byte parity unsigned borrow your capture: 3 − 5 → eflags 0x293 → CF=1 (3<5 unsigned) · SF=1 (result −2) · ZF=0 (not equal) · OF=0 Only four matter for reading code: ZF, CF, SF, OF. IF/TF/DF are OS & string machinery.
RFLAGS after your real cmp esi, edi with 3 and 5. One subtraction feeds every possible question about a vs b.
The crown jewel of x86 reverse engineering: condition codes leak the C types. The CPU doesn't know whether your bytes are signed or unsigned — but the jump mnemonic does. After cmp a, b, "greater/less" (jg jge jl jle) test the signed story (SF, OF), while "above/below" (ja jae jb jbe) test the unsigned story (CF). The compiler chose the mnemonic based on the C declaration — so when you read ja, you just learned the variable was unsigned (or a pointer/size_t); when you read jg, it was signed. Disassembly leaks the type system. Nothing tells you more about the original source for less effort.
After cmp a, b, jump if…Signed (g/l family)Flags testedUnsigned (a/b family)Flags tested
equal / not equalje (jz) · jne (jnz) — ZF, shared by both worlds
a > bjgZF=0 and SF=OFjaCF=0 and ZF=0
a ≥ bjgeSF=OFjae (jnc)CF=0
a < bjlSF≠OFjb (jc)CF=1
a ≤ bjleZF=1 or SF≠OFjbeCF=1 or ZF=1
sign / overflow / parityjs · jo · jp — direct single-flag tests (jp after float compare = NaN!)

Why "SF≠OF" means "less than", in one sentence: normally a negative result (SF=1) means a<b, but if the subtraction overflowed (OF=1) the sign is a lie, so the truth is the XOR of the two. Your capture shows the honest case: SF=1, OF=0 → 3<5 signed; and CF=1 → 3<5 unsigned too.

The same condition, three consumers

Every condition code cc exists in three instruction forms, and your specimen has all three:

; 1. jcc — branch on it (day_name's bounds check)
    1424:  cmp  edi, 0x6
    1427:  ja   14a8            ; unsigned trick: one ja handles d<0 AND d>6 (see Step 7)

; 2. cmovcc — select on it, no branch (maxi: return a > b ? a : b)
    13f4:  cmp  esi, edi
    13f6:  mov  eax, edi
    13f8:  cmovge eax, esi      ; if (b >= a) result = b — a ternary with zero branches

; 3. setcc — materialize it as 0/1 (is_zero: return x == 0)
    1404:  xor  eax, eax         ; clear first — sete only writes 1 byte (al)!
    1406:  test rdi, rdi
    1409:  sete al               ; al = (x == 0)
test vs cmp, the two flag-setters you'll see most: cmp a,b is a−b discarded; test a,b is a&b discarded. test reg, reg (same register twice) is the idiomatic "is it zero / is it negative?" — after it, je means "was zero" (NULL check after loading a pointer, error check after a call), js means "was negative". test al, al after a call = checking a returned bool/char. Also remember both xor reg,reg and sub/add/and/or update flags as a side effect, but mov, lea, push/pop never touch flags — that's why compilers can schedule them between a cmp and its jcc, and you must read past them to find which comparison a jump belongs to.

5The System V calling convention: the social contract

Nothing in the hardware defines "arguments". The System V AMD64 ABI (Linux, macOS, BSD — Windows differs!) says: first six integer/pointer args ride in rdi rsi rdx rcx r8 r9, first eight float/double args in xmm0–xmm7 (counted separately), the rest go on the stack right-to-left, and the return comes back in rax (or xmm0). We froze your many_args(1,2,3,4,5,6,7,8) at its first instruction and photographed the contract in action:

(gdb) break *many_args ; run
reg args:  rdi=1  rsi=2  rdx=3  rcx=4  r8=5  r9=6
0x7fffffffd6d8:  0x0000555555555186   ← [rsp]      return address (pushed by call)
0x7fffffffd6e0:  0x0000000000000007   ← [rsp+0x08] argument 7
0x7fffffffd6e8:  0x0000000000000008   ← [rsp+0x10] argument 8
many_args(a,…,h) rdi=1 rsi=2 rdx=3 rcx=4 r8=5 r9=6 integer/pointer args 1–6 (verified live above) [rsp+0x10] = 8 arg 8 [rsp+0x08] = 7 arg 7 [rsp] = return address args 7+ pushed right-to-left; callee reads, never pops xmm0=1.5 xmm1=0.25 rdi=4 mixed(double x, long n, double y): two counters run in parallel — doubles take xmm0, xmm1 while the long takes rdi ret: rax=36 ret: xmm0=6.25 integers return in rax (rdx:rax if 128-bit), floats in xmm0 rsp ≡ 16-aligned at every call  ·  al = # of xmm regs used, set before calling varargs (printf)  ·  leaf functions may use 128 bytes below rsp (the red zone) without moving it volatile (caller must not trust): rax rcx rdx rsi rdi r8–r11, all xmm  ·  preserved (callee must restore): rbx rbp rsp r12–r15
The System V contract with your live values: many_args got 1–6 in registers, 7 and 8 at [rsp+8] and [rsp+0x10], and returned 36 in rax.

The volatile/preserved split is where reverse engineering gets real leverage. Your callee_saved_demo needs a and b to survive the call to greet — watch what the compiler does:

long callee_saved_demo(long a, long b) { greet("demo"); return a*3 + b; }
    1584:  push rbp              ; save caller's rbp — I'm about to use it
    1585:  mov  rbp, rsi         ; b moves into a PRESERVED register
    1588:  push rbx              ; save caller's rbx too
    1589:  mov  rbx, rdi         ; a moves into a preserved register
    158c:  lea  rdi, [rip+0xa9b] ; arg1 = "demo"
    1593:  sub  rsp, 0x8         ; realign: 2 pushes + retaddr = 24 bytes → +8 = 32 ✔
    1597:  call 1520 <greet>     ; may destroy rax rcx rdx rsi rdi r8-r11…
    159c:  lea  rax, [rbx+rbx*2] ; …but rbx still holds a: rax = a*3
    15a4:  add  rax, rbp         ; + b, still alive in rbp
    15a7:  pop  rbx              ; restore in reverse order
    15a8:  pop  rbp
    15a9:  ret
Read functions backwards from this contract. Values marching into rdi/rsi/rdx… before a call are the arguments — that's how you recover prototypes of unnamed functions. A value moved into rbx/r12–r15 at the top is a long-lived local. eax examined right after a call is the return value being checked. And at -O2, note what's absent: no frame pointer, no wasted moves — rbp here is just a spare preserved register holding b, nothing to do with frames.

6Stack frames, the canary, and where everything lives

The stack grows downward (push = rsp -= 8 then store), and a function's frame is just the slice of stack between the return address and wherever rsp stops. Compare the two builds of greet — same C, two dialects:

;— -O0: the textbook frame —————————————;    ;— -O2: frame pointer optimized out ——————
push rbp                                      push rbp            ; (just preserving it)
mov  rbp, rsp     ; anchor the frame          ; args staged into r9, r8, rcx, rdx, rsi…
sub  rsp, 0x60    ; room for locals           sub  rsp, 0x50
mov  [rbp-0x58], rdi                          mov  rax, fs:0x28   ; canary in…
mov  rax, fs:0x28                             mov  [rsp+0x48], rax
mov  [rbp-0x8], rax   ; canary                ; locals addressed from RSP, not rbp
; locals at [rbp-X] — easy to read            mov  rbp, rsp       ; rbp = &buf (reused!)
…                                             …
leave             ; = mov rsp,rbp; pop rbp    add  rsp, 0x50
ret                                           pop  rbp
                                              ret

We froze the -O2 frame live, right after the canary store, and mapped every byte (all values real):

higher addresses ↑ (caller's territory) 0x55555555520f → main+335 (return address) saved rbp = 0x2 canary = 0x0d07838cdfa94700 8 bytes padding char buf[64] "hello, %s" lands here — attacker-influenced bytes red zone: 128 bytes below rsp, usable by leaf functions [rsp+0x58] [rsp+0x50] [rsp+0x48] [rsp]…+0x3f ← rsp = 0x7fffffffd6b0 overflow writes grow upward — must trample the canary before the return address epilogue: mov rax,[rsp+0x48] ; sub rax, fs:0x28 ; jne __stack_chk_fail — mismatch = abort, no ret
Your greet frame, photographed live. The canary's low byte is always 00 — it terminates C strings so string functions can't leak or write past it quietly.

And the widest view — where that frame sits in the whole process. This is info proc mappings from your actual run:

[stack] 0x7ffffffdd000 ↓ grows down ld-linux + [vdso] 0x7ffff7fc… libc.so.6 0x7ffff7c00000 (r/rx/r/rw) [heap] 0x555555559000 ↑ grows up specimen 0x555555554000 (PIE base) ← greet's frame at 0x7fffffffd6b0 lives here ← the dynamic linker; vsyscall page at 0xffffffffff600000 ← puts really lives here; PLT jumps land here ← ~46 TB of unmapped nothing between ← malloc's territory ← five mappings: r--/r-x/r--/r--/rw- = headers/code/rodata/relro/data These exact addresses repeat every GDB run because GDB disables ASLR — 0x555555554000 and 0x7ffff7… are its signature no-randomization bases. Outside GDB they shuffle every execution.
Your process, from info proc mappings. Code low, heap above it, libraries high, stack highest — with a canyon between.

7Control-flow shapes: if, loop, switch — learn the silhouettes

Compilers emit a small number of recognizable shapes. Once you know the silhouettes, you read structure instead of instructions.

The loop silhouette — sum_array

    14b4:  test rsi, rsi          ; ── guard: n <= 0?
    14b7:  jle  14d0              ;    skip loop entirely (early-out block below)
    14b9:  lea  rcx, [rdi+rsi*4]  ; ── setup: rcx = one-past-the-end pointer!
    14bd:  xor  eax, eax          ;    s = 0
    14c0:  movsxd rdx, DWORD PTR [rdi]   ; ─┐ body: load a[i] (sign-extend int→long)
    14c3:  add  rdi, 0x4          ;  │ pointer walks — the compiler killed i!
    14c7:  add  rax, rdx          ;  │ s += a[i]
    14ca:  cmp  rdi, rcx          ;  │ pointer == end?
    14cd:  jne  14c0              ; ─┘ BACKWARD jump = "this is a loop"
    14cf:  ret
    14d0:  xor  eax, eax          ; the n<=0 path: return 0
    14d2:  ret
guard: jle setup body14c0–14cd ret jne 14c0 — the backward edge IS the loop n ≤ 0: skip everything
Scan any function for backward conditional jumps first — each one is a loop, and its target is the loop head.

Two habits of optimized loops to recognize: the compiler replaced the index with a walking pointer (no i exists — rdi strides by 4 toward a precomputed end pointer), and the loop condition moved to the bottom (do-while shape) with a guard on top — cheaper than jumping to a test each iteration.

The switch silhouette — day_name's jump table

    1424:  cmp  edi, 0x6
    1427:  ja   14a8                       ; unsigned trick: d<0 wraps to huge → one test kills both sides
    1429:  lea  rdx, [rip+0xc1c]           # 204c — table base, in .rodata
    1430:  mov  edi, edi                    ; zero upper 32 (Step 1's "useless" mov!)
    1432:  movsxd rax, DWORD PTR [rdx+rdi*4]  ; fetch 32-bit OFFSET, not a pointer
    1436:  add  rax, rdx                    ; target = base + offset (PIE-friendly)
    1439:  notrack jmp rax                  ; indirect jump into the case blocks
edi = d cmp edi, 6 ja → default 14a8: "???" [204c] 0: −0xbfc → 1450 "Sun" [2050] 1: −0xc0c → 1440 "Mon" [2054] 2: −0xbec → 1460 "Tue" [2058] 3: −0xbdc → 1470 "Wed" [205c] 4: −0xbbc → 1490 "Thu" [2060] 5: −0xbac → 14a0 "Fri" [2064] 6: −0xbcc → 1480 "Sat" d ≤ 6 1460: lea rax,"Tue" ret jmp rax each case block: 16 bytes, nop-padded, order shuffled (Mon's block sits first!) table entries are signed 32-bit offsets from the table itself (movsxd + add rdx) — position-independent, half the size of pointers
Your jump table at .rodata:0x204c, decoded entry by entry. cmp/ja + lea rip-rel + movsxd + jmp rax = "this was a switch".
The ja bounds trick is a mini signed/unsigned masterclass: d is a signed int, but the compiler tests cmp edi,6; ja default — an unsigned compare. A negative d reinterpreted as unsigned becomes huge (≥ 0x80000000), so one unsigned test rejects both d > 6 and d < 0. When you see ja right before an indirect jump, it's a switch bounds check, and the case count is the immediate + 1.

The if/else silhouette barely needs a diagram once you've internalized Step 4: cmp/test → jcc over the "then" block, optional jmp over the "else". At -O2 short ifs usually vanish into cmov/setcc (branchless), so a real branch surviving in optimized code hints the compiler thought it unpredictable or the sides too big.

8Widths and extensions: how C's integer rules look in metal

Mixing widths is where silent bugs and RE misreadings live. The machine has exactly three widening strategies — and your specimen used all of them:

u32 in edi i32 in edi i8 at [rdi] 00000000│value ssssssss│value ssssssssssssss│v mov eax, edi — free zero-extend movsxd rax, edi — sign copies movsx rax, BYTE PTR [rdi] your three widen functions, verbatim: unsigned→mov, signed→movsx(d) — the mnemonic tells you the C type, again
The extension family. movzx exists for 8/16-bit sources; 32→64 zero-extension is just a plain 32-bit mov.

The in-place cousins act on the accumulator family only, and the division ritual depends on them:

InstructionEffectWhere you'll meet it
cbw / cwde / cdqeal→ax → eax → rax (sign, in place)widening a local before pointer math
cwd / cdq / cqosign of ax/eax/rax fills dx/edx/rdxthe herald of signed division — your absolute uses cqo
cqo ; idiv rcxsigned divide rdx:rax → quot rax, rem rdxcqo/cdq before idiv = signed / or %
xor edx,edx ; div rcxunsigned dividezeroed rdx before div = unsigned — type leak #3

9The compiler idiom Rosetta stone

Optimizing compilers speak in idioms — fixed phrases that mean something quite different from their literal instructions. This table is the heart of "reading assembly like C". Every row marked ● was pulled from your own specimen's disassembly:

You seeIt meansWhy / notes
● xor eax, eaxeax = 02 bytes vs 5 for mov eax,0; also breaks CPU dependency chains. Ubiquitous.
● test rdi, rdi ; je …if (!ptr) / if (x == 0)AND with itself only sets flags. js variant = "if negative".
● lea rax, [rdi+rdi*4]x * 5lea = mul/add without flags. ×3, ×5, ×9 direct; ×10 = two leas (your matrix).
● imul rax, rdi ; shr rax, 35
(after mov edi, 0xcccccccd)
x / 10Magic-number division — walkthrough below. Constant divisors never emit div.
● sar edx, 31sign mask: 0 or −1Arithmetic shift smears the sign bit; the all-ones mask feeds branchless tricks.
● cqo ; xor rax,rdx ; sub rax,rdxlabs(x)Your absolute: (x XOR mask) − mask negates iff mask=−1. No branch.
● (sar,shr,lea) ; and ; subx % 8 (signed)Signed modulo must round toward zero, so a bias built from the sign mask wraps the simple and eax,7. Unsigned % 2ⁿ is just the and.
● cmovge / sete / ja-vs-jgternary / bool / type leakSteps 4 & 7.
● mov rax, fs:0x28stack canaryAny fs: access = thread-local storage; offset 0x28 is glibc's canary slot.
● mov edi, edizero bits 63–32 of rdiNot a no-op! (Step 1.)
● endbr64 / nop WORD PTR […] / xchg ax,axnothingCET landing pads and alignment padding. Train your eyes to skip them.
shl rax, 4 ; add/or rax, …x*16 + yIndex into array of 16-byte structs; shifts are multiplies by 2ⁿ.
rep movsb / rep stosqmemcpy / memsetrdi=dest rsi=src rcx=count; DF decides direction (cld = forward).
pxor xmm0, xmm00.0 (or zero vector)The float twin of xor eax,eax — your mixed does it before cvtsi2sd.
add rsp, 8 … or sub rsp, 8 around a callalignment, not dataKeeping rsp ≡ 16 at the call (your callee_saved_demo).
mov reg, 0x7fffffff… / 0x80000000…INT_MAX / sign-bit maskBig round hex constants are almost always masks or limits — decode them.
Crown-jewel walkthrough: why x / 10 is a multiply. Division is ~20–40 cycles; multiplication ~3. So for any constant divisor the compiler precomputes a fixed-point reciprocal. Your div10:
    1384:  mov  eax, edi              ; zero-extend x to 64 bits (Step 1!)
    1386:  mov  edi, 0xcccccccd       ; ⌈2³⁵ / 10⌉ — "0.1" in fixed-point
    138b:  imul rax, rdi              ; 64-bit product = x · (2³⁵/10)
    138f:  shr  rax, 0x23             ; ÷ 2³⁵  ⇒  floor(x/10), exact for all u32
The math: 0xcccccccd = 3435973837 = ⌈2³⁵/10⌉, so (x·0xcccccccd) >> 35 = x·(2³⁵/10)/2³⁵ ≈ x/10, and the ceiling guarantees exactness for every 32-bit x. Verified live: your program printed div10(1234) = 123. RE recipe: when you meet imul with an ugly constant followed by a shift, the divisor is round(2^(32+shift) / magic) — for 0xcccccccd, shift 3 (35−32): 2³⁵/0xcccccccd ≈ 9.99999… → 10. Repeating-pattern constants are the tell: 0xcccccccd→10, 0xaaaaaaab→3 or 6, 0x92492493→7, 0x38e38e39→9 or 18.

10SIMD: xmm/ymm/zmm and how to decode the alphabet soup

All floating point in x86-64 goes through the vector registers — there is no separate "float unit" in normal compiled code (the ancient x87 stack only appears for long double). So even scalar double math lives in xmm0…, and the ABI returns floats in xmm0. The registers nest like the GPRs do:

511 255 127 0 zmm0 512 bits — AVX-512 (zmm0–31, +k0–7 masks) ymm0 256 — AVX xmm0 128 — SSE f7 f6 f5 f4 f3 f2 f1 f0 ymm as 8 packed floats (your dot product) — scalar (…ss/…sd) ops touch ONLY lane f0 and ignore the rest writing xmm with VEX (v-prefixed) zeroes the upper ymm/zmm — same spirit as eax→rax zero-extension
The vector register file. Same nesting idea as rax/eax/ax — including a zero-the-top rule.

The mnemonic decoder ring

v mul p s AVX/VEX form (absent = old SSE) operation add sub mul div sqrt min max cmp mov p = packed (all lanes) s = scalar (lane 0 only) s = single (float, 4B) d = double (8B) mulsd = scalar double · vmulps = AVX packed floats · addss = scalar float integer vectors spell it differently: p-prefix + width — paddd (4×i32), paddq (2×i64), pcmpeqb (16×i8)
Four glyphs decode 90% of vector mnemonics.

Scalar reality first — your mixed(double x, long n, double y) is pure lane-0 code:

    1504:  movapd xmm2, xmm0        ; save x (xmm0 is needed as the result reg)
    1508:  pxor   xmm0, xmm0        ; zero idiom — breaks false dependency
    150c:  cvtsi2sd xmm0, rdi       ; ConVerT Signed Int 2 Scalar Double: n → double
    1511:  mulsd  xmm0, xmm2        ; x * n     (scalar double)
    1515:  addsd  xmm0, xmm1        ; + y  → returned in xmm0 = 6.25 ✔ (printed by your run)

Packed reality — dot() from simd.c built with -O3 -mavx2. Eight multiplications happen in one instruction; the rest is folding the lanes down to one scalar:

    154:  vmovups xmm4, [rsi]                    ; load b[0..3]
    160:  vinsertf128 ymm0, ymm5, [rdi+0x10], 1  ; build a[0..7] in one ymm
    16e:  vmulps  ymm1, ymm1, ymm0               ; a[i]*b[i] — ×8 AT ONCE
    172:  vaddss  xmm3, xmm3, xmm1               ; ── horizontal reduction begins:
    176:  vshufps xmm2, xmm1, xmm1, 0x55         ;    shuffle lane k to lane 0,
    …     (vshufps / vunpckhps / vextractf128 …) ;    add, repeat for all 8 lanes
    1b4:  vzeroupper                             ; drop ymm halves before returning

Field notes for reading vector code:

11The startup chain & RE workflow: where the code you care about hides

Your main is not the entry point. The ELF header's e_entry points at _start; main's address is merely handed to __libc_start_main as an argument. Straight from your _start:

    1254:  xor  ebp, ebp              ; mark the outermost frame (rbp=0 ends backtraces)
    1256:  mov  r9,  rdx              ; rtld_fini
    1259:  pop  rsi                   ; argc  (kernel put it on the stack)
    125a:  mov  rdx, rsp              ; argv
    125d:  and  rsp, 0xfffffffffffffff0  ; 16-align the stack (the ABI mandate)
    1268:  lea  rdi, [rip+0xfffffffffffffe51]  # 10c0 <main> — main is just an ARGUMENT
    126f:  call [rip+0x2d63]          # __libc_start_main — never returns
    1275:  hlt
kernelexecve, mmap ld-linuxrelocs, libc _start0x1250 __libc_start_mainctors, INIT_ARRAY main0x10c0 exit In any new binary, set your first breakpoint on main, not the entry point — skip the boilerplate.
Startup path for your specimen. Constructors (__attribute__((constructor)), C++ static init) run in __libc_start_main — malware hides here to execute before main.
The reverse-engineering loop, distilled to a toolkit you can run on this specimen right now:
file specimen                          # arch, PIE?, stripped?, static/dynamic
objdump -d -M intel specimen            # the static disassembly (Intel syntax!)
readelf -h -l -d specimen               # entry point, segments, libraries needed
strings -a -t x specimen                # literals + their offsets — the fastest clues
nm -C specimen  /  readelf --dyn-syms    # symbols, imports (UND = comes from libc)
ltrace ./specimen self                   # library-call trace: the API storyline
strace ./specimen self                   # syscall trace: what it does to the OS
gdb -q ./specimen                        # dynamic: break main, stepi, x/i $pc, watch
  gef> break main ; run ; context         # GEF/pwndbg give regs+stack+code at a glance
Static tells you what could happen; dynamic tells you what did. Alternate between them. For serious work, a disassembler with a decompiler (Ghidra is free, IDA/Binary Ninja commercial) turns all of Steps 4–10 into pseudo-C automatically — but only someone who can read the raw form knows when the decompiler is lying.

Tips, tricks & gotchas most tutorials skip

★Exercises to cement it — all runnable on your specimen

  1. Prove the 32-bit zero-extension yourself: gdb -q ./specimen -ex 'break *maxi+6' -ex run -ex 'set $rax=0xffffffffffffffff' -ex stepi -ex 'p/x $rax'. Watch mov eax,edi blow away the top half. Then try the same idea after an ax-writing instruction and confirm it does not.
  2. Recover a prototype cold: objdump -d -M intel specimen | less, jump to many_args, and reconstruct its signature purely from which registers/stack slots feed the adds. Check against the source.
  3. Read the flags live: break at *maxi+4, set edi/esi to different signed and unsigned extremes (e.g. 0x80000000 vs 1), step the cmp, and predict eflags before revealing it. Feel why jg and ja disagree on the same bytes.
  4. Decode a magic divisor: add unsigned d7(unsigned x){return x/7;} to specimen, recompile, and reverse the emitted imul constant + shift back to "7" using the recipe in Step 9. Try a signed version and spot the extra sign-fix instructions.
  5. Walk the jump table: objdump -s -j .rodata specimen, find the 7 little-endian dwords at 0x204c, add each (as a signed offset) to 0x204c, and confirm the targets match the lea rax,[rip+…] case blocks in day_name. You'll see the cases are physically out of order.
  6. Watch a canary die: break in greet after the canary store, set a byte inside buf's canary slot, continue, and watch __stack_chk_fail abort the program instead of returning — the mitigation firing in real time.
  7. Break the disassembler: objdump -d -M intel --start-address=0x1345 specimen (one byte into a real instruction) and watch plausible garbage appear — the variable-length desync from Step 3, first-hand.
  8. See SIMD widen: compile simd.c at -O3 -mavx2 vs -O3 -mno-avx vs -O0 and diff dot. One math loop becomes eight-wide, four-wide, or scalar — the compiler's vectorization decision made visible.