| rax | |
| rcx | |
| rsp |
Layout tell: declared a, b, c — stored a, c, b. Locals need not
keep their source order. &b[i] = rsp + 10h + i·4.
a, b, c; laid out
a, c, b.)int) times the index desired —
hence mov eax,4 + imul rax,rax,index, even for index 1.movsxd widened b[1] to 64 bits to add it to
long long c).short a → int a, add a (short)c castThe takeaways slide steps a modified version. Two tiny C edits, and the compiler's choice of extension instructions changes completely (differences highlighted):
short main(){ short a; int b[6]; long long c; a = 0xbabe; c = 0xba1b0ab1edb100d; b[1] = a; b[4] = b[1] + c; return b[4]; } mov eax, 0FFFFBABEh mov word ptr [rsp], ax … movsx ecx, word ptr [rsp] ; load short a mov dword ptr [rsp+rax+10h], ecx … movsxd rax, dword ptr [rsp+rax+10h] add rax, qword ptr [rsp+8] ; 64-bit add … b[4] = 1EDACACB → returns AX = CACB
short main(){ int a; int b[6]; long long c; a = 0xbabe; c = 0xba1b0ab1edb100d; b[1] = a; b[4] = b[1] + (short)c; return b[4]; } mov dword ptr [rsp], 0BABEh ; int a: plain dword store … mov ecx, dword ptr [rsp] ; plain load, no movsx mov dword ptr [rsp+rax+10h], ecx … movsx ecx, word ptr [rsp+8] ; (short)c = low word 100Dh add ecx, dword ptr [rsp+rax+10h] ; 32-bit add … b[4] = 0000CACB → returns AX = CACB
int a → the store is a full dword 0BABEh = 0x0000BABE. As an
int, 0xbabe stays the positive 47,806 — no sign-extended
0FFFFBABEh constant, and b[1] = a becomes a plain int-to-int
mov instead of movsx. So variant 2's b[1] is
0000BABE, not FFFFBABE.(short)c → take only c's low word: movsx ecx, word ptr [rsp+8]
grabs 0x100D (positive → sign-extends to 0000100D) and the addition happens in 32-bit
ecx, not 64-bit rax. b[4] = 0x0000BABE + 0x0000100D = 0x0000CACB.FYI: Visual Studio seems to have a predilection for imul over
mul (unsigned multiply). You'll see it showing up in places you expect
mul — Xeno reports never getting MSVC to emit plain mul for simple
examples. (That's a fingerprint, by the way.)
Three forms — one, two, or three operands. Three operands? Possibly the only "basic" (non-added-on-instruction-set: MMX/AVX/VMX/etc.) instruction of its kind — and note the destination doesn't have to be one of the sources.
| form | effect |
|---|---|
| IMUL r/m8 | AX = AL × r/m8 |
| IMUL r/m16 | DX:AX = AX × r/m16 |
| IMUL r/m32 | EDX:EAX = EAX × r/m32 |
| IMUL r/m64 | RDX:RAX = RAX × r/m64 |
The result register is twice the operand width, so nothing is lost. The other groups aren't so lucky:
| IMUL r16, r/m16 | r16 = r16 × r/m16 |
| IMUL r32, r/m32 | r32 = r32 × r/m32 |
| IMUL r64, r/m64 | r64 = r64 × r/m64 |
| IMUL r16, r/m16, imm8 | r16 = r/m16 × sign-extended imm8 |
| IMUL r32, r/m32, imm8 | r32 = r/m32 × sign-extended imm8 |
| IMUL r64, r/m64, imm8 | r64 = r/m64 × sign-extended imm8 |
(Group 4 is the same idea with a
16-bit immediate: IMUL r16, r/m16, imm16.)
| IMUL r32, r/m32, imm32 | r32 = r/m32 × imm32 |
| IMUL r64, r/m64, imm32 | r64 = r/m64 × sign-extended imm32 |
All three start from the same registers: r12 = 0x84,
rax = 0x609966C1A977E177. (In the video these run live in a Visual Studio
myasm.asm scratchpad — comment/uncomment one imul, breakpoint on
ret, watch the Registers window.)
MOV.Encoding nuance (the slide glosses this): there is no
movzx r64, r/m32 — a plain mov ebx, eax already zero-extends
(§3.4.1.1, as ever). That's why only the sign-extending 32→64 form needs its own mnemonic:
movsxd. movzx/movsx exist for 8-bit and 16-bit sources.
All three appear in this lecture's listing:
movsx ecx, word ptr [rsp] (short → int, signed),
movsxd rax, dword ptr […] (int → long long, signed),
movzx eax, word ptr […] (16 defined bits are all a short-returning
function owes its caller).
"Visual Studio likes imul" is more than trivia: compilers have styles, and
with practice you can often name the compiler (sometimes the version and flags) from a page of
disassembly. Reverse engineers use this to pick the right FLIRT/type signatures, date a binary,
or cluster malware families. Some tells:
| signal | MSVC (this course) | GCC | Clang |
|---|---|---|---|
| multiply idiom | imul everywhere, even for unsigned; unoptimized index math as
mov eax,4; imul rax,rax,1 — literally ×1 |
prefers lea tricks for small constants (lea eax,[rax+rax*4] = ×5);
imul only when lea/shift can't |
similar to GCC; at -O0 often still folds the ×1 away |
| -O0 frame style | no frame pointer: locals addressed [rsp+disp]; frame sizes like 0x28/0x38/0x58
(shadow space + 16-byte chunks + 8) |
push rbp; mov rbp,rsp; locals at [rbp−disp]; leaf functions may use
the 128-byte red zone with no sub rsp at all |
rbp frame too, but different slot-assignment order; often materializes a
dedicated "return value" slot |
| calling convention | Win64: args in rcx,rdx,r8,r9, caller reserves 32-byte shadow space |
System V: args in rdi,rsi,rdx,rcx,r8,r9, no shadow space, red zone
below rsp — the ABI alone splits Windows from Linux/macOS builds instantly |
|
| stack protection | __security_cookie / __security_check_cookie (/GS) |
__stack_chk_fail, canary at fs:0x28 |
__stack_chk_fail, slightly different canary placement |
| metadata artifacts | PE Rich header (undocumented linker/tool build IDs — a whole fingerprint on its own);
.pdata unwind info; CRT names like mainCRTStartup |
ELF .comment section literally says GCC: (GNU) 13.2.0;
endbr64 prologues on CET distros; __libc_start_main |
.comment says clang version …; distinctive
ltmp/ODR artifacts; same _start chain as GCC on Linux |
| big-picture idioms | division-by-constant magic numbers, memcpy expansion style, jump-table shape, and instruction scheduling all differ per compiler and per version — that's what automated provenance tools classify on | same categories, different fingerprints — diff the same C on all three in Compiler Explorer and the styles jump out in minutes | |
readelf -p .comment ./binary on ELF;
strings | grep -i "GCC:"; on PE, parse the Rich header — see
"The devil's in the Rich header" (used to attribute OlympicDestroyer's
false flag).