ArrayLocalVariable.c — arrays on the stack

OST2 Arch1001 · adding and accessing an array local variable · new instructions: imul · movsx · movzx (+ movsxd)
RIP = 00000001`40001000 — no instruction yet executed
☒executed instruction
♍modified value
⌘start value
RIP▶next instruction
NEWfirst appearance (red in the slides)

C source

Disassembly

⤴ execution continues at 00000001`40001379 — the code that called main()

Registers

rax
rcx
rsp

Stack (4-byte granularity, like the course diagram)

↑higher addresses
lower addresses (stack grows down)↓
rsp

Layout tell: declared a, b, c — stored a, c, b. Locals need not keep their source order. &b[i] = rsp + 10h + i·4.

←/→ step · Space play/pause · Home/End jump

ArrayLocalVariable.c takeaways

Variant 2 — change short a → int a, add a (short)c cast

The takeaways slide steps a modified version. Two tiny C edits, and the compiler's choice of extension instructions changes completely (differences highlighted):

Variant 1 (stepped above)

short main(){
    short a;
    int b[6];
    long long c;
    a = 0xbabe;
    c = 0xba1b0ab1edb100d;
    b[1] = a;
    b[4] = b[1] + c;
    return b[4];
}

mov   eax, 0FFFFBABEh
mov   word ptr [rsp], ax
…
movsx ecx, word ptr [rsp]      ; load short a
mov   dword ptr [rsp+rax+10h], ecx
…
movsxd rax, dword ptr [rsp+rax+10h]
add   rax, qword ptr [rsp+8]   ; 64-bit add
…
b[4] = 1EDACACB  →  returns AX = CACB

Variant 2 (takeaways slide)

short main(){
    int a;
    int b[6];
    long long c;
    a = 0xbabe;
    c = 0xba1b0ab1edb100d;
    b[1] = a;
    b[4] = b[1] + (short)c;
    return b[4];
}

mov   dword ptr [rsp], 0BABEh  ; int a: plain dword store
…
mov   ecx, dword ptr [rsp]     ; plain load, no movsx
mov   dword ptr [rsp+rax+10h], ecx
…
movsx ecx, word ptr [rsp+8]    ; (short)c = low word 100Dh
add   ecx, dword ptr [rsp+rax+10h] ; 32-bit add
…
b[4] = 0000CACB  →  returns AX = CACB

IMUL — Signed Multiplyinstruction ⭐ #9 · book p.65

FYI: Visual Studio seems to have a predilection for imul over mul (unsigned multiply). You'll see it showing up in places you expect mul — Xeno reports never getting MSVC to emit plain mul for simple examples. (That's a fingerprint, by the way.)

Three forms — one, two, or three operands. Three operands? Possibly the only "basic" (non-added-on-instruction-set: MMX/AVX/VMX/etc.) instruction of its kind — and note the destination doesn't have to be one of the sources.

Group 1 — single operand (widening: full product kept)

formeffect
IMUL r/m8AX = AL × r/m8
IMUL r/m16DX:AX = AX × r/m16
IMUL r/m32EDX:EAX = EAX × r/m32
IMUL r/m64RDX:RAX = RAX × r/m64

The result register is twice the operand width, so nothing is lost. The other groups aren't so lucky:

Group 2 — two operands

IMUL r16, r/m16r16 = r16 × r/m16
IMUL r32, r/m32r32 = r32 × r/m32
IMUL r64, r/m64r64 = r64 × r/m64
⚠ TRUNCATION WARNING! ⚠ — the full product may not fit; only the low half is kept (CF/OF signal the loss)

Group 3 — three operands, 8-bit immediate

IMUL r16, r/m16, imm8r16 = r/m16 × sign-extended imm8
IMUL r32, r/m32, imm8r32 = r/m32 × sign-extended imm8
IMUL r64, r/m64, imm8r64 = r/m64 × sign-extended imm8
⚠ TRUNCATION WARNING! ⚠

(Group 4 is the same idea with a 16-bit immediate: IMUL r16, r/m16, imm16.)

Group 5 — three operands, 32-bit immediate

IMUL r32, r/m32, imm32r32 = r/m32 × imm32
IMUL r64, r/m64, imm32r64 = r/m64 × sign-extended imm32
⚠ TRUNCATION WARNING! ⚠

IMUL worked examples — with the math spelled out

All three start from the same registers: r12 = 0x84, rax = 0x609966C1A977E177. (In the video these run live in a Visual Studio myasm.asm scratchpad — comment/uncomment one imul, breakpoint on ret, watch the Registers window.)

Example 1 — imul r12b (the "hard math" one)

Group 1: AX = AL × r/m8 — everything is signed
r12
0x84
rax
0x609966C1A977E177
↓  imul r12b  ↓
r12
0x84
rax
0x609966C1A977C65C
AL = 0x77 = +119 r12b = 0x84 (sign bit set!) → 0x84 − 0x100 = −124 (the slide's calculator shows this as 0xFFFFFFFFFFFFFF7C = −124) +119 × −124 = −14,756 −14,756 in 16-bit two's complement: 0x10000 − 14,756 = 50,780 = 0xC65C AX ← 0xC65C … and ONLY AX: 16-bit writes never zero-extend, so bits 16–63 of rax are untouched.

Example 2 — imul r12d, eax

Group 2: r32 = r32 × r/m32 — truncation in action
r12
0x84
rax
0x609966C1A977E177
↓  imul r12d, eax  ↓
r12
0x0000000061D0415C
rax
0x609966C1A977E177
r12d = 0x84 = +132 eax = 0xA977E177 (sign bit set) = −1,451,761,289 +132 × −1,451,761,289 = −191,632,490,148 = 0xFFFFFFD3`61D0415C (full 64-bit) Truncated to 32 bits → 0x61D0415C — positive-looking! The sign lived in the discarded upper half. That's the TRUNCATION WARNING made real. 32-bit register write → upper half of r12 zero-extended: 0x00000000`61D0415C.

Example 3 — imul r12, rax, 12341234h

Group 5: r64 = r/m64 × sign-extended imm32 — dest ≠ source!
r12
0x84
rax
0x609966C1A977E177
↓  imul r12, rax, 12341234h  ↓
r12
0xE5A3577504602A2C
rax
0x609966C1A977E177
imm32 = 0x12341234 (positive, sign-extension changes nothing) 0x609966C1A977E177 × 0x12341234 = a 91-bit product — 27 bits too big. Kept: low 64 bits = 0xE5A3577504602A2C (r12's old 0x84 is simply replaced: with 3 operands the destination doesn't participate in the multiply) Note rax is untouched — three-operand imul is the only "basic" instruction that works like this.

MOVZX / MOVSX — move with zero / sign extendinstructions ⭐ #10 & #11 · book p.53

mov eax, 0F00DFACEh movzx → rbx = 0x00000000`F00DFACE (high bits ← 0, always) movsxd rbx, eax → rbx = 0xFFFFFFFF`F00DFACE (high bits ← 1, because 0xF00DFACE's sign bit is 1)

Encoding nuance (the slide glosses this): there is no movzx r64, r/m32 — a plain mov ebx, eax already zero-extends (§3.4.1.1, as ever). That's why only the sign-extending 32→64 form needs its own mnemonic: movsxd. movzx/movsx exist for 8-bit and 16-bit sources.

All three appear in this lecture's listing: movsx ecx, word ptr [rsp] (short → int, signed), movsxd rax, dword ptr […] (int → long long, signed), movzx eax, word ptr […] (16 defined bits are all a short-returning function owes its caller).

Curiosity corner — compiler fingerprinting 🔍

"Visual Studio likes imul" is more than trivia: compilers have styles, and with practice you can often name the compiler (sometimes the version and flags) from a page of disassembly. Reverse engineers use this to pick the right FLIRT/type signatures, date a binary, or cluster malware families. Some tells:

signalMSVC (this course)GCCClang
multiply idiom imul everywhere, even for unsigned; unoptimized index math as mov eax,4; imul rax,rax,1 — literally ×1 prefers lea tricks for small constants (lea eax,[rax+rax*4] = ×5); imul only when lea/shift can't similar to GCC; at -O0 often still folds the ×1 away
-O0 frame style no frame pointer: locals addressed [rsp+disp]; frame sizes like 0x28/0x38/0x58 (shadow space + 16-byte chunks + 8) push rbp; mov rbp,rsp; locals at [rbp−disp]; leaf functions may use the 128-byte red zone with no sub rsp at all rbp frame too, but different slot-assignment order; often materializes a dedicated "return value" slot
calling convention Win64: args in rcx,rdx,r8,r9, caller reserves 32-byte shadow space System V: args in rdi,rsi,rdx,rcx,r8,r9, no shadow space, red zone below rsp — the ABI alone splits Windows from Linux/macOS builds instantly
stack protection __security_cookie / __security_check_cookie (/GS) __stack_chk_fail, canary at fs:0x28 __stack_chk_fail, slightly different canary placement
metadata artifacts PE Rich header (undocumented linker/tool build IDs — a whole fingerprint on its own); .pdata unwind info; CRT names like mainCRTStartup ELF .comment section literally says GCC: (GNU) 13.2.0; endbr64 prologues on CET distros; __libc_start_main .comment says clang version …; distinctive ltmp/ODR artifacts; same _start chain as GCC on Linux
big-picture idioms division-by-constant magic numbers, memcpy expansion style, jump-table shape, and instruction scheduling all differ per compiler and per version — that's what automated provenance tools classify on same categories, different fingerprints — diff the same C on all three in Compiler Explorer and the styles jump out in minutes