An interactive tour of process memory on x86-64 Linux โ with real-looking addresses.
All addresses below are typical values you would see in gdb (ASLR changes them each run, but the shape is always this).
Yes and no. A Stack you write with malloc is just bytes in the heap โ you can make 1000 of them. But THE stack is one special region the kernel creates per thread, and the CPU itself helps manage it: the RSP register and the push / pop / call / ret instructions are hardware support for exactly one stack at a time. See section 4.
One stack per thread. Thread stacks are just mmap'd blocks in the middle of the address space. All threads share the heap, globals and code โ each thread only gets a private stack + private registers. See section 3.
It doesn't โ that's the trick. Functions live in the code segment, not the stack. The stack only holds the frames of calls that haven't returned yet โ usually a chain of 5โ50 frames, a few KB. Every ret frees a frame instantly, and the same bytes get reused by the next call. Big data goes to the heap. Watch it happen in section 2.
This is the "ultimate shape" you asked for. Hover / tap each segment. High addresses on top, like in your slide.
main() calls square_sum(3,4) which calls add(3,4).
Step through and watch frames get pushed, used, and recycled. Keys: โ โ
The stack survives because frames are constantly freed by ret. Break that contract with infinite recursion and 8 MiB fills up:
Click pthread_create() and watch new stacks appear in the mmap region. Each thread runs a different call chain at the same time โ that's only possible because each has a private RSP.
pthread_create with plain mmap() โ they're ordinary memory, just used as a stack because the new thread's RSP points into them.Stack struct vs. THE stackThe word "stack" is overloaded โ this is the root of your confusion. Both of these exist at the same time, in different places:
Created by the kernel (not by you) ยท one per thread ยท managed by the CPU's RSP register and push/pop/call/ret instructions ยท used automatically for every function call, local variable, and return address. You never malloc it and never free it.
Machine code sits in .text forever. A function only borrows stack space (a frame) while a call to it is in flight, and gives it back at ret. 10,000 functions in your app โ 10,000 frames โ only the current call chain exists.
You saw the same addresses (0x7FFFFFFFDF38โฆ) reused by add's dead frame, then overwritten later. That reuse is why one small region handles millions of calls โ and why reading uninitialized locals gives garbage.
Per process: code + globals + one heap + mmap area. Per thread: one stack + one register set. Your 1000 Stack structs? All just heap bytes. THE stack is the one with hardware and kernel support.