← Part 1: 32-bit · Part 2 of 2
In 64-bit mode the hardware forces the flat model. This part covers what still matters: CS as the mode-and-privilege switch, FS and GS as per-thread and per-CPU pointers, and the paging that does everything else. Values marked measured were recorded on this Linux 5.10 machine.
Part 1 ended with every serious 32-bit OS using the flat model: all segments at base 0 with a 4 GiB limit, paging doing the real work, and only FS or GS pointing at per-thread data. When AMD designed x86-64 in 2003, they turned that convention into hardware rules:
| Thing | 32-bit protected mode | 64-bit mode |
|---|---|---|
| CS / DS / ES / SS base | from the descriptor | always 0; the descriptor base is ignored |
| Segment limits | checked on every access | not checked |
| FS / GS base | 32-bit, from the descriptor | 64-bit, stored in MSRs, independent of the GDT |
| Null selector in DS/ES/FS/GS | loadable, faults when used | loadable and usable (the base is 0 anyway) |
| What CS still decides | code base/limit, 16/32-bit, CPL | 64-bit vs 32-bit decoding (L bit) and CPL |
| Instruction pointer | EIP | RIP, with RIP-relative addressing |
| Address size | 32-bit linear → 32/36-bit physical | 48-bit linear → up to 52-bit physical |
Every x86-64 CPU still powers on in 16-bit real mode, just like in 1978. Firmware walks it up the modes by setting control bits and reloading CS. Long mode is active once paging is on and the EFER.LME bit is set. From then on, each code segment's L bit chooses between two sub-modes: 64-bit mode (L = 1) and compatibility mode (L = 0) for running 32-bit programs.
Two things to notice. First, paging is mandatory in long mode: you cannot turn on 64-bit mode without page tables, which is exactly why segmentation could be simplified. Second, the last step into 64-bit code is a CS reload. The mode lives in CS.
Machine code is just bytes, and the same bytes mean different instructions depending on the current code segment. In 64-bit mode the bytes 0x40–0x4F are REX prefixes, which select 64-bit operands and registers R8–R15. In 32-bit mode those same bytes are inc and dec instructions. On Linux, CS 0x33 has L = 1 and CS 0x23 has L = 0, D = 1.
Because CS alone decides the mode, a 64-bit process can run 32-bit code by jumping to a CS with L = 0, and jump back the same way. This is allowed in user mode because both code segments have DPL 3. The program below does that. It was tested on this machine and exits with status 35 = 0x23, the CS value it read while in 32-bit mode measured:
; gate.asm — nasm -f elf64 gate.asm && ld gate.o -o gate && ./gate; echo $? -> 35 bits 64 global _start _start: lea rax, [rel compat] push 0x23 ; new CS: 32-bit user code (L=0, D=1) push rax ; new RIP retfq ; far return pops RIP then CS -> compatibility mode bits 32 compat: mov ebx, cs ; ebx = 0x23 (we are executing 32-bit code now) jmp 0x33:back64 ; far jump back to the 64-bit code segment bits 64 back64: mov edi, ebx ; exit status = 0x23 mov eax, 60 syscall
ld binary loads below 4 GiB (at 0x401000), and compatibility mode only forms 32-bit addresses. The same selector-swap trick is used by 32-bit Windows to reach 64-bit code, where it is nicknamed "Heaven's Gate."An x86-64 memory operand looks like seg:[base + index*scale + disp]. The CPU first computes the effective address from the registers, then adds the segment base to get the linear address. In 64-bit mode the base is 0 for CS, DS, ES and SS even with an explicit prefix, and comes from an MSR for FS and GS:
effective = base + index*scale + displacement ; or RIP + displacement
linear = effective + { FS_BASE if the fs: prefix is present
GS_BASE if the gs: prefix is present
0 otherwise (cs: ds: es: ss: prefixes add nothing) }
then: the linear address must be canonical, or the CPU raises #GP (chapter 11)
This is why the flat model stops being a choice. Even if you deliberately load a data selector whose descriptor has a non-zero base, the CPU ignores it. The only segment bases that exist in 64-bit mode are the two MSR-backed ones.
The GDT still exists, and CS/SS selectors still index into it, but the entries are simpler because base and limit no longer matter for code and data. Linux keeps one GDT per CPU. Click a row to see what each selector is for:
0x33, user data 0x2b. The odd numeric order is forced by the syscall/sysret instructions, explained in chapter 7.Only a handful of these are ever loaded into a segment register. In user space you will always see CS = 0x33 and SS = 0x2b, with DS, ES, FS and GS holding 0. The FS and GS bases are set separately (chapter 10), not by loading one of these selectors.
A common question: if a program is position-independent, how does it get the right CS? The answer is that CS has nothing to do with where your code lives. Its base is 0 for everyone. The kernel assigns the same selector value to every process, and the CPU loads it on the return to user mode. Position independence is handled by paging and RIP-relative addressing, never by CS.
iretq loads them.start_thread(), the kernel writes regs->cs = 0x33 and regs->ss = 0x2b into the register frame it will restore, and clears DS/ES/FS/GS to 0. The CPU makes them live when it returns to user mode. Every 64-bit process on the machine starts with the identical CS, because CS no longer encodes an address.There is no mov cs, ... instruction (it raises #UD). CS changes only through far jump/call/return, interrupts, iret, and the syscall/sysret pair. All of those change CS and RIP together, which is the whole point: the mode and the next instruction must switch atomically.
Entering the kernel through the old int gate is slow. x86-64 added syscall/sysret, which switch privilege without touching memory: the new CS and SS come straight from the STAR MSR, not from the GDT. This is why Linux's selectors are ordered the way they are.
syscall : CS = STAR[47:32] SS = STAR[47:32] + 8 ; kernel 0x10, 0x18 sysret : CS = STAR[63:48] + 16 SS = STAR[63:48] + 8 ; user 0x33, 0x2b STAR[63:48] = 0x23 -> SS = 0x2b, CS = 0x33
But a problem appears the moment the kernel starts running: it needs a pointer to its own per-CPU data, and every register including GS still holds the user's values. The answer is swapgs, a single instruction that exchanges GS_BASE with a hidden MSR (KERNEL_GS_BASE). The kernel runs swapgs on entry and again on exit.
You can still read and even load these registers, but it changes nothing about addressing. Reading is unprivileged, loading is checked against the GDT exactly as in 32-bit mode, and then the base is discarded. The following was run on this machine, showing which loads succeed and which fault measured:
0x2b (user data) and 0x33 (readable user code) succeed. 0x10 (kernel data, DPL 0) and 0x50 (past the table) fault with SIGSEGV. Null into SS faults; null into DS/ES/FS/GS is fine. Even on success, the base stays 0.So the tutorial phrase "the four data segments keep data elements from overlapping" no longer applies at all. DS, ES and SS have base 0 and no limit, so they cover the entire address space and overlap completely. Keeping data apart is the job of page permissions.
FS earns its keep. glibc points FS_BASE at each thread's thread control block (TCB). A single instruction with the fs: prefix then reaches per-thread data with no pointer chasing. Two uses appear in almost every compiled program:
mov rax, fs:0x28 ; the stack-protector canary (start of most functions) mov rax, fs:0 ; pointer to the TCB itself -> pthread_self() mov eax, fs:0xfffffffffffffffc ; a __thread int, i.e. fs:-4
The magic is that the same instruction bytes read different memory in different threads, because each thread has its own FS_BASE. On every context switch the kernel reloads it. This was measured on this machine with a two-thread program measured: the two threads reported different FS bases but the identical canary value, and a __thread variable sat at fs:-4 in both.
00 stops C string functions from copying past it, so a stray strcpy cannot leak or overwrite the whole value. __thread variables live just below the TCB, at negative offsets from FS_BASE.Since the base is a 64-bit MSR, not a descriptor, loading a selector cannot set it. There are two ways:
arch_prctl (syscall 158). The classic route. glibc calls arch_prctl(ARCH_SET_FS, tcb) once per thread during startup. This was verified on this machine: after mov fs, 0x2b, the FS base actually dropped to 0, because loading a selector overwrites the MSR-backed base with the descriptor's (zero) base measured.
wrfsbase / wrgsbase. If the CPU supports FSGSBASE and the kernel enabled it (CR4.FSGSBASE), user code can write its own FS base in one unprivileged instruction, no syscall. This CPU and kernel 5.10 support it, reported through AT_HWCAP2 = 0x2 measured.
// set FS base from user space, two ways #include <asm/prctl.h> #include <sys/syscall.h> syscall(SYS_arch_prctl, ARCH_SET_FS, (unsigned long)my_tcb); // always works unsigned long b; __asm__("rdfsbase %0" : "=r"(b)); // read (needs FSGSBASE) __asm__("wrfsbase %0" :: "r"(new_base)); // write (needs FSGSBASE)
errno all read through it. Point it somewhere wrong and the next protected function will fail its canary check and abort.A 64-bit register can name 264 addresses, an absurd 16 EiB. Current CPUs only wire up 48 bits (256 TiB), or 57 with the newer 5-level paging. To keep the unused top bits from being misused, x86-64 requires addresses to be canonical: bits 63 down to 47 must all be copies of bit 47. That splits the usable space into a low half and a high half with a vast non-canonical hole between them.
Linux gives user space the low half (up to 0x00007fffffffffff, 128 TiB) and the kernel the high half. Because the CPU rejects non-canonical addresses immediately, the classic 32-bit "3 GiB user / 1 GiB kernel" squeeze is gone. There is no 4 GiB wall and no PAE-style trickery: physical frames can sit anywhere the 52-bit physical address allows.
la57 flag), so the last user address is 0x00007fffffffffff. Kernel space starts at 0xffff800000000000, and a direct map of all physical RAM lives at 0xffff888000000000, letting the kernel touch any frame.With base 0 forced, the linear address equals the effective address (or FS/GS base plus offset), and paging turns it into physical. In 64-bit mode the walk has four levels instead of two. The 48-bit address splits into four 9-bit table indices and a 12-bit page offset. CR3 points at the top table (the PML4).
Everything the old segment limits and data segments tried to do is now a paging bit, applied per 4 KiB page: isolation (each process has its own CR3), read-only pages (R/W), user vs kernel (U/S), and non-executable data (NX). Huge pages let a level stop early and map 2 MiB or 1 GiB at once.
Reading the selectors is unprivileged. The FS/GS bases need gdb, a syscall, or the FSGSBASE instructions. There is no /proc file for segment registers; the closest sources are /proc/PID/maps for the memory layout and a core dump for a full register snapshot. This C program reads everything and was run on this machine:
// seg64.c — gcc -O0 seg64.c -o seg64 && ./seg64 #include <stdio.h> #include <sys/syscall.h> #include <asm/prctl.h> int main(void){ unsigned short cs,ds,es,ss,fs,gs; __asm__("mov %%cs,%0":"=r"(cs)); __asm__("mov %%ds,%0":"=r"(ds)); __asm__("mov %%es,%0":"=r"(es)); __asm__("mov %%ss,%0":"=r"(ss)); __asm__("mov %%fs,%0":"=r"(fs)); __asm__("mov %%gs,%0":"=r"(gs)); unsigned long fsb=0,gsb=0; syscall(SYS_arch_prctl, ARCH_GET_FS, &fsb); syscall(SYS_arch_prctl, ARCH_GET_GS, &gsb); unsigned long canary; __asm__("mov %%fs:0x28,%0":"=r"(canary)); printf("cs=%#x ds=%#x es=%#x ss=%#x fs=%#x gs=%#x CPL=%d\n", cs,ds,es,ss,fs,gs, cs&3); printf("fs_base=%#lx gs_base=%#lx canary=%#lx\n", fsb, gsb, canary); }
In a disassembler like Binary Ninja or objdump, you see which segment an instruction uses (for example fs:0x28) but not its runtime base. When you spot fs:0x28 at the top of a function, that function has a stack-protector canary. A gs: operand in a Windows x64 binary is a TEB access instead.
| Reg | Value on 64-bit Linux | What it does in 64-bit mode |
|---|---|---|
| CS | 0x33 user, 0x10 kernel | Selects 64-bit vs 32-bit decoding (L bit); low 2 bits = CPL. Changed only by far transfers, int/iret, syscall/sysret |
| DS | 0 | Base forced to 0; default for data operands but adds nothing |
| ES | 0 | Base forced to 0; still the string destination |
| SS | 0x2b user, 0x18 kernel | Base forced to 0; still the default for [rsp]/[rbp] |
| FS | 0 + FS_BASE MSR | Thread-local storage. fs:0x28 = canary, fs:0 = TCB |
| GS | 0 + GS_BASE MSR | Kernel per-CPU data via swapgs (Windows: user TEB) |
arch_prctl or wrfsbase.swapgs.That completes the two-part tour. Part 1 built the segmented model from the 8086 up to the 32-bit flat model; this part showed how x86-64 froze that model into hardware and kept only FS, GS and CS doing real work.