← Part 1: 32-bit · Part 2 of 2

Segment Registers in x86-64

In 64-bit mode the hardware forces the flat model. This part covers what still matters: CS as the mode-and-privilege switch, FS and GS as per-thread and per-CPU pointers, and the paging that does everything else. Values marked measured were recorded on this Linux 5.10 machine.

1What changed from 32-bit

Part 1 ended with every serious 32-bit OS using the flat model: all segments at base 0 with a 4 GiB limit, paging doing the real work, and only FS or GS pointing at per-thread data. When AMD designed x86-64 in 2003, they turned that convention into hardware rules:

Thing32-bit protected mode64-bit mode
CS / DS / ES / SS basefrom the descriptoralways 0; the descriptor base is ignored
Segment limitschecked on every accessnot checked
FS / GS base32-bit, from the descriptor64-bit, stored in MSRs, independent of the GDT
Null selector in DS/ES/FS/GSloadable, faults when usedloadable and usable (the base is 0 anyway)
What CS still decidescode base/limit, 16/32-bit, CPL64-bit vs 32-bit decoding (L bit) and CPL
Instruction pointerEIPRIP, with RIP-relative addressing
Address size32-bit linear → 32/36-bit physical48-bit linear → up to 52-bit physical
In 64-bit mode, segmentation does two jobs and nothing else. CS selects the execution mode and the privilege level. FS and GS each add a 64-bit base, used as a pointer to "this thread's" or "this CPU's" data. Everything else about memory is paging.

2Getting into long mode

Every x86-64 CPU still powers on in 16-bit real mode, just like in 1978. Firmware walks it up the modes by setting control bits and reloading CS. Long mode is active once paging is on and the EFER.LME bit is set. From then on, each code segment's L bit chooses between two sub-modes: 64-bit mode (L = 1) and compatibility mode (L = 0) for running 32-bit programs.

Figure 1. The boot path from real mode into 64-bit mode. Setting CR0.PE does not switch the CPU by itself; the code keeps running the old way until a far jump reloads CS with a descriptor for the new mode. Watch the ladder on the right.

Two things to notice. First, paging is mandatory in long mode: you cannot turn on 64-bit mode without page tables, which is exactly why segmentation could be simplified. Second, the last step into 64-bit code is a CS reload. The mode lives in CS.

3CS decides how bytes decode

Machine code is just bytes, and the same bytes mean different instructions depending on the current code segment. In 64-bit mode the bytes 0x40–0x4F are REX prefixes, which select 64-bit operands and registers R8–R15. In 32-bit mode those same bytes are inc and dec instructions. On Linux, CS 0x33 has L = 1 and CS 0x23 has L = 0, D = 1.

decode with
byte string:
Figure 2. One byte string, two meanings. Switch CS and watch the bytes regroup into different instructions.

Because CS alone decides the mode, a 64-bit process can run 32-bit code by jumping to a CS with L = 0, and jump back the same way. This is allowed in user mode because both code segments have DPL 3. The program below does that. It was tested on this machine and exits with status 35 = 0x23, the CS value it read while in 32-bit mode measured:

; gate.asm  —  nasm -f elf64 gate.asm && ld gate.o -o gate && ./gate; echo $?   -> 35
bits 64
global _start
_start:
    lea  rax, [rel compat]
    push 0x23               ; new CS: 32-bit user code (L=0, D=1)
    push rax                ; new RIP
    retfq                    ; far return pops RIP then CS -> compatibility mode
bits 32
compat:
    mov  ebx, cs              ; ebx = 0x23  (we are executing 32-bit code now)
    jmp  0x33:back64          ; far jump back to the 64-bit code segment
bits 64
back64:
    mov  edi, ebx             ; exit status = 0x23
    mov  eax, 60
    syscall
This works here because a static ld binary loads below 4 GiB (at 0x401000), and compatibility mode only forms 32-bit addresses. The same selector-swap trick is used by 32-bit Windows to reach 64-bit code, where it is nicknamed "Heaven's Gate."

4The address rule: only FS and GS count

An x86-64 memory operand looks like seg:[base + index*scale + disp]. The CPU first computes the effective address from the registers, then adds the segment base to get the linear address. In 64-bit mode the base is 0 for CS, DS, ES and SS even with an explicit prefix, and comes from an MSR for FS and GS:

effective = base + index*scale + displacement        ; or RIP + displacement
linear    = effective + { FS_BASE  if the  fs:  prefix is present
                          GS_BASE  if the  gs:  prefix is present
                          0        otherwise  (cs: ds: es: ss: prefixes add nothing) }
then: the linear address must be canonical, or the CPU raises #GP  (chapter 11)
: [ base + index * + disp ]
FS_BASE GS_BASE
Figure 3. Effective address vs linear address. Change the prefix and watch which base is added. Only fs: and gs: move the result; the others add 0.

This is why the flat model stops being a choice. Even if you deliberately load a data selector whose descriptor has a non-zero base, the CPU ignores it. The only segment bases that exist in 64-bit mode are the two MSR-backed ones.

5Linux's 64-bit GDT

The GDT still exists, and CS/SS selectors still index into it, but the entries are simpler because base and limit no longer matter for code and data. Linux keeps one GDT per CPU. Click a row to see what each selector is for:

Click a row.
Figure 4. The selectors a 64-bit Linux process and kernel actually use. User code is 0x33, user data 0x2b. The odd numeric order is forced by the syscall/sysret instructions, explained in chapter 7.

Only a handful of these are ever loaded into a segment register. In user space you will always see CS = 0x33 and SS = 0x2b, with DS, ES, FS and GS holding 0. The FS and GS bases are set separately (chapter 10), not by loading one of these selectors.

6Who sets CS: execve

A common question: if a program is position-independent, how does it get the right CS? The answer is that CS has nothing to do with where your code lives. Its base is 0 for everyone. The kernel assigns the same selector value to every process, and the CPU loads it on the return to user mode. Position independence is handled by paging and RIP-relative addressing, never by CS.

Figure 5. What happens to CS, SS, CPL and FS_BASE when you launch a program. The ELF file is never consulted for the segment values; the kernel writes fixed constants into the saved register frame, and iretq loads them.
In Linux's start_thread(), the kernel writes regs->cs = 0x33 and regs->ss = 0x2b into the register frame it will restore, and clears DS/ES/FS/GS to 0. The CPU makes them live when it returns to user mode. Every 64-bit process on the machine starts with the identical CS, because CS no longer encodes an address.

There is no mov cs, ... instruction (it raises #UD). CS changes only through far jump/call/return, interrupts, iret, and the syscall/sysret pair. All of those change CS and RIP together, which is the whole point: the mode and the next instruction must switch atomically.

7syscall, sysret and swapgs

Entering the kernel through the old int gate is slow. x86-64 added syscall/sysret, which switch privilege without touching memory: the new CS and SS come straight from the STAR MSR, not from the GDT. This is why Linux's selectors are ordered the way they are.

syscall :  CS = STAR[47:32]           SS = STAR[47:32] + 8      ; kernel 0x10, 0x18
sysret  :  CS = STAR[63:48] + 16      SS = STAR[63:48] + 8      ; user 0x33, 0x2b
           STAR[63:48] = 0x23  ->  SS = 0x2b, CS = 0x33

But a problem appears the moment the kernel starts running: it needs a pointer to its own per-CPU data, and every register including GS still holds the user's values. The answer is swapgs, a single instruction that exchanges GS_BASE with a hidden MSR (KERNEL_GS_BASE). The kernel runs swapgs on entry and again on exit.

Figure 6. A system call from user to kernel and back. Watch CS, CPL, and the two GS bases trade places. The GDT is never read during any of this.
swapgs is why GS, not FS, is the kernel's per-CPU register. Getting the swapgs pairing wrong (for example, on an exception that arrives just after entry) has been the source of real CPU errata and the "SWAPGS" speculative-execution issue. FS was left for user thread-local storage.

8DS, ES, SS: present but idle

You can still read and even load these registers, but it changes nothing about addressing. Reading is unprivileged, loading is checked against the GDT exactly as in 32-bit mode, and then the base is discarded. The following was run on this machine, showing which loads succeed and which fault measured:

← selector
Figure 7. Loading a segment register in user mode. 0x2b (user data) and 0x33 (readable user code) succeed. 0x10 (kernel data, DPL 0) and 0x50 (past the table) fault with SIGSEGV. Null into SS faults; null into DS/ES/FS/GS is fine. Even on success, the base stays 0.

So the tutorial phrase "the four data segments keep data elements from overlapping" no longer applies at all. DS, ES and SS have base 0 and no limit, so they cover the entire address space and overlap completely. Keeping data apart is the job of page permissions.

9FS and thread-local storage

FS earns its keep. glibc points FS_BASE at each thread's thread control block (TCB). A single instruction with the fs: prefix then reaches per-thread data with no pointer chasing. Two uses appear in almost every compiled program:

mov  rax, fs:0x28      ; the stack-protector canary (start of most functions)
mov  rax, fs:0         ; pointer to the TCB itself -> pthread_self()
mov  eax, fs:0xfffffffffffffffc   ; a __thread int, i.e. fs:-4

The magic is that the same instruction bytes read different memory in different threads, because each thread has its own FS_BASE. On every context switch the kernel reloads it. This was measured on this machine with a two-thread program measured: the two threads reported different FS bases but the identical canary value, and a __thread variable sat at fs:-4 in both.

running thread:
execute:
Figure 8. Three threads, one FS register, three bases. Pick a thread, then run an access and watch which TCB it lands in. The canary is shared; the TLS variable and the self-pointer are per-thread.
The canary always ends in a zero byte. That low 00 stops C string functions from copying past it, so a stray strcpy cannot leak or overwrite the whole value. __thread variables live just below the TCB, at negative offsets from FS_BASE.

10Setting the FS and GS bases

Since the base is a 64-bit MSR, not a descriptor, loading a selector cannot set it. There are two ways:

arch_prctl (syscall 158). The classic route. glibc calls arch_prctl(ARCH_SET_FS, tcb) once per thread during startup. This was verified on this machine: after mov fs, 0x2b, the FS base actually dropped to 0, because loading a selector overwrites the MSR-backed base with the descriptor's (zero) base measured.

wrfsbase / wrgsbase. If the CPU supports FSGSBASE and the kernel enabled it (CR4.FSGSBASE), user code can write its own FS base in one unprivileged instruction, no syscall. This CPU and kernel 5.10 support it, reported through AT_HWCAP2 = 0x2 measured.

// set FS base from user space, two ways
#include <asm/prctl.h>
#include <sys/syscall.h>
syscall(SYS_arch_prctl, ARCH_SET_FS, (unsigned long)my_tcb);   // always works

unsigned long b;
__asm__("rdfsbase %0" : "=r"(b));          // read  (needs FSGSBASE)
__asm__("wrfsbase %0" :: "r"(new_base));      // write (needs FSGSBASE)
Do not overwrite FS_BASE while running under glibc. Every TLS access, the canary check, and errno all read through it. Point it somewhere wrong and the next protected function will fail its canary check and abort.

11The 64-bit address space

A 64-bit register can name 264 addresses, an absurd 16 EiB. Current CPUs only wire up 48 bits (256 TiB), or 57 with the newer 5-level paging. To keep the unused top bits from being misused, x86-64 requires addresses to be canonical: bits 63 down to 47 must all be copies of bit 47. That splits the usable space into a low half and a high half with a vast non-canonical hole between them.

probe address
Figure 9. The x86-64 virtual layout on Linux (not to scale; the hole is astronomically larger than the used parts). Click a region or type an address. A non-canonical address faults with #GP before any page lookup, which is why the fault reports no address.

Linux gives user space the low half (up to 0x00007fffffffffff, 128 TiB) and the kernel the high half. Because the CPU rejects non-canonical addresses immediately, the classic 32-bit "3 GiB user / 1 GiB kernel" squeeze is gone. There is no 4 GiB wall and no PAE-style trickery: physical frames can sit anywhere the 52-bit physical address allows.

This machine has 48-bit VAs (no la57 flag), so the last user address is 0x00007fffffffffff. Kernel space starts at 0xffff800000000000, and a direct map of all physical RAM lives at 0xffff888000000000, letting the kernel touch any frame.

12Four-level paging

With base 0 forced, the linear address equals the effective address (or FS/GS base plus offset), and paging turns it into physical. In 64-bit mode the walk has four levels instead of two. The 48-bit address splits into four 9-bit table indices and a 12-bit page offset. CR3 points at the top table (the PML4).

address try
Figure 10. A four-level page walk. Each entry holds the physical address of the next table plus permission bits (P present, R/W writable, U/S user, NX no-execute). The frame numbers are illustrative; the index math is exact. The CPU caches results in the TLB.

Everything the old segment limits and data segments tried to do is now a paging bit, applied per 4 KiB page: isolation (each process has its own CR3), read-only pages (R/W), user vs kernel (U/S), and non-executable data (NX). Huge pages let a level stop early and map 2 MiB or 1 GiB at once.

13Seeing it yourself

Reading the selectors is unprivileged. The FS/GS bases need gdb, a syscall, or the FSGSBASE instructions. There is no /proc file for segment registers; the closest sources are /proc/PID/maps for the memory layout and a core dump for a full register snapshot. This C program reads everything and was run on this machine:

// seg64.c  —  gcc -O0 seg64.c -o seg64 && ./seg64
#include <stdio.h>
#include <sys/syscall.h>
#include <asm/prctl.h>
int main(void){
    unsigned short cs,ds,es,ss,fs,gs;
    __asm__("mov %%cs,%0":"=r"(cs)); __asm__("mov %%ds,%0":"=r"(ds));
    __asm__("mov %%es,%0":"=r"(es)); __asm__("mov %%ss,%0":"=r"(ss));
    __asm__("mov %%fs,%0":"=r"(fs)); __asm__("mov %%gs,%0":"=r"(gs));
    unsigned long fsb=0,gsb=0;
    syscall(SYS_arch_prctl, ARCH_GET_FS, &fsb);
    syscall(SYS_arch_prctl, ARCH_GET_GS, &gsb);
    unsigned long canary; __asm__("mov %%fs:0x28,%0":"=r"(canary));
    printf("cs=%#x ds=%#x es=%#x ss=%#x fs=%#x gs=%#x  CPL=%d\n", cs,ds,es,ss,fs,gs, cs&3);
    printf("fs_base=%#lx gs_base=%#lx canary=%#lx\n", fsb, gsb, canary);
}
# measured on this machine: cs=0x33 ds=0 es=0 ss=0x2b fs=0 gs=0 CPL=3 fs_base=0x7f407b3b8540 gs_base=0 canary=0xf23424f90fbf7300 # in gdb the bases have their own convenience variables: (gdb) info registers cs ss fs gs cs 0x33 ss 0x2b fs 0x0 gs 0x0 (gdb) p/x $fs_base $1 = 0x7ffff7fab540 # TLS pointer (gdb) x/gx $fs_base+0x28 0x7ffff7fab568: 0x98ef6660276d9b00 # the canary

In a disassembler like Binary Ninja or objdump, you see which segment an instruction uses (for example fs:0x28) but not its runtime base. When you spot fs:0x28 at the top of a function, that function has a stack-protector canary. A gs: operand in a Windows x64 binary is a TEB access instead.

14Summary

RegValue on 64-bit LinuxWhat it does in 64-bit mode
CS0x33 user, 0x10 kernelSelects 64-bit vs 32-bit decoding (L bit); low 2 bits = CPL. Changed only by far transfers, int/iret, syscall/sysret
DS0Base forced to 0; default for data operands but adds nothing
ES0Base forced to 0; still the string destination
SS0x2b user, 0x18 kernelBase forced to 0; still the default for [rsp]/[rbp]
FS0 + FS_BASE MSRThread-local storage. fs:0x28 = canary, fs:0 = TCB
GS0 + GS_BASE MSRKernel per-CPU data via swapgs (Windows: user TEB)
  1. 64-bit mode forces the flat model: CS/DS/ES/SS bases are 0 and limits are not checked.
  2. CS selects the execution mode (L bit) and holds the CPL in its low 2 bits. It is assigned by the kernel and identical for every process.
  3. FS and GS are the only real segment bases, stored in 64-bit MSRs, set by arch_prctl or wrfsbase.
  4. FS = user thread-local storage; GS = kernel per-CPU data, switched with swapgs.
  5. syscall/sysret get their selectors from the STAR MSR, never touching the GDT.
  6. Addresses must be canonical; the space splits into a user half, a kernel half, and a non-canonical hole.
  7. Isolation, protection and per-process address spaces come entirely from 4-level paging.

That completes the two-part tour. Part 1 built the segmented model from the 8086 up to the 32-bit flat model; this part showed how x86-64 froze that model into hardware and kept only FS, GS and CS doing real work.