Part 1 of 2 · Part 2: x86-64 →
A tutorial from first principles: why segments exist, what the CPU does when you load one, and why modern 32-bit systems turned them off and used paging instead.
In 1978 Intel's 8086 had 16-bit registers, which can hold 216 = 65,536 different values, so a register alone can name 64 KiB of memory. But the chip had 20 address lines, enough for 220 bytes = 1 MiB. Intel needed a way to form a 20-bit address from 16-bit parts.
The answer was a second register that says which 64 KiB window you are looking through. That register is a segment register. The ordinary register (the offset) says where you are inside the window.
That design outlived the 8086. Every x86 CPU, including the one you are using now, still computes addresses this way. What changed over the years is what the segment part means, and that is what the rest of this tutorial is about.
When an x86 CPU powers on, it starts in real mode, which is 8086-compatible. Here the rule is simple arithmetic: shift the segment left by 4 bits (multiply by 16) and add the offset.
FFFF:0010.Two consequences of this formula matter later:
Segments overlap. Segment 0x1000 starts at 0x10000, and segment 0x1001 starts only 16 bytes later. The two windows share almost all their memory. The same physical byte can be named by up to 4096 different segment:offset pairs.
There is no protection. Any program can load any value into a segment register and read or write anywhere in the 1 MiB. Keeping data areas separate was only a programming convention.
FFFF:0010 = 0x100000, one byte past 1 MiB. The 8086 had only 20 address lines, so this wrapped around to 0, and some old software depended on it. Later PCs added the "A20 gate" to switch the wrap on or off. With it off, real-mode code could reach almost 64 KiB above 1 MiB (the "HMA").The 8086 had four segment registers: CS, DS, SS and ES. The 386 added FS and GS. Instructions rarely name a segment register. The CPU picks one automatically based on what kind of access is happening:
| Register | Used automatically for |
|---|---|
| CS code | Every instruction fetch: the CPU reads the next instruction from CS:EIP. |
| SS stack | push, pop, call, ret, and any memory operand whose base register is ESP or EBP. |
| DS data | All other data accesses, including the source of string instructions (DS:ESI). |
| ES extra | The destination of string instructions (ES:EDI). This one cannot be overridden. |
| FS | Nothing by default. Used only when an instruction asks for it with a prefix. |
| GS | Same as FS. |
You can override the default with a segment prefix byte in front of the instruction: 26=ES, 2E=CS, 36=SS, 3E=DS, 64=FS, 65=GS. In assembly you write it as mov eax, fs:[0x18].
Real mode's lack of protection became a problem once multitasking operating systems arrived. The 286 and then the 386 introduced protected mode. In protected mode, the 16-bit value in a segment register is no longer an address. It is a selector, meaning an index into a table of segment definitions kept by the operating system.
The selector's three parts:
Index (13 bits): which entry of the table to use, 0 to 8191. Each entry is 8 bytes, so the entry sits at table_base + index × 8. Conveniently, that is just the selector with its low 3 bits cleared.
TI (table indicator): 0 selects the Global Descriptor Table (GDT), shared by everyone. 1 selects the current Local Descriptor Table (LDT), which can be per-process. Linux almost never uses LDTs.
RPL (requested privilege level): a privilege level from 0 (most privileged, the kernel) to 3 (least, user programs). In CS, these two bits are the current privilege level (CPL) of the running code. That is how the CPU knows whether it is running kernel code or user code.
Each table entry is an 8-byte segment descriptor. It defines one segment with three things: where it starts (base, 32 bits), how big it is (limit, 20 bits), and what it is allowed to be used for (access rights). The CPU finds the GDT through the GDTR register, which holds the table's address and size and is loaded by the kernel with the privileged lgdt instruction.
The format looks strange because the base and limit are split into pieces. That is a leftover of the 286, whose descriptors had only a 24-bit base and 16-bit limit, and the 386 fitted the extra bits into what was left. Build a descriptor below and watch the fields land in their bits:
0xFFFFF means 4 GiB. S = 1 marks a code/data segment (0 would be a system descriptor such as a TSS or gate).The access-rights fields:
Type says whether this is code (executable) or data, and whether it is readable or writable. You can never write to a code segment, and never execute from a data segment.
DPL (descriptor privilege level) is the privilege needed to use the segment. We check it in chapter 6.
P (present): if 0, using the segment raises a "segment not present" fault. Old operating systems used this to swap whole segments to disk.
D/B selects 32-bit (1) or 16-bit (0). For a code segment it decides how the CPU decodes instructions. This matters again in 64-bit mode.
Consider mov ds, ax. In real mode it just copies a number. In protected mode it sets off a sequence of lookups and checks, and if one fails the CPU raises a fault instead of loading:
It checks that the selector is inside the table (index × 8 + 7 ≤ GDTR.limit), reads the descriptor, checks the type (DS must be data or readable code), checks privilege, checks the present bit, and finally copies the descriptor's base, limit and rights into a hidden part of the register.
0x2B (a small data segment), 0x10 (kernel data from user mode), 0x3B (not present), 0x48 (outside the table), 0x18 (code segment into SS), and 0x00 (null). The error code of the fault is the offending selector.The privilege rule for data segments is that the less privileged of CPL and RPL must still be allowed by the descriptor: max(CPL, RPL) ≤ DPL. User code (CPL 3) therefore cannot load a kernel data segment (DPL 0). SS is stricter: RPL = CPL = DPL, because the stack must always belong to the current privilege level.
The null selector (0) is special. It can be loaded into DS, ES, FS or GS, which means "no segment", but any access through it faults. It cannot be loaded into SS or CS.
0xF000 but a cached base of 0xFFFF0000, so the first instruction is fetched from 0xFFFFFFF0, just below 4 GiB, where the firmware ROM is.Once a register is loaded, each memory access through it computes the linear address and checks the limit:
linear = segment.base + offset ; offset = what the instruction computed, e.g. [ebx+8] if offset + size − 1 > segment.limit → #GP(0) ; #SS(0) if the segment is SS if writing and segment is read-only → #GP(0) if segment register holds null → #GP(0)
The segments you loaded in Figure 6 are still there. Use them now:
0x2B into DS in Figure 6 (base 4 MiB, limit 64 KiB) and drag the slider past the end. Load 0x33 into ES (read-only) and try a write.This is the "separate data segments" model the old textbooks describe: each segment is a protected box. A buffer overflow past a segment's limit faults immediately instead of corrupting a neighbor. In practice it was slow and awkward, and chapter 9 explains why everyone abandoned it.
CS is loaded differently from the others. There is no mov cs, ax: the encoding exists but always raises an "invalid opcode" (#UD) exception. The reason is that changing CS changes where the next instruction comes from, so CS and EIP must change together in one atomic step. Only control-transfer instructions do that:
| Instruction | What it does with CS |
|---|---|
jmp far sel:off, call far | Load a new CS (same privilege level only, unless going through a call gate). |
retf | Pop EIP and CS (can return to lower privilege). |
int n / hardware interrupts / exceptions | CS comes from the IDT gate. This is how user code enters the kernel. |
iret | Pop EIP, CS, EFLAGS (and ESP, SS if returning to a lower privilege). |
sysenter / sysexit | Fast system call: CS is taken from an MSR (Linux i386 uses it through the vDSO). |
lgdt, mov cr3, hlt, …) are allowed.So "no program can change CS" is only half true. A program can far-jump to any code segment at its own privilege level. What it cannot do is raise its privilege. That only happens through gates the kernel set up (interrupts, call gates, sysenter), and those always enter at a kernel-chosen address.
Segmented protected mode works, but it hurts in practice. A C pointer would need to carry a segment and an offset (a "far pointer"), segment loads are slow because of all the checks, and compilers struggle with memory that lives in different segments.
Every mainstream 32-bit OS (Linux, Windows NT, macOS, the BSDs) made the same choice: make every segment start at 0 and cover all 4 GiB. This is the flat memory model. With base 0, the linear address equals the offset, so a pointer is just a 32-bit number and segmentation effectively disappears.
[0x0804A010] means the same byte whichever segment is used.32-bit Linux defines just four flat segments in its GDT, plus a few special ones:
| Selector | GDT entry | Descriptor | Used as |
|---|---|---|---|
| 0x60 | 12 | base 0, 4 GiB, code, DPL 0 | kernel CS |
| 0x68 | 13 | base 0, 4 GiB, data, DPL 0 | kernel DS/ES/SS |
| 0x73 | 14 | base 0, 4 GiB, code, DPL 3 | user CS |
| 0x7b | 15 | base 0, 4 GiB, data, DPL 3 | user DS/ES/SS |
| 0x33 | 6 | base = this thread's TLS | user GS (chapter 12) |
| 0xd8 | 27 | base = this CPU's data | kernel FS (per-CPU variables) |
With flat segments, segmentation no longer protects anything between processes: every process uses the same selectors 0x73 and 0x7b with the same base 0. The only protection segmentation still provides is privilege: user code runs at CPL 3 and cannot execute kernel instructions. Keeping processes apart moved to a different mechanism, paging.
.text and .data end up is decided by the ELF program headers and the page tables. The kernel gives every process the same selector values no matter what is in the executable.The 386 introduced a second translation step after segmentation. It splits the linear address space into 4 KiB pages and uses page tables to map each page to a physical frame of RAM, or to nothing at all. The complete path of an address is:
In classic 32-bit paging, the linear address is cut into three fields. The top 10 bits select an entry in the page directory, the next 10 bits select an entry in a page table, and the last 12 bits are the offset inside the 4 KiB page. The physical address of the page directory is in control register CR3.
Everything that segmentation tried to do is done better by paging, one page at a time:
Isolation. Each process has its own page directory. The kernel loads a new CR3 on every process switch, so the same linear address maps to different RAM in different processes.
Protection. The U/S bit keeps user code out of kernel pages, and R/W marks read-only pages. With PAE there is also an NX bit that marks pages non-executable.
Demand loading and sharing. A page that is not present causes a page fault. The kernel then loads the page from disk and retries the instruction. Shared libraries are mapped into many processes but use the same physical frames.
A 32-bit register can hold 232 = 4,294,967,296 different values, so a 32-bit process can name at most 4 GiB of linear addresses. On 32-bit Linux that space is split into 3 GiB for the user (0x00000000–0xBFFFFFFF) and 1 GiB for the kernel (0xC0000000–0xFFFFFFFF). The kernel part is mapped in every process but protected by the U/S bit.
Those 4 GiB are virtual. Each process has its own page tables, so each gets its own private 4 GiB, and the frames behind them can be anywhere in RAM. Classic page-table entries are 32 bits wide and hold a 20-bit frame number, so they can only point to frames below 4 GiB. PAE (Physical Address Extension) widens the entries to 64 bits, which allows 36-bit physical addresses and up to 64 GiB of RAM. Each process still sees only 4 GiB, but together the processes can fill much more RAM than that.
Flat mode makes CS, DS, ES and SS boring. FS and GS have no default use, so an OS is free to give them a non-zero base as a pointer to something that differs per thread or per CPU. That gives a fast way to reach "my own data" in one instruction, without passing a pointer around.
32-bit Linux gives every thread its own descriptor in GDT slot 6 (selector 0x33), whose base is the address of that thread's thread control block. User code keeps GS = 0x33 forever. On every thread switch the kernel rewrites that GDT slot and reloads GS, so gs:[0] always refers to the current thread:
mov eax, gs:[0x14] ; the stack-protector canary (every protected function starts this way) mov eax, gs:[0x0] ; pointer to this thread's own control block → pthread_self() mov eax, gs:[-8] ; a __thread variable (thread-local storage sits below the TCB)
The kernel does the same for itself with FS (selector 0xd8, a per-CPU base). Windows 32-bit uses FS in user mode: fs:[0x18] is the Thread Environment Block and fs:[0x30] the Process Environment Block.
set_thread_area() to fill their TLS slot. 32-bit Linux offers three such slots (GDT 6–8, selectors 0x33, 0x3b, 0x43). Programs like Wine that need their own custom segments can create LDT entries with modify_ldt(); those selectors have TI = 1, for example 0x0f.Reading a segment register is not privileged. mov ax, ds works in user mode. What user code cannot do is change the GDT, the IDT or the page tables. Compile a 32-bit program (on Debian: apt install gcc-multilib):
// seg32.c — gcc -m32 seg32.c -o seg32 && ./seg32 #include <stdio.h> int main(void) { unsigned short cs, ds, es, ss, fs, gs; __asm__("mov %%cs,%0" : "=r"(cs)); __asm__("mov %%ds,%0" : "=r"(ds)); __asm__("mov %%es,%0" : "=r"(es)); __asm__("mov %%ss,%0" : "=r"(ss)); __asm__("mov %%fs,%0" : "=r"(fs)); __asm__("mov %%gs,%0" : "=r"(gs)); unsigned canary; __asm__("mov %%gs:0x14,%0" : "=r"(canary)); printf("cs=%#x ds=%#x es=%#x ss=%#x fs=%#x gs=%#x CPL=%d canary=%#x\n", cs, ds, es, ss, fs, gs, cs & 3, canary); }
The values depend on which kernel runs the program:
| CS | DS / ES / SS | FS | GS | |
|---|---|---|---|---|
| native 32-bit Linux kernel | 0x73 | 0x7b | 0 | 0x33 |
| 32-bit program on a 64-bit kernel | 0x23 | 0x2b | 0 | 0x63 |
In gdb, info registers lists all six selectors. gdb cannot show the hidden cache directly, but for CS/DS/ES/SS you already know it: base 0, limit 4 GiB. Disassemblers (objdump, Binary Ninja, IDA) show which segment an instruction uses, such as gs:0x14, but not its runtime value. There is no /proc file for segment registers. The closest things are /proc/PID/maps, which shows the memory layout, and core dumps, which store all registers.
(Typical output for a 32-bit program on a 64-bit kernel. Addresses vary.)
| Reg | Real mode | 32-bit protected, flat (Linux) |
|---|---|---|
| CS | code window = CS×16 | flat code, base 0. Its low 2 bits = CPL. Changed only by far jmp/call/ret, int, iret, sysenter/sysexit |
| DS | data window | flat data, base 0 |
| ES | string destination window | flat data, base 0 (still the string destination) |
| SS | stack window | flat data, base 0 |
| FS | spare window | 0 in user mode · kernel per-CPU data · Windows TEB |
| GS | spare window | thread-local storage, base = TCB; canary at gs:0x14 |
In Part 2: x86-64, the flat model stops being a convention and becomes the hardware rule, FS and GS get 64-bit bases stored outside the GDT, and paging grows to four levels.