Part 1 of 2 · Part 2: x86-64 →

Segment Registers in 32-bit x86

A tutorial from first principles: why segments exist, what the CPU does when you load one, and why modern 32-bit systems turned them off and used paging instead.

1The problem segments solved

In 1978 Intel's 8086 had 16-bit registers, which can hold 216 = 65,536 different values, so a register alone can name 64 KiB of memory. But the chip had 20 address lines, enough for 220 bytes = 1 MiB. Intel needed a way to form a 20-bit address from 16-bit parts.

The answer was a second register that says which 64 KiB window you are looking through. That register is a segment register. The ordinary register (the offset) says where you are inside the window.

A memory address on x86 always has two parts: a segment and an offset. The instruction usually supplies only the offset, and the segment is picked implicitly. This pair is called a logical address.

That design outlived the 8086. Every x86 CPU, including the one you are using now, still computes addresses this way. What changed over the years is what the segment part means, and that is what the rest of this tutorial is about.

2Real mode: segment × 16 + offset

When an x86 CPU powers on, it starts in real mode, which is 8086-compatible. Here the rule is simple arithmetic: shift the segment left by 4 bits (multiply by 16) and add the offset.

segment : offset
Figure 1. Real-mode address translation. The segment slides one hex digit to the left (×16), a zero fills in, and the offset is added. Try FFFF:0010.

Two consequences of this formula matter later:

Segments overlap. Segment 0x1000 starts at 0x10000, and segment 0x1001 starts only 16 bytes later. The two windows share almost all their memory. The same physical byte can be named by up to 4096 different segment:offset pairs.

There is no protection. Any program can load any value into a segment register and read or write anywhere in the 1 MiB. Keeping data areas separate was only a programming convention.

Probe physical address → how each register would have to name it
Figure 2. Each segment register opens a 64 KiB window into the 1 MiB address space. Drag the sliders and watch the windows move and overlap. The white line is the probed byte.
FFFF:0010 = 0x100000, one byte past 1 MiB. The 8086 had only 20 address lines, so this wrapped around to 0, and some old software depended on it. Later PCs added the "A20 gate" to switch the wrap on or off. With it off, real-mode code could reach almost 64 KiB above 1 MiB (the "HMA").

3The six registers and their jobs

The 8086 had four segment registers: CS, DS, SS and ES. The 386 added FS and GS. Instructions rarely name a segment register. The CPU picks one automatically based on what kind of access is happening:

RegisterUsed automatically for
CS codeEvery instruction fetch: the CPU reads the next instruction from CS:EIP.
SS stackpush, pop, call, ret, and any memory operand whose base register is ESP or EBP.
DS dataAll other data accesses, including the source of string instructions (DS:ESI).
ES extraThe destination of string instructions (ES:EDI). This one cannot be overridden.
FSNothing by default. Used only when an instruction asks for it with a prefix.
GSSame as FS.

You can override the default with a segment prefix byte in front of the instruction: 26=ES, 2E=CS, 36=SS, 3E=DS, 64=FS, 65=GS. In assembly you write it as mov eax, fs:[0x18].

Pick an instruction.
Figure 3. Which segment register does each instruction use? Click an instruction to highlight the segment(s) it uses.

4Protected mode: the selector

Real mode's lack of protection became a problem once multitasking operating systems arrived. The 286 and then the 386 introduced protected mode. In protected mode, the 16-bit value in a segment register is no longer an address. It is a selector, meaning an index into a table of segment definitions kept by the operating system.

Index (bits 15–3): entry number in the tableTI (bit 2): 0 = GDT, 1 = LDTRPL (bits 1–0): requested privilege level
Selector
Figure 4. Anatomy of a selector. Click individual bits to flip them, or type a value. The presets are real selectors used by 32-bit Linux.

The selector's three parts:

Index (13 bits): which entry of the table to use, 0 to 8191. Each entry is 8 bytes, so the entry sits at table_base + index × 8. Conveniently, that is just the selector with its low 3 bits cleared.

TI (table indicator): 0 selects the Global Descriptor Table (GDT), shared by everyone. 1 selects the current Local Descriptor Table (LDT), which can be per-process. Linux almost never uses LDTs.

RPL (requested privilege level): a privilege level from 0 (most privileged, the kernel) to 3 (least, user programs). In CS, these two bits are the current privilege level (CPL) of the running code. That is how the CPU knows whether it is running kernel code or user code.

5Descriptors and the GDT

Each table entry is an 8-byte segment descriptor. It defines one segment with three things: where it starts (base, 32 bits), how big it is (limit, 20 bits), and what it is allowed to be used for (access rights). The CPU finds the GDT through the GDTR register, which holds the table's address and size and is loaded by the kernel with the privileged lgdt instruction.

The format looks strange because the base and limit are split into pieces. That is a leftover of the 286, whose descriptors had only a 24-bit base and 16-bit limit, and the 386 fitted the extra bits into what was left. Build a descriptor below and watch the fields land in their bits:

base limit
type DPL
presets:
or decode a qword
bits 63 … 32 (upper dword)
bits 31 … 0 (lower dword)
Figure 5. Descriptor builder. The limit is 20 bits. With G = 1 it counts 4 KiB pages instead of bytes, so a limit of 0xFFFFF means 4 GiB. S = 1 marks a code/data segment (0 would be a system descriptor such as a TSS or gate).

The access-rights fields:

Type says whether this is code (executable) or data, and whether it is readable or writable. You can never write to a code segment, and never execute from a data segment.

DPL (descriptor privilege level) is the privilege needed to use the segment. We check it in chapter 6.

P (present): if 0, using the segment raises a "segment not present" fault. Old operating systems used this to swap whole segments to disk.

D/B selects 32-bit (1) or 16-bit (0). For a code segment it decides how the CPU decodes instructions. This matters again in 64-bit mode.

6Loading a segment register

Consider mov ds, ax. In real mode it just copies a number. In protected mode it sets off a sequence of lookups and checks, and if one fails the CPU raises a fault instead of loading:

It checks that the selector is inside the table (index × 8 + 7 ≤ GDTR.limit), reads the descriptor, checks the type (DS must be data or readable code), checks privilege, checks the present bit, and finally copies the descriptor's base, limit and rights into a hidden part of the register.

Every segment register has a visible part (the 16-bit selector) and a hidden part (the descriptor cache: base, limit, access rights). The CPU reads the GDT only at load time. After that, every memory access uses the cached copy, which is why segmentation adds no memory traffic per access.
← selector running at
Checks performed by the CPU
    GDT in memory GDTR = { base 0x00100000, limit 0x3F }
    Segment registers: visible selector │ hidden descriptor cache
    Figure 6. A simulated CPU with a small GDT. Try loading 0x2B (a small data segment), 0x10 (kernel data from user mode), 0x3B (not present), 0x48 (outside the table), 0x18 (code segment into SS), and 0x00 (null). The error code of the fault is the offending selector.

    The privilege rule for data segments is that the less privileged of CPL and RPL must still be allowed by the descriptor: max(CPL, RPL) ≤ DPL. User code (CPL 3) therefore cannot load a kernel data segment (DPL 0). SS is stricter: RPL = CPL = DPL, because the stack must always belong to the current privilege level.

    The null selector (0) is special. It can be loaded into DS, ES, FS or GS, which means "no segment", but any access through it faults. It cannot be loaded into SS or CS.

    Because the cache is filled only at load time, changing the GDT afterwards does not affect registers that are already loaded. This is a documented part of the architecture, not a bug, and it enables tricks like "unreal mode": load a 4 GiB segment in protected mode, switch back to real mode, and the 4 GiB limit stays in the cache. It also explains the reset state: the CPU starts with CS selector 0xF000 but a cached base of 0xFFFF0000, so the first instruction is fetched from 0xFFFFFFF0, just below 4 GiB, where the firmware ROM is.

    7Using a segment: base, limit, faults

    Once a register is loaded, each memory access through it computes the linear address and checks the limit:

    linear = segment.base + offset          ; offset = what the instruction computed, e.g. [ebx+8]
    if offset + size − 1 > segment.limit  → #GP(0)   ; #SS(0) if the segment is SS
    if writing and segment is read-only   → #GP(0)
    if segment register holds null        → #GP(0)

    The segments you loaded in Figure 6 are still there. Use them now:

    through at offset
    Figure 7. The bar is the segment itself, from offset 0 to its limit, with the region past the limit in red. Load 0x2B into DS in Figure 6 (base 4 MiB, limit 64 KiB) and drag the slider past the end. Load 0x33 into ES (read-only) and try a write.

    This is the "separate data segments" model the old textbooks describe: each segment is a protected box. A buffer overflow past a segment's limit faults immediately instead of corrupting a neighbor. In practice it was slow and awkward, and chapter 9 explains why everyone abandoned it.

    8CS and privilege levels

    CS is loaded differently from the others. There is no mov cs, ax: the encoding exists but always raises an "invalid opcode" (#UD) exception. The reason is that changing CS changes where the next instruction comes from, so CS and EIP must change together in one atomic step. Only control-transfer instructions do that:

    InstructionWhat it does with CS
    jmp far sel:off, call farLoad a new CS (same privilege level only, unless going through a call gate).
    retfPop EIP and CS (can return to lower privilege).
    int n / hardware interrupts / exceptionsCS comes from the IDT gate. This is how user code enters the kernel.
    iretPop EIP, CS, EFLAGS (and ESP, SS if returning to a lower privilege).
    sysenter / sysexitFast system call: CS is taken from an MSR (Linux i386 uses it through the vDSO).
    Figure 8. Step through a user program calling the kernel. Watch the CPL (= CS & 3), the CS value, and the stack frame the CPU builds on the kernel stack.
    The current privilege level is not stored in a separate register. It is the low two bits of CS. Kernel code runs with a CS whose descriptor has DPL 0, and user code with DPL 3. The CPU uses the CPL to decide whether privileged instructions (lgdt, mov cr3, hlt, …) are allowed.

    So "no program can change CS" is only half true. A program can far-jump to any code segment at its own privilege level. What it cannot do is raise its privilege. That only happens through gates the kernel set up (interrupts, call gates, sysenter), and those always enter at a kernel-chosen address.

    9The flat model

    Segmented protected mode works, but it hurts in practice. A C pointer would need to carry a segment and an offset (a "far pointer"), segment loads are slow because of all the checks, and compilers struggle with memory that lives in different segments.

    Every mainstream 32-bit OS (Linux, Windows NT, macOS, the BSDs) made the same choice: make every segment start at 0 and cover all 4 GiB. This is the flat memory model. With base 0, the linear address equals the offset, so a pointer is just a 32-bit number and segmentation effectively disappears.

    Figure 9. Four segment registers in the 4 GiB linear address space. In the segmented model they are separate boxes. In the flat model they are all the same box, so [0x0804A010] means the same byte whichever segment is used.

    32-bit Linux defines just four flat segments in its GDT, plus a few special ones:

    SelectorGDT entryDescriptorUsed as
    0x6012base 0, 4 GiB, code, DPL 0kernel CS
    0x6813base 0, 4 GiB, data, DPL 0kernel DS/ES/SS
    0x7314base 0, 4 GiB, code, DPL 3user CS
    0x7b15base 0, 4 GiB, data, DPL 3user DS/ES/SS
    0x336base = this thread's TLSuser GS (chapter 12)
    0xd827base = this CPU's datakernel FS (per-CPU variables)

    With flat segments, segmentation no longer protects anything between processes: every process uses the same selectors 0x73 and 0x7b with the same base 0. The only protection segmentation still provides is privilege: user code runs at CPL 3 and cannot execute kernel instructions. Keeping processes apart moved to a different mechanism, paging.

    Old tutorials say segment registers "point to the start of the .text or .data section". In the flat model they don't: CS and DS both have base 0. Where .text and .data end up is decided by the ELF program headers and the page tables. The kernel gives every process the same selector values no matter what is in the executable.

    10Paging: the real memory manager

    The 386 introduced a second translation step after segmentation. It splits the linear address space into 4 KiB pages and uses page tables to map each page to a physical frame of RAM, or to nothing at all. The complete path of an address is:

    EBX = 0x0804A008
    Figure 10. The two translation stages. In the flat model the segmentation stage adds 0, so all the real work is done by paging.

    In classic 32-bit paging, the linear address is cut into three fields. The top 10 bits select an entry in the page directory, the next 10 bits select an entry in a page table, and the last 12 bits are the offset inside the 4 KiB page. The physical address of the page directory is in control register CR3.

    linear try
    Figure 11. A page-table walk. Each table entry holds the physical address of the next level plus permission bits: P present, R/W writable, U/S user-accessible. The frame addresses are made up but the index math is exact. The CPU caches finished translations in the TLB so it rarely has to walk.

    Everything that segmentation tried to do is done better by paging, one page at a time:

    Isolation. Each process has its own page directory. The kernel loads a new CR3 on every process switch, so the same linear address maps to different RAM in different processes.

    Protection. The U/S bit keeps user code out of kernel pages, and R/W marks read-only pages. With PAE there is also an NX bit that marks pages non-executable.

    Demand loading and sharing. A page that is not present causes a page fault. The kernel then loads the page from disk and retries the instruction. Shared libraries are mapped into many processes but use the same physical frames.

    114 GB, many processes, and PAE

    A 32-bit register can hold 232 = 4,294,967,296 different values, so a 32-bit process can name at most 4 GiB of linear addresses. On 32-bit Linux that space is split into 3 GiB for the user (0x00000000–0xBFFFFFFF) and 1 GiB for the kernel (0xC0000000–0xFFFFFFFF). The kernel part is mapped in every process but protected by the U/S bit.

    Those 4 GiB are virtual. Each process has its own page tables, so each gets its own private 4 GiB, and the frames behind them can be anywhere in RAM. Classic page-table entries are 32 bits wide and hold a 20-bit frame number, so they can only point to frames below 4 GiB. PAE (Physical Address Extension) widens the entries to 64 bits, which allows 36-bit physical addresses and up to 64 GiB of RAM. Each process still sees only 4 GiB, but together the processes can fill much more RAM than that.

    process:
    Figure 12. Three processes, each with its own 4 GiB virtual space, mapped onto 16 GiB of physical RAM. They use the same virtual addresses but different frames. libc is one set of frames shared by everyone, and the kernel is mapped identically in each. Turn PAE off and everything above the 4 GiB line becomes unreachable.
    A common misconception is that "segment registers arrange the 4 GB region within physical memory." They don't: in the flat model they add 0. Placing each process's pages in physical RAM is done entirely by paging (CR3 and the page tables), under the kernel's control.

    12What's left: GS and thread-local storage

    Flat mode makes CS, DS, ES and SS boring. FS and GS have no default use, so an OS is free to give them a non-zero base as a pointer to something that differs per thread or per CPU. That gives a fast way to reach "my own data" in one instruction, without passing a pointer around.

    32-bit Linux gives every thread its own descriptor in GDT slot 6 (selector 0x33), whose base is the address of that thread's thread control block. User code keeps GS = 0x33 forever. On every thread switch the kernel rewrites that GDT slot and reloads GS, so gs:[0] always refers to the current thread:

    mov eax, gs:[0x14]      ; the stack-protector canary (every protected function starts this way)
    mov eax, gs:[0x0]       ; pointer to this thread's own control block → pthread_self()
    mov eax, gs:[-8]        ; a __thread variable (thread-local storage sits below the TCB)

    The kernel does the same for itself with FS (selector 0xd8, a per-CPU base). Windows 32-bit uses FS in user mode: fs:[0x18] is the Thread Environment Block and fs:[0x30] the Process Environment Block.

    Threads call set_thread_area() to fill their TLS slot. 32-bit Linux offers three such slots (GDT 6–8, selectors 0x33, 0x3b, 0x43). Programs like Wine that need their own custom segments can create LDT entries with modify_ldt(); those selectors have TI = 1, for example 0x0f.

    13Seeing it yourself

    Reading a segment register is not privileged. mov ax, ds works in user mode. What user code cannot do is change the GDT, the IDT or the page tables. Compile a 32-bit program (on Debian: apt install gcc-multilib):

    // seg32.c — gcc -m32 seg32.c -o seg32 && ./seg32
    #include <stdio.h>
    int main(void) {
        unsigned short cs, ds, es, ss, fs, gs;
        __asm__("mov %%cs,%0" : "=r"(cs)); __asm__("mov %%ds,%0" : "=r"(ds));
        __asm__("mov %%es,%0" : "=r"(es)); __asm__("mov %%ss,%0" : "=r"(ss));
        __asm__("mov %%fs,%0" : "=r"(fs)); __asm__("mov %%gs,%0" : "=r"(gs));
        unsigned canary; __asm__("mov %%gs:0x14,%0" : "=r"(canary));
        printf("cs=%#x ds=%#x es=%#x ss=%#x fs=%#x gs=%#x  CPL=%d  canary=%#x\n",
               cs, ds, es, ss, fs, gs, cs & 3, canary);
    }

    The values depend on which kernel runs the program:

    CSDS / ES / SSFSGS
    native 32-bit Linux kernel0x730x7b00x33
    32-bit program on a 64-bit kernel0x230x2b00x63

    In gdb, info registers lists all six selectors. gdb cannot show the hidden cache directly, but for CS/DS/ES/SS you already know it: base 0, limit 4 GiB. Disassemblers (objdump, Binary Ninja, IDA) show which segment an instruction uses, such as gs:0x14, but not its runtime value. There is no /proc file for segment registers. The closest things are /proc/PID/maps, which shows the memory layout, and core dumps, which store all registers.

    (gdb) info registers cs ss ds es fs gs cs 0x23 35 ss 0x2b 43 ds 0x2b 43 es 0x2b 43 fs 0x0 0 gs 0x63 99 ← TLS descriptor (GDT 12, RPL 3) (gdb) x/i $pc => 0x565561ad <main+20>: mov eax,gs:0x14 ← canary load

    (Typical output for a 32-bit program on a 64-bit kernel. Addresses vary.)

    14Summary

    RegReal mode32-bit protected, flat (Linux)
    CScode window = CS×16flat code, base 0. Its low 2 bits = CPL. Changed only by far jmp/call/ret, int, iret, sysenter/sysexit
    DSdata windowflat data, base 0
    ESstring destination windowflat data, base 0 (still the string destination)
    SSstack windowflat data, base 0
    FSspare window0 in user mode · kernel per-CPU data · Windows TEB
    GSspare windowthread-local storage, base = TCB; canary at gs:0x14
    1. An x86 address is segment + offset. The segment is usually chosen implicitly by the kind of access.
    2. In real mode, the segment is a number: address = segment × 16 + offset, with no protection.
    3. In protected mode, the segment register holds a selector: an index into the GDT/LDT plus an RPL.
    4. Loading a selector makes the CPU check the descriptor and copy base/limit/rights into the hidden descriptor cache.
    5. Every access computes linear = base + offset and checks the limit and rights.
    6. The low 2 bits of CS are the CPL. Privilege can only be raised through kernel-defined gates.
    7. Real operating systems use the flat model (base 0, 4 GiB) and leave isolation, protection and the 4 GiB-per-process illusion to paging.
    8. Only FS/GS keep a job: pointing at per-thread or per-CPU data.

    In Part 2: x86-64, the flat model stops being a convention and becomes the hardware rule, FS and GS get 64-bit bases stored outside the GDT, and paging grows to four levels.