x86-64 Segment Registers, interactive

The tutorial text you pasted describes 32-bit protected mode, and a few of its sentences mix in 16-bit real mode. On x86-64 Linux (the machine you are using), segment registers have been reduced to a few leftover jobs. This page shows all three eras. Everything marked measured was read on your machine (Linux 5.10, CPU has fsgsbase).

One-sentence summary: in 64-bit mode CS only carries mode and privilege (64-bit code, ring 3). DS, ES and SS have their bases forced to 0 and are ignored. FS and GS are the only ones still used as real "segment bases": FS points to thread-local storage and GS to per-CPU data in the kernel. Memory isolation is done by paging, not by segments.

⚠ Read first: where the tutorial text is outdated or wrong

The tutorial is not bad, but it describes an older model. If you keep these corrections in mind, everything else on this page will make sense.

The text says…What is actually true
"CS stores the base location of the code section (.text), which is used for data access"CS is used for instruction fetch, not data. In 64-bit mode its base is always 0. What still matters in CS is the L bit (64-bit code or 32-bit compat code) and its low 2 bits, which are your current privilege level (CPL).
"Each segment register … contains the pointer to the start of the segment"That is only true in real mode, where the register holds a paragraph number and the base is value × 16. In protected and long mode the register holds a selector, which is an index into the GDT/LDT. The actual base and limit sit in a hidden descriptor cache (see anatomy).
"No program can explicitly load or change the CS register"There is no mov cs, ax (it raises #UD) and no pop cs. But CS is loaded by far jmp, far call, far ret, iret, int, syscall and sysret. A user program can even far-jump to selector 0x23 and run 32-bit code inside a 64-bit process (the "Heaven's Gate" trick).
"Segment registers can neither be read nor changed directly"Reading them is unprivileged: mov ax, ds works in ring 3 (we ran it on your machine). You can also load any valid ring-3 selector into DS, ES, FS, GS or SS. The part that is protected is the GDT contents and the MSRs, not the registers.
"The four data segments help the program separate data elements so they do not overlap"In the flat model every segment has base 0 and covers the whole address space, so they all overlap completely. Even in real mode, segments overlap freely (you will see this in the simulator). Separation comes from paging: page tables plus the R/W/NX and U/S bits.
"The OS must arrange a 4 GB region … this task is completed by the segment registers"This is done by paging (CR3 → page tables). Every process gets its own 4 GB virtual space, and its pages can live anywhere in physical RAM. With PAE a 32-bit OS can use up to 64 GB of RAM (see Q5).
"EIP"On x86-64 it is RIP (64-bit). RIP-relative addressing is the reason PIE code works without segments.
(your Q0) "there are 15 general registers"There are 16: RAX, RBX, RCX, RDX, RSI, RDI, RBP, RSP and R8–R15. RSP counts as general-purpose even though the stack uses it. Intel APX adds R16–R31.
"SS is used when implicitly using the stack pointer or explicitly using the base pointer"Correct: any memory operand whose base register is RSP or RBP defaults to SS. That no longer matters in 64-bit mode because SS.base = 0.

Q0Where do segment registers fit in the big picture?

Yes, they are just one category among many. Click any register chip for details. The filters highlight what your program can touch versus what only the kernel (ring 0) can touch.

Click a register

Details appear here.

GPR anatomy: how one 64-bit register splits into smaller names

Writing a 32-bit sub-register (e.g. mov eax, 1) zeroes the upper 32 bits of RAX. Writing 16-bit or 8-bit parts (ax, al, ah) leaves the other bits alone. That is why compilers write xor eax, eax to clear RAX: it is shorter and equivalent.
AH/BH/CH/DH cannot be encoded in an instruction that has a REX prefix, so mov ah, sil or mov ah, r8b is impossible. SIL, DIL, BPL and SPL only exist because of REX.

Anatomy of a segment register (the part the text leaves out)

Each segment register is 16 bits that you can see, plus a larger hidden descriptor cache that the CPU fills whenever you load it.

Visible 16 bits
Selector
Hidden descriptor cache (filled by CPU from GDT/LDT on load)
Base (32/64 bit) · Limit (20 bit + G) · Access rights (type, DPL, P, L, D/B)
Selector bit layout
■ Index (13 bits) → entry number in the table (×8 = byte offset)
■ TI → 0 = GDT, 1 = LDT
■ RPL → requested privilege level (for CS this is the CPL: 3 = user, 0 = kernel)
Decode a selector:
Linux x86-64 GDT (per CPU) from arch/x86/include/asm/segment.h
The order USER32_CS (0x23), USER_DS (0x2b), USER_CS (0x33) looks strange but is forced by the hardware. sysret computes SS = STAR[63:48] + 8 and CS = STAR[63:48] + 16. With STAR[63:48] = 0x23 you get SS = 0x2b and CS = 0x33.
Selector 0x7b (GDT index 15, "CPUNODE") holds nothing useful as memory. Its limit field encodes (node << 12) | cpu. The vDSO's getcpu() reads it with the unprivileged lsl instruction. On your machine lsl returned 0xa, meaning CPU 10, NUMA node 0 measured.

Q1Does DS get its value from the ELF .data section?

No. DS does not know your ELF file exists. The kernel sets DS = 0 (null selector) for every 64-bit process, and in 64-bit mode its base is 0 no matter what it holds. What places .data in memory is the ELF program header (PT_LOAD, with its p_vaddr), which the kernel mmaps. Your code then reaches it with an absolute address (as in your lab/01_hello.asm: mov rax, [myvar]) or RIP-relative (PIE: mov rax, [rip+0x2ec0]). DS base 0 plus that address gives the linear address, and paging turns it into a physical one.
The word "segment" has two unrelated meanings here. An ELF segment is a program header (PT_LOAD), a chunk of the file that gets mmapped. An x86 segment is a CS/DS/… descriptor. readelf -l's "Section to Segment mapping" is about ELF, not about DS.

ELF file on disk

kernel
load_elf_binary()
→ mmap()
per PT_LOAD

Process virtual address space

Press the button to load the program.

Try it on your own binaries

# ELF sections (linker view)  vs  ELF segments (loader view)
readelf -SW ./a.out
readelf -lW ./a.out            # PT_LOAD = what the kernel maps

# Where did .data really land? (run while the process is alive)
cat /proc/$(pidof a.out)/maps
#  555555558000-555555559000 rw-p 00003000 ... a.out   ← .data/.bss

# And the CPU side: DS is just 0
gdb -q ./a.out -ex starti -ex 'p $ds'   # $1 = 0

Your nasm lab (non-PIE, static)

section .data
    myvar   dq  0x1122334455667788   ; → ELF PT_LOAD RW at e.g. 0x402000
section .text
_start:
    mov rax, [myvar]  ; = mov rax, ds:[0x402000]
                      ;   DS.base = 0 → linear 0x402000
    mov rbx, myvar    ; just the number 0x402000
In nasm, add default rel (or write [rel myvar]) to get RIP-relative addressing. It is 1 byte shorter than a 32-bit absolute address, and it is required for PIE/shared libraries. Check the encoding in your .lst file.

Q2Who assigns CS? And what if the program is position independent?

The kernel picks the selector, and the CPU loads it. In execve(), Linux's start_thread() writes cs = __USER_CS (0x33) and ss = __USER_DS (0x2b) into the saved register frame. When the kernel returns to user mode (iretq/sysretq), the CPU loads those into the real CS and SS. Every 64-bit process on the system gets the same CS value, 0x33, because CS no longer says where your code is. It only says "64-bit, ring 3".

PIE has nothing to do with CS. Position independence is handled by (1) ASLR, which chooses a random load base for the mmapped PT_LOADs, (2) RIP-relative addressing, where code refers to data as "RIP + constant", and (3) the dynamic loader patching the GOT with relocations. CS.base stays 0 throughout.
32-bit programs on the same 64-bit kernel run with CS = 0x23 (a descriptor with L=0, D=1), so the CPU decodes their bytes as 32-bit code. That is the only difference between a 32-bit and a 64-bit process as far as the CPU is concerned. The same bytes decode differently: 48 is a REX prefix in 64-bit mode but dec eax in 32-bit mode.

Q3Segment registers in action: loading, using, "unloading"

The paragraph you quoted ("the program loads the data segment registers … then references memory using an offset") describes real mode and non-flat protected mode programs. Step through all three eras and watch what happens to the address math. Click any instruction to jump to it.

Formula: physical = segment × 16 + offset. Each register opens a 64 KiB "window" into 1 MiB. The colored bands below are those windows. The markers show every memory access: ■ fetch (CS:IP) · ■ read · ■ write.

Formula: loading a selector makes the CPU index the GDT, check the descriptor and privilege, and copy base/limit into the hidden cache. Every access then computes linear = cache.base + offset after checking offset ≤ limit. The bar shows the first 16 MiB of the 4 GiB linear space.

Rule in 64-bit mode: CS/DS/ES/SS bases are treated as 0 and limits are not checked. FS and GS keep a real 64-bit base, stored in the MSRs IA32_FS_BASE and IA32_GS_BASE, not in the GDT. "Loading and unloading" now happens per thread, at every context switch.

segment : offset
FFFF:0010 = 0x100000, which is 1 byte past 1 MiB. On the 8086 this wrapped around to 0x00000, and some DOS programs relied on it. The PC/AT therefore added the A20 gate (famously routed through the keyboard controller) to emulate the wrap. With A20 on, FFFF:0010–FFFF:FFFF reach the 65,520-byte "HMA" above 1 MiB, which DOS=HIGH used.

Q4How can I see and read them? gdb, Binary Ninja, /proc

Short answer: the selectors can be read by anyone (mov ax, ds, gdb info registers, ptrace). The FS/GS bases can be read with gdb ($fs_base, $gs_base), arch_prctl(ARCH_GET_FS), or the rdfsbase instruction (enabled on your CPU and kernel). There is no /proc file that shows segment registers. /proc shows the memory map (maps) and a few other registers (syscall). The hidden descriptor cache for CS/DS/SS cannot be dumped from user space, but in 64-bit mode you already know its contents: base 0.
output recorded on your machine with /tmp/segp/p measured
Right after starti, fs_base is 0. At that point you are inside ld.so's _start and TLS does not exist yet. ld.so later calls arch_prctl(ARCH_SET_FS, tcb). You can catch it with catch syscall arch_prctl.
Useful gdb commands: info registers (general and segment), info all-registers (includes SIMD), tui reg system / tui reg all, p/x $fs_base, x/gx $fs_base+0x28 (canary), set $fs_base = … (writes it through ptrace; you will probably crash libc). gdb also has $_tlb on Windows targets and info threads + thread N to compare fs_base between threads.
// seg.c — gcc -O0 seg.c -o seg && ./seg      (run it twice: fs_base changes → ASLR)
#include <stdio.h>
#include <stdint.h>
#include <unistd.h>
#include <sys/syscall.h>
#include <sys/auxv.h>
#include <asm/prctl.h>
#include <pthread.h>
#ifndef HWCAP2_FSGSBASE
#define HWCAP2_FSGSBASE (1 << 1)
#endif

static void dump(const char *who){
    uint16_t cs, ds, es, ss, fs, gs;            // 1) selectors: plain unprivileged MOV
    __asm__ volatile("mov %%cs,%0" : "=r"(cs));  __asm__ volatile("mov %%ds,%0" : "=r"(ds));
    __asm__ volatile("mov %%es,%0" : "=r"(es));  __asm__ volatile("mov %%ss,%0" : "=r"(ss));
    __asm__ volatile("mov %%fs,%0" : "=r"(fs));  __asm__ volatile("mov %%gs,%0" : "=r"(gs));

    unsigned long fsb = 0, gsb = 0;             // 2) bases via syscall (always works)
    syscall(SYS_arch_prctl, ARCH_GET_FS, &fsb);
    syscall(SYS_arch_prctl, ARCH_GET_GS, &gsb);

    unsigned long self, canary;                  // 3) use FS like the compiler does
    __asm__("mov %%fs:0,%0"    : "=r"(self));      // TCB self-pointer (== pthread_self())
    __asm__("mov %%fs:0x28,%0" : "=r"(canary));    // stack-protector canary

    printf("[%s] cs=%#x ds=%#x es=%#x ss=%#x fs=%#x gs=%#x\n", who, cs, ds, es, ss, fs, gs);
    printf("      CPL=%d  fs_base=%#lx  gs_base=%#lx  fs:0=%#lx  canary=%#lx\n",
           cs & 3, fsb, gsb, self, canary);

    if (getauxval(AT_HWCAP2) & HWCAP2_FSGSBASE) {   // 4) rdfsbase: kernel ≥5.9 + CPU support
        unsigned long r;
        __asm__ volatile("rdfsbase %0" : "=r"(r));
        printf("      rdfsbase=%#lx\n", r);
    }
}

static void *thr(void *a){ dump("thread"); return 0; }

int main(void){
    dump("main  ");
    pthread_t t; pthread_create(&t, 0, thr, 0); pthread_join(t, 0);  // different fs_base, same canary!

    unsigned p;                                   // 5) bonus: CPU/node via segment limit
    __asm__ volatile("lsl %1,%0" : "=r"(p) : "r"(0x7bu));
    printf("running on cpu %u node %u\n", p & 0xfff, p >> 12);
}
# measured output on your machine (single-thread version): cs=0x33 ds=0 es=0 ss=0x2b fs=0 gs=0 fs_base=0x7f407b3b8540 gs_base=0 self=0x7f407b3b8540 canary=0xf23424f90fbf7300 lsl(0x7b)=0xa cpu=10 node=0 hwcap2=0x2 &counter=0x5597bc530040
The canary always ends in 00. The low byte is deliberately zero so that string functions such as strcpy stop at it and cannot leak or overwrite the whole value.
; segs.asm — nasm -f elf64 segs.asm && ld segs.o -o segs && gdb -q ./segs
section .text
global _start
_start:
    mov  ax, cs          ; ax = 0x33   (reading is allowed!)
    mov  bx, ss          ; bx = 0x2b
    mov  cx, ds          ; cx = 0
    and  ax, 3           ; ax = CPL = 3

    mov  dx, 0x2b
    mov  ds, dx          ; legal: valid DPL-3 data selector, base still 0 → no effect
    ;mov cs, dx         ; ← nasm refuses; hand-encoded 8E CA → #UD (SIGILL)
    ;mov ds, 0x2b       ; ← no such instruction: sreg ← imm doesn't exist
    ;mov dx, 0x50       ; GDT index 10 (LDT descriptor, not data)
    ;mov ds, dx         ; ← #GP → SIGSEGV   (dmesg: "general protection fault")

    mov  rax, [fs:0]     ; static binary without libc: fs_base = 0 → SIGSEGV reading address 0
                          ; (comment this out, or set fs_base first with arch_prctl = syscall 158)
    mov  eax, 60
    xor  edi, edi
    syscall
A static nasm binary has no libc, so nothing ever sets fs_base. You can set it yourself: mov edi, 0x1002 (ARCH_SET_FS) · lea rsi, [rel mybuf] · mov eax, 158 · syscall. After that, mov rax, [fs:0] reads mybuf. This is exactly what ld.so/libc do for TLS.
WhereWhat it tells you about segments
/proc/PID/maps, /proc/PID/smapsThe real memory layout (ELF PT_LOADs, heap, stack, TLS lives in an anonymous mapping next to libc). This is what "segments" mean today.
/proc/PID/syscallSyscall number, 6 args, SP and PC of a blocked task. No segment registers.
/proc/PID/stat fields 29–30kstkesp/kstkeip, zeroed for security unless the task is dumping core.
/proc/cpuinfo flagslm (long mode), pae, fsgsbase (rd/wrfsbase), umip (blocks sgdt in user mode), la57 (5-level paging). Your CPU: lm pae fsgsbase, no umip, no la57 measured.
LD_SHOW_AUXV=1 /bin/trueAT_HWCAP2: 0x2 means the kernel lets user space use FSGSBASE.
Core dump: gcore PID → eu-readelf -n core.PIDThe NT_PRSTATUS note contains cs, ss, ds, es, fs, gs, fs_base, gs_base for each thread. It is a frozen user_regs_struct.
ptrace(PTRACE_GETREGS)This is how gdb gets them. struct user_regs_struct (sys/user.h) has fields cs ss ds es fs gs fs_base gs_base.
dmesg after a crashtraps: seg[1234] general protection fault ip:… sp:… error:50. The error code is the selector that failed (0x50 here).
sgdt in user modeReturns the GDT address. It is unprivileged on old CPUs. With UMIP, Linux emulates it and returns a dummy value (or sends SIGSEGV). The KASLR hardening happened partly because of this leak.
pid=$(pgrep -n seg); cat /proc/$pid/maps; cat /proc/$pid/syscall
grep -o -w -E 'fsgsbase|umip|la57|pae|lm' /proc/cpuinfo | sort | uniq -c
LD_SHOW_AUXV=1 /bin/true | grep HWCAP2

Binary Ninja, IDA, Ghidra and objdump are static: they show you which segment an instruction uses, not its value. The patterns to recognize:

; objdump -d -M intel ./a.out   (Binary Ninja shows the same)
64 48 8b 04 25 28 00 00 00   mov  rax, QWORD PTR fs:0x28     ; stack canary load (prologue)
64 48 2b 04 25 28 00 00 00   sub  rax, QWORD PTR fs:0x28     ; canary check (epilogue) → __stack_chk_fail
64 48 8b 04 25 00 00 00 00   mov  rax, QWORD PTR fs:0x0      ; TLS self pointer / pthread_self
64 8b 04 25 f8 ff ff ff      mov  eax, DWORD PTR fs:0xfffffffffffffff8 ; __thread var (local-exec TLS, negative offset)
48 8b 05 c0 2e 00 00         mov  rax, QWORD PTR [rip+0x2ec0]  ; global in .data (PIE) — no segment at all
2e 0f 1f 84 00 00 00 00 00   cs nop WORD PTR [rax+rax*1]    ; padding! 2E prefix is meaningless here
3e ff e0                     notrack jmp rax                 ; 3E (DS prefix) reused as CET "notrack"
When you see fs:0x28 in a function, it has a stack protector, so -fstack-protector was on. A gs: operand in a Windows x64 binary means TEB access: gs:0x30 = TEB self, gs:0x60 = PEB, a classic anti-debug and shellcode pattern. On 32-bit Windows the same thing is fs:0x18 / fs:0x30.
Binary Ninja's debugger (Debugger sidebar → Registers) shows live values, and it uses gdb/lldb backends underneath, so you see the same thing as in the gdb tab. In static analysis, look at the "segment" field in the instruction's disassembly tokens, and at the fsbase/gsbase pseudo-registers in the IL.
Register64-bit process32-bit process (gcc -m32) on the same kernel
CS0x33 (GDT 6, L=1)0x23 (GDT 4, L=0 D=1)
SS0x2b0x2b
DS / ES0 (null, allowed in 64-bit)0x2b (must be a valid selector in 32-bit mode)
FS0 + fs_base MSR = TLS0
GS00x63 = GDT 12 (a TLS slot set by set_thread_area()). The canary is at gs:0x14.
Why does 32-bit use GS for TLS and 64-bit use FS? On i386, glibc took GS because Wine and Windows use FS. On x86-64 the kernel needed GS for per-CPU data (because of swapgs), so user-space TLS moved to FS.
To build 32-bit programs you need apt install gcc-multilib. On your machine gcc -m32 produced no working binary, so the multilib files are probably missing.

Q54 GB address space, more than 4 GB of RAM, and 64-bit

The real mechanism is paging, with the segment step fixed at base 0. Every address goes through logical (seg:offset) → [segmentation: + base 0] → linear → [paging: CR3 page tables] → physical. Each process has its own page tables, so each gets its own private 4 GB (in 32-bit) of virtual addresses. The OS scatters those 4 KiB pages anywhere in physical RAM. With PAE, page-table entries are 64-bit, so physical frames can sit above 4 GB (up to 64 GB). A single process still only sees 4 GB at a time, but different processes together can use all the RAM.

The address pipeline (animated)

mov eax, [ebx+8] / mov rax, fs:[0x28]

Many processes → one physical RAM

Mode: Process:

Page-table walk: take a virtual address apart

VA
32-bit (no PAE)
VA 32 bit → PA 32 bit
2 levels · 4 KiB or 4 MiB pages
4 GiB virtual per process
4 GiB physical max
Linux: 3 GiB user / 1 GiB kernel ("3G/1G split")
32-bit PAE
VA 32 bit → PA 36 bit (up to 52)
3 levels, 8-byte PTEs, NX bit appears
4 GiB virtual per process
64 GiB physical
Kernel must use HIGHMEM tricks to reach RAM above ~896 MiB
x86-64
VA 48 bit (57 with LA57) → PA up to 52 bit
4 (or 5) levels, 2 MiB/1 GiB huge pages
128 TiB user + 128 TiB kernel
(64 PiB each with LA57)
Addresses must be canonical: bits 63..47 all equal
Why does 0x00007fffffffffff look so random? It is the last canonical user address with 48-bit VAs. Everything from 0x0000800000000000 to 0xffff7fffffffffff is the non-canonical hole. Touching it raises #GP, not a page fault, so you get SIGSEGV or SIGBUS with no faulting address. Kernel addresses start at 0xffff800000000000, and the direct map of all RAM is at 0xffff888000000000.
CR3 holds the physical address of the top-level table. On a context switch the kernel reloads CR3, and that is what "gives each process its own 4 GB". With PCID the CPU tags TLB entries per address space, so a switch does not have to flush the TLB. KPTI (the Meltdown fix) gives each process two page tables, and CR3 switches on every syscall.

Cheat sheet

Default segment per access type
AccessDefaultOverride?
Instruction fetchCSnever
push/pop/call/ret, [rsp+…], [rbp+…]SSyes (except push/pop's own stack access)
String source [rsi] (movs, lods, cmps, outs)DSyes
String destination [rdi] (movs, stos, scas, ins)ESno
Everything elseDSyes
Override prefix bytes
26 = ES2E = CS36 = SS3E = DS64 = FS65 = GS

In 64-bit mode 26/2E/36/3E are ignored as overrides. 2E/3E are reused as old branch hints and as CET notrack, and cs nop is just padding.

Instructions that touch segment state
mov r16, sreg / mov sreg, r/m16read (any ring) / load with checks (not CS)
push fs · pop fspush/pop of CS/DS/ES/SS is invalid in 64-bit mode; FS/GS still allowed
lds lss lfs lgs lesload a far pointer (seg + offset). LDS/LES are invalid in 64-bit mode (those opcodes became VEX)
jmp far, call far, retf, iretqthe legal ways to change CS
syscall / sysretCS/SS from the STAR MSR (kernel 0x10/0x18, user 0x33/0x2b)
swapgs ring 0swaps GS_BASE ↔ KERNEL_GS_BASE
rdfsbase/wrfsbase/rdgsbase/wrgsbasering 3 only if CR4.FSGSBASE is set (Linux ≥ 5.9)
arch_prctl(ARCH_SET_FS / GET_FS)syscall 158. Codes: SET_GS 0x1001, SET_FS 0x1002, GET_FS 0x1003, GET_GS 0x1004
lgdt lidt lldt ltr ring 0load descriptor tables
lsl lar verr verwquery a descriptor (unprivileged)
Faults you will meet
#GP(sel)loading a bad selector (outside the GDT, wrong type, privilege violation); error code = selector
#GP(0)using a null DS/ES/FS/GS in 32-bit mode, exceeding a limit, a non-canonical address
#SSstack-segment limit violation, or a non-canonical address through RSP/RBP
#NPdescriptor Present bit = 0
#UDmov cs, …, or rdfsbase when not enabled
Linux signal#GP → SIGSEGV (si_addr = 0!), #UD → SIGILL
What each register is used for today
Linux x86-64Windows x64
CS0x33 user / 0x10 kernel. Mode + CPL0x33 / 0x10
SS0x2b user / 0x18 kernel (or 0)0x2b / 0x18
DS00x2b
ES00x2b
FSuser TLS (glibc TCB)0x53 (32-bit compat TEB)
GSkernel per-CPU (via swapgs)user TEB / kernel KPCR

All tips & tricks, collected

  • Mental model for 64-bit: linear address = RIP/registers math + (FS or GS base if prefixed, otherwise 0). That is all segmentation does now.
  • CPL = CS & 3. To know which ring you are in, look at CS: 0x33 means ring 3 and 0x10 means ring 0.
  • There is no mov sreg, imm. Always go through a GPR: mov ax, 0x1000 / mov ds, ax.
  • mov ss, … (and pop ss) block interrupts and debug traps for one instruction so mov ss/mov sp is atomic. That quirk caused CVE-2018-8897 ("POP SS" debug exception bug) in several OS kernels.
  • Loading a selector reads the GDT once. If the GDT is changed later, the hidden cache is stale until the register is reloaded. "Unreal mode" (big real mode) abuses this: load a 4 GiB limit in protected mode, drop back to real mode, and keep the 4 GiB limit.
  • In real mode, up to 4096 different seg:off pairs point at the same byte. The normalized form puts the offset in 0..15.
  • The BIOS jumps to the boot sector at 0000:7C00, but some BIOSes use 07C0:0000. Good bootloaders do a far jump first to normalize CS.
  • At reset the CPU starts at F000:FFF0 but with a hidden CS base of 0xFFFF0000, so the first fetch is at physical 0xFFFFFFF0, just below 4 GiB. That is the hidden cache again.
  • 8086 opcode 0F was pop cs. It was removed on the 286 and became the two-byte opcode escape used by syscall (0F 05), cpuid (0F A2) and many others.
  • glibc's TCB layout on x86-64: fs:0x00 self pointer (tcb), fs:0x10 self (struct pthread), fs:0x28 stack guard (canary), fs:0x30 pointer guard (used to mangle setjmp/atexit pointers). __thread variables sit at negative offsets from fs_base.
  • All threads share one canary value, but each thread has its own copy in its own TCB, reached through its own fs_base.
  • gdb -ex starti stops before ld.so runs: fs_base = 0 and TLS is not set up yet. catch syscall arch_prctl shows the moment TLS appears.
  • gdb disables ASLR by default (set disable-randomization on), which is why PIE binaries always load at 0x555555554000 in gdb. Use set disable-randomization off to see real behavior.
  • Kernel entry uses swapgs to reach its per-CPU area via GS. Forgetting or mis-speculating it led to the "SWAPGS" Spectre variant (CVE-2019-1125).
  • Heaven's Gate: a 32-bit Windows (WoW64) process can jmp far 0x33:addr to run 64-bit code. On Linux a 64-bit process can ljmp/lret to 0x23 and run 32-bit code. Malware uses this to hide from 32-bit hooks.
  • You can create your own segments on Linux: modify_ldt() installs LDT entries (selector TI=1, e.g. 0x07). Wine and DOSEMU use this. It is disabled on many kernels (CONFIG_MODIFY_LDT_SYSCALL).
  • The #GP error code tells you which selector was bad. Look for it in dmesg as error:NN.
  • lsl on selector 0x7b returns the CPU and NUMA node without a syscall. The vDSO uses it on CPUs without RDPID.
  • Paging, not segments, gives you R/W/X protection today. NX (bit 63 of a PTE) exists only with PAE or 64-bit page tables. That is one reason PAE mattered even on machines with less than 4 GB.
  • Huge pages cut the walk short: the PD entry with PS=1 maps 2 MiB, and the PDPT entry with PS=1 maps 1 GiB. Look for AnonHugePages in /proc/PID/smaps.
  • ELF "segments" (program headers) are unrelated to CPU segments. Don't let the shared name confuse you.