summaryrefslogtreecommitdiff
path: root/gemfeed/DRAFT-unveiling-ior-ng-part-2.md
diff options
context:
space:
mode:
Diffstat (limited to 'gemfeed/DRAFT-unveiling-ior-ng-part-2.md')
-rw-r--r--gemfeed/DRAFT-unveiling-ior-ng-part-2.md35
1 files changed, 19 insertions, 16 deletions
diff --git a/gemfeed/DRAFT-unveiling-ior-ng-part-2.md b/gemfeed/DRAFT-unveiling-ior-ng-part-2.md
index c4b81886..35a4fdf9 100644
--- a/gemfeed/DRAFT-unveiling-ior-ng-part-2.md
+++ b/gemfeed/DRAFT-unveiling-ior-ng-part-2.md
@@ -2,9 +2,9 @@
> Draft — not in the gemfeed yet. Promote with the usual rename + index dance.
-This is Part 2 of three. Part 1 is the demo-driven tour — what ior looks like, how the dashboard tabs work, how filtering and recording behave. This part is about the install dance for Rocky Linux 9 (with one annoying kernel-backport caveat) and, more interestingly, why you only have to do that dance on a single machine: the resulting binary is portable to every other Linux box thanks to CO-RE — Compile Once, Run Everywhere — plus full static linking. Part 3 is the under-the-hood companion (per-event schema, async-syscall caveats, the syscall-coverage probe generator, and post-mortem SQL on the parquet output).
+This is Part 2 of three. Part 1 is the demo-driven tour: what ior looks like, how the dashboard tabs work, how filtering and recording behave. This part is about the install dance for Rocky Linux 9 (with one annoying kernel-backport caveat) and, more interestingly, why you only have to do that dance on a single machine: the resulting binary is portable to every other Linux box thanks to CO-RE (Compile Once, Run Everywhere) plus full static linking. Part 3 is the under-the-hood companion (per-event schema, async-syscall caveats, the syscall-coverage probe generator, and post-mortem SQL on the parquet output).
-If you came here for the dashboard tour, that's Part 1. If you want to know how the data pipeline is shaped after you've got ior running, that's Part 3. This one is for the moment between "I want to try this" and "OK, it's running on the box I care about."
+If you came here for the dashboard tour, that's Part 1. If you want to know how the data pipeline is shaped, that's Part 3. This one is for the moment between "I want to try this" and "OK, it's running on the box I care about."
[Part 1: a guided tour](./DRAFT-unveiling-ior-ng-part-1.md)
[Part 3: under the hood (schema, probe generator, ClickHouse)](./DRAFT-unveiling-ior-ng-part-3.md)
@@ -13,6 +13,7 @@ If you came here for the dashboard tour, that's Part 1. If you want to know how
[![I/O Riot NG logo](./unveiling-ior-ng/00-logo.png "I/O Riot NG logo")](./unveiling-ior-ng/00-logo.png)
+[2026-05-08 Unveiling I/O Riot NG 1.0.0 — Part 1: a guided tour](./2026-05-08-unveiling-ior-ng-part-1.md)
## Table of Contents
@@ -44,11 +45,11 @@ That's the officially supported install path, and it's the right one for anyone
If you're curious why Docker became the answer, the native install on Rocky Linux 9 illustrates the problem well. Three separate things bite you before you even get to `mage build`:
-Rocky 9 ships neither `libelf.a` nor `libzstd.a` — there are no `*-static` subpackages for either, only the dynamic `.so` files. Both have to be compiled from source. `libelf` from the elfutils source RPM, `libzstd` from the upstream GitHub release tarball.
+Rocky 9 ships neither `libelf.a` nor `libzstd.a`. There are no `*-static` subpackages for either, only the dynamic `.so` files. Both have to be compiled from source. `libelf` from the elfutils source RPM, `libzstd` from the upstream GitHub release tarball.
Rocky 9 also only ships Go 1.25.x, but ior requires 1.26+. So Go itself has to be installed from go.dev in parallel with the library builds.
-And there's a kernel quirk that used to make this section much longer. Pre-fix, ior would happily load on a stock 5.14 RHEL kernel and then die on the very first tracepoint attach with `BPF_LINK_CREATE`/`BPF_PERF_EVENT` returning `EACCES` — even as root, with SELinux permissive, with every BPF-related sysctl wide open. The cause is that RHEL 9 carries an `rt`-tree backport that adds `preempt_lazy_count` to `struct trace_entry`. That widens the BTF-emitted alias `trace_event_raw_sys_enter`/`_exit` by 8 bytes and shifts the `args`/`ret` offsets — but the actual context the kernel hands the BPF program is still `struct syscall_trace_enter`/`_exit`, where the offsets did not move. Programs written against `trace_event_raw_sys_*` (the conventional choice; bcc, libbpf-tools, and ior all used to do this) end up reading past `max_ctx_offset`, so the verifier rejects the attach. The fix — also what bcc shipped in [PR #4920](https://github.com/iovisor/bcc/pull/4920) and what inspektor-gadget did — is to type the BPF context as `syscall_trace_enter`/`_exit` directly. ior now generates its handlers that way, and stock 5.14 RHEL/Rocky/Alma works without an ElRepo kernel.
+And there's a kernel quirk that used to make this section much longer. Pre-fix, ior would happily load on a stock 5.14 RHEL kernel and then die on the very first tracepoint attach with `BPF_LINK_CREATE`/`BPF_PERF_EVENT` returning `EACCES`, even as root, with SELinux permissive, with every BPF-related sysctl wide open. The cause is that RHEL 9 carries an `rt`-tree backport that adds `preempt_lazy_count` to `struct trace_entry`. That widens the BTF-emitted alias `trace_event_raw_sys_enter`/`_exit` by 8 bytes and shifts the `args`/`ret` offsets, but the actual context the kernel hands the BPF program is still `struct syscall_trace_enter`/`_exit`, where the offsets did not move. Programs written against `trace_event_raw_sys_*` (the conventional choice; bcc, libbpf-tools, and ior all used to do this) end up reading past `max_ctx_offset`, so the verifier rejects the attach. The fix (also what bcc shipped in [PR #4920](https://github.com/iovisor/bcc/pull/4920) and what inspektor-gadget did) is to type the BPF context as `syscall_trace_enter`/`_exit` directly. ior now generates its handlers that way, and stock 5.14 RHEL/Rocky/Alma works without an ElRepo kernel.
### What the Docker build is actually doing
@@ -114,32 +115,32 @@ If you see `Probing for 5s` followed by CSV rows, the build is good. `mage build
If you haven't touched eBPF before: it's a small in-kernel bytecode VM. You compile a tiny C program, the kernel verifies it can't crash or loop forever, and then it runs every time some hook fires — a syscall enter/exit, a kprobe, a tracepoint, a network packet. The program writes events into a ring buffer that userspace mmaps and drains. No kernel module, no patched kernel, no debug symbols required.
-ior plugs into the syscall tracepoints — `sys_enter_openat`, `sys_exit_read`, etc. — and the BPF side does the bare minimum: timestamp the event, copy a few fields, push to a perf ring buffer. All the heavy lifting (string interning, latency math, aggregation, the dashboard) is in Go on the userspace side.
+ior plugs into the syscall tracepoints (`sys_enter_openat`, `sys_exit_read`, etc.) and the BPF side does the bare minimum: timestamp the event, copy a few fields, push to a perf ring buffer. All the heavy lifting (string interning, latency math, aggregation, the dashboard) is in Go on the userspace side.
The kernel ships a C library called libbpf that handles loading the program, attaching it to hooks, managing maps, and reading the ring buffer. There are two well-known ways to drive that from Go:
* libbpfgo (Aqua Security): a thin cgo wrapper around libbpf. You ship libbpf along with your binary and call into the same C API that `bpftool` and `perf` use.
-* cilium/ebpf: a from-scratch pure-Go reimplementation of everything libbpf does — ELF parser, BTF resolver, syscall layer, the lot.
+* cilium/ebpf: a from-scratch pure-Go reimplementation of everything libbpf does (ELF parser, BTF resolver, syscall layer, the lot).
-I went with libbpfgo specifically because it's a wrapper, not a reimplementation. Whatever lands in libbpf upstream — new map types, new attach kinds, CO-RE fixes — I get for free the next kernel cycle. The pure-Go variant has to chase libbpf's feature set in parallel, and any divergence is on me to debug. For a tracer that's mostly value-add on the userspace side, "be a thin client of the kernel's own library" wins.
+I went with libbpfgo specifically because it's a wrapper, not a reimplementation. Whatever lands in libbpf upstream (new map types, new attach kinds, CO-RE fixes) I get for free the next kernel cycle. The pure-Go variant has to chase libbpf's feature set in parallel, and any divergence is on me to debug. For a tracer that's mostly value-add on the userspace side, "be a thin client of the kernel's own library" wins.
## CO-RE — the part that makes the binary actually portable
The headline fact about ior's deployment story: build it once on one box, then `scp ior other-host:/usr/local/bin/` to anywhere else and it just runs. No recompile per kernel, no kernel-debuginfo dance, no DKMS hooks. Two mechanisms make that work, and they reinforce each other.
-The first is plain old static linking on the userspace side. A quick refresher on what that means, since it's central to why "scp the binary anywhere" works: when you build a normal Linux executable, the linker has two ways to wire library code into your program. Dynamic linking ("shared library") leaves a placeholder in the binary that says "at run time, find `libfoo.so.6` somewhere on `LD_LIBRARY_PATH` and pull in its symbols." Static linking pastes the library's machine code directly into your binary at build time, so there's nothing to look up later. Dynamic is smaller on disk and lets distros patch shared libs without rebuilding everything; static is bigger but self-contained — no surprise about which version of the library the target box happens to have, no `error while loading shared libraries: libwhatever.so.6: cannot open shared object file` when the target ships a newer ABI.
+The first is plain old static linking on the userspace side. A quick refresher on what that means, since it's central to why "scp the binary anywhere" works: when you build a normal Linux executable, the linker has two ways to wire library code into your program. Dynamic linking ("shared library") leaves a placeholder in the binary that says "at run time, find `libfoo.so.6` somewhere on `LD_LIBRARY_PATH` and pull in its symbols." Static linking pastes the library's machine code directly into your binary at build time, so there's nothing to look up later. Dynamic is smaller on disk and lets distros patch shared libs without rebuilding everything; static is bigger but self-contained, with no surprise about which version of the library the target box happens to have, no `error while loading shared libraries: libwhatever.so.6: cannot open shared object file` when the target ships a newer ABI.
-For Go, this is mostly a non-issue. A pure-Go binary (no cgo) is statically linked by default — the Go toolchain produces a single self-contained ELF file with no `.dynamic` section and no `NEEDED` entries. You can `scp` it to any Linux box of the same architecture and it just runs. That's one of the quietly nice things about Go.
+For Go, this is mostly a non-issue. A pure-Go binary (no cgo) is statically linked by default. The Go toolchain produces a single self-contained ELF file with no `.dynamic` section and no `NEEDED` entries. You can `scp` it to any Linux box of the same architecture and it just runs. That's one of the quietly nice things about Go.
-ior is the not-quite-pure case: it goes through cgo to call into libbpf, libelf, and libzstd, and each of those has its own .so on the build host. By default cgo links those C dependencies dynamically, which would defeat the "scp the binary anywhere" property — the target box would need to have matching `.so` files at matching versions, which is exactly the kind of dependency hell Go usually saves you from. The fix is the line `-extldflags "-static"` in ior's Magefile: it tells the external (C) linker to resolve `-lbpf -lelf -lzstd -lz` against the static archives (`.a` files) instead of the dynamic ones. That's why the install procedure above is so picky about having `libelf.a` and `libzstd.a` actually present on the build host — without them the C-side static link fails outright.
+ior is the not-quite-pure case: it goes through cgo to call into libbpf, libelf, and libzstd, and each of those has its own .so on the build host. By default cgo links those C dependencies dynamically, which would defeat the "scp the binary anywhere" property: the target box would need to have matching `.so` files at matching versions, which is exactly the kind of dependency hell Go usually saves you from. The fix is the line `-extldflags "-static"` in ior's Magefile: it tells the external (C) linker to resolve `-lbpf -lelf -lzstd -lz` against the static archives (`.a` files) instead of the dynamic ones. That's why the install procedure above is so picky about having `libelf.a` and `libzstd.a` actually present on the build host. Without them the C-side static link fails outright.
The result is a single ~23 MB binary with libbpf, libelf, libzstd, and zlib all baked in. None of them are looked up dynamically at runtime. The build host's library versions stay on the build host. (A couple of glibc resolver functions — `getpwnam_r` and friends — do still fall back to the target's libc, which is fine on any reasonable distro and is what the linker warnings during the build are about.)
-The second, and the one that's actually unusual, is CO-RE — Compile Once, Run Everywhere. CO-RE is the eBPF feature that solves the "the kernel changed its struct layout between releases" problem.
+The second, and the one that's actually unusual, is CO-RE (Compile Once, Run Everywhere). CO-RE is the eBPF feature that solves the "the kernel changed its struct layout between releases" problem.
-The old I/O Riot was Systemtap. Systemtap programs are translated into a kernel module against the running kernel's exact headers, and that module then has to be loaded with `insmod`. That meant: the user has to install a kernel-debuginfo package matching their running kernel, and a fresh build per host (or per kernel update). On the BSD-style "you only run what you compiled here" laptop crowd that was tolerable; on a fleet of distros + kernel versions it was a recurring tax. Half of the original I/O Riot's README was about kernel-debuginfo dance steps.
+The old I/O Riot was Systemtap. Systemtap programs are translated into a kernel module against the running kernel's exact headers, and that module then has to be loaded with `insmod`. That meant the user has to install a kernel-debuginfo package matching their running kernel, and a fresh build per host (or per kernel update). On the BSD-style "you only run what you compiled here" laptop crowd that was tolerable; on a fleet of distros + kernel versions it was a recurring tax. Half of the original I/O Riot's README was about kernel-debuginfo dance steps.
-CO-RE throws all of that out. The idea, in one paragraph: when you write a BPF program that reads `task->mm->start_stack`, you don't bake the offsets of those fields into the compiled program. Instead, the compiler emits relocation records ("at this instruction, fetch the offset of `mm` inside `task_struct`"). At load time, libbpf looks up the actual offsets in the target kernel's BTF (BPF Type Format — a description of every kernel struct, embedded in `/sys/kernel/btf/vmlinux` on any modern kernel) and patches the program in place. The same `.bpf.o` that ran on a 5.10 Debian kernel runs on a 6.8 Fedora kernel without recompilation.
+CO-RE throws all of that out. The idea, in one paragraph: when you write a BPF program that reads `task->mm->start_stack`, you don't bake the offsets of those fields into the compiled program. Instead, the compiler emits relocation records ("at this instruction, fetch the offset of `mm` inside `task_struct`"). At load time, libbpf looks up the actual offsets in the target kernel's BTF (BPF Type Format, a description of every kernel struct embedded in `/sys/kernel/btf/vmlinux` on any modern kernel) and patches the program in place. The same `.bpf.o` that ran on a 5.10 Debian kernel runs on a 6.8 Fedora kernel without recompilation.
Pictorially, the contrast looks like this:
@@ -191,9 +192,11 @@ The runtime shape of a trace pipeline lines up with that:
## A note on cgo overhead
-The cost of being a libbpf wrapper rather than a pure-Go reimplementation is cgo. Every call from Go into libbpf crosses the cgo boundary, which historically meant tens to ~hundred-ish nanoseconds of overhead per call — register save/restore, a stack switch onto g0, goroutine state bookkeeping. Cheap in absolute terms, but it adds up if you call into C inside a tight loop. ior keeps the actual hot path on the kernel side and only crosses into Go once per drained batch of events from the ring buffer, so the per-call cost is amortized over thousands of events. In practice it doesn't show up in profiles.
+The cost of being a libbpf wrapper rather than a pure-Go reimplementation is cgo. Every call from Go into libbpf crosses the cgo boundary, which historically meant tens to ~hundred-ish nanoseconds of overhead per call: register save/restore, a stack switch onto g0, goroutine state bookkeeping. Cheap in absolute terms, but it adds up if you call into C inside a tight loop. ior keeps the actual hot path on the kernel side and only crosses into Go once per drained batch of events from the ring buffer, so the per-call cost is amortized over thousands of events. In practice it doesn't show up in profiles.
-Go 1.26, the current release at the time of writing (early May 2026), is the one that finally took a serious bite out of cgo's per-call cost — the runtime can elide a chunk of the bookkeeping for calls that don't need it. Real-world wins depend heavily on the workload, but the rough direction is that cgo now feels closer to "an unusually expensive function call" than to "a context switch", which is the right mental model for almost everyone touching a C library from Go. The shorter version: cgo overhead used to be a real footgun for ports that called into C in the inner loop. With Go 1.26 it's a footnote unless you're doing many millions of small calls per second, in which case batching across the boundary still fixes it.
+Go 1.26, the current release at the time of writing (early May 2026), is the one that finally took a serious bite out of cgo's per-call cost. The runtime can elide a chunk of the bookkeeping for calls that don't need it. Real-world wins depend heavily on the workload, but the rough direction is that cgo now feels closer to "an unusually expensive function call" than to "a context switch", which is the right mental model for almost everyone touching a C library from Go. The shorter version: cgo overhead used to be a real footgun for ports that called into C in the inner loop. With Go 1.26 it's a footnote unless you're doing many millions of small calls per second, in which case batching across the boundary still fixes it.
+
+If you want to verify any of that on your own workload, ior bakes in a small benchmark suite (`mage bench`, plus `mage prReview` which is the full clean+generate+test+build+benchmark baseline) that drops CPU/mem/block profiles into `bench-profiles/` for `go tool pprof` to chew on. That's the same path I run before tagging any release I plan to copy to the homelab.
## If you want to go deeper
@@ -206,7 +209,7 @@ Between the two, Rice teaches you the moving parts and Gregg teaches you what to
## Wrapping up
-That's the install dance and the why-it's-portable story. Part 3 is the bottom of the data stack — what's actually in each event row, the syscall-coverage safeguard against new kernels, async-syscall caveats, and how to query the parquet output with ClickHouse Local. Part 1, if you haven't read it, is the demo-driven tour with all the GIFs of the dashboard.
+That's the install dance and the why-it's-portable story. Part 3 is the bottom of the data stack: what's actually in each event row, the syscall-coverage safeguard against new kernels, async-syscall caveats, the integration test harness that pins coverage in place, and how to query the parquet output with ClickHouse Local. Part 1, if you haven't read it, is the demo-driven tour with all the GIFs of the dashboard.
[Part 1: a guided tour](./DRAFT-unveiling-ior-ng-part-1.md)
[Part 3: under the hood (schema, probe generator, ClickHouse)](./DRAFT-unveiling-ior-ng-part-3.md)