diff options
Diffstat (limited to 'gemfeed/DRAFT-ior-guided-tour.md')
| -rw-r--r-- | gemfeed/DRAFT-ior-guided-tour.md | 213 |
1 files changed, 0 insertions, 213 deletions
diff --git a/gemfeed/DRAFT-ior-guided-tour.md b/gemfeed/DRAFT-ior-guided-tour.md deleted file mode 100644 index df52e070..00000000 --- a/gemfeed/DRAFT-ior-guided-tour.md +++ /dev/null @@ -1,213 +0,0 @@ -# I/O Riot NG: a guided tour - -> Draft — not in the gemfeed yet. Promote with the usual rename + index dance. - -I rewrote I/O Riot. The old one was C + Systemtap and dates from 2017. The new one — call it ior — is Go + C + BPF via libbpfgo, runs on Linux, and is mostly a TUI dashboard rather than a record/replay box. Since pictures are worth more than yet another README table of key bindings, I built a demo. - -``` - .---. - / \ - \.@-@./ - /`\_/`\ - // _ \\ - | \ )|_ - /`\_`> <_/ \ -jgs\__/'---'\__/ -``` - -[I/O Riot NG on Codeberg](https://codeberg.org/snonux/ior) -[the original I/O Riot post (2018)](./2018-06-01-realistic-load-testing-with-ioriot-for-linux.md) - -## Table of Contents - -* [⇢ I/O Riot NG: a guided tour](#io-riot-ng-a-guided-tour) -* [⇢ ⇢ What it does](#what-it-does) -* [⇢ ⇢ A short detour: eBPF and libbpfgo](#a-short-detour-ebpf-and-libbpfgo) -* [⇢ ⇢ The whole thing as a tape pipeline](#the-whole-thing-as-a-tape-pipeline) -* [⇢ ⇢ First launch](#first-launch) -* [⇢ ⇢ The seven tabs, in 30 seconds each](#the-seven-tabs-in-30-seconds-each) -* [⇢ ⇢ The Stream tab is the good one](#the-stream-tab-is-the-good-one) -* [⇢ ⇢ Filtering, more thoroughly](#filtering-more-thoroughly) -* [⇢ ⇢ Recording](#recording) -* [⇢ ⇢ Querying a parquet trace with ClickHouse](#querying-a-parquet-trace-with-clickhouse) -* [⇢ ⇢ Reproducing the whole demo](#reproducing-the-whole-demo) -* [⇢ ⇢ What's still missing](#what-s-still-missing) - -## What it does - -ior attaches BPF tracepoints to a chunk of the synchronous-I/O syscall surface — open, read, write, stat, mmap, sync, link, fcntl, dup, the obvious ones. Each enter/exit pair becomes an event with a duration plus an inter-syscall gap, and the events feed a Bubble Tea dashboard with seven tabs: a live flamegraph, an overview, sortable per-syscall / per-file / per-process tables, latency histograms, and a live event stream with a stackable filter UI on top. - -Same shape as the old I/O Riot in spirit: capture what the system is actually doing, not synthetic load. Different shape in execution: no replay engine, no separate record file unless you ask for one, no kernel-debug-info dance. - -## A short detour: eBPF and libbpfgo - -If you haven't touched eBPF before: it's a small in-kernel bytecode VM. You compile a tiny C program, the kernel verifies it can't crash or loop forever, and then it runs every time some hook fires — a syscall enter/exit, a kprobe, a tracepoint, a network packet. The program writes events into a ring buffer that userspace mmaps and drains. No kernel module, no patched kernel, no debug symbols required. - -ior plugs into the syscall tracepoints — `sys_enter_openat`, `sys_exit_read`, etc. — and the BPF side does the bare minimum: timestamp the event, copy a few fields, push to a perf ring buffer. All the heavy lifting (string interning, latency math, aggregation, the dashboard) is in Go on the userspace side. - -The kernel ships a C library called libbpf that handles loading the program, attaching it to hooks, managing maps, and reading the ring buffer. There are two well-known ways to drive that from Go: - -* libbpfgo (Aqua Security): a thin cgo wrapper around libbpf. You ship libbpf along with your binary and call into the same C API that `bpftool` and `perf` use. -* cilium/ebpf: a from-scratch pure-Go reimplementation of everything libbpf does — ELF parser, BTF resolver, syscall layer, the lot. - -I went with libbpfgo specifically because it's a wrapper, not a reimplementation. Whatever lands in libbpf upstream — new map types, new attach kinds, CO-RE fixes — I get for free the next kernel cycle. The pure-Go variant has to chase libbpf's feature set in parallel, and any divergence is on me to debug. For a tracer that's mostly value-add on the userspace side, "be a thin client of the kernel's own library" wins. - -The cost is cgo. Every call from Go into libbpf crosses the cgo boundary, which historically meant tens to ~hundred-ish nanoseconds of overhead per call — register save/restore, a stack switch onto g0, goroutine state bookkeeping. Cheap in absolute terms, but it adds up if you call into C inside a tight loop. ior keeps the actual hot path on the kernel side and only crosses into Go once per drained batch of events from the ring buffer, so the per-call cost is amortized over thousands of events. In practice it doesn't show up in profiles. - -Go 1.26, the current release at the time of writing (late April 2026), is the one that finally took a serious bite out of cgo's per-call cost — the runtime can elide a chunk of the bookkeeping for calls that don't need it. Real-world wins depend heavily on the workload, but the rough direction is that cgo now feels closer to "an unusually expensive function call" than to "a context switch", which is the right mental model for almost everyone touching a C library from Go. The shorter version: cgo overhead used to be a real footgun for ports that called into C in the inner loop. With Go 1.26 it's a footnote unless you're doing many millions of small calls per second, in which case batching across the boundary still fixes it. - -## The whole thing as a tape pipeline - -The demo isn't a screencast I sat through. It's 14 VHS tapes that drive the TUI deterministically, with a background workload generator producing real syscall traffic for the trace to chew on. One `mage demo` and every GIF below regenerates from scratch. The boring part of "make a demo" — having to re-record everything when the UI shifts — goes away. - -## First launch - -```sh -sudo ./ior -``` - -You land on the PID picker. The default selection is "All PIDs", so Enter just dumps you straight at the dashboard. - -[](./ior-guided-tour/01-launch.gif) - -The dashboard opens on the **live flamegraph**. Bars grow as new events arrive. The whole thing is keyboard-driven: `h`/`l` walk siblings at the current depth, `j`/`k` step deeper or shallower, `enter` zooms into the selected subtree (the rest of the chart greys out and the selection becomes the new root), `u` or `Esc` undoes the zoom. `o` cycles the stack ordering — `comm/path/tracepoint`, `path/tracepoint/comm`, etc. — and `b` toggles the metric driving bar width between event count and total bytes. - -[](./ior-guided-tour/13-tui-flamegraph.gif) - -## The seven tabs, in 30 seconds each - -The number keys jump between tabs. `tab` and `shift+tab` step. - -`2` is **Overview** — a sparkline plus the top syscalls and top paths, the at-a-glance view. - -[](./ior-guided-tour/02-overview-tab.gif) - -`3` is **Syscalls** — a sortable table. `s` sorts by the selected column, `S` reverses. - -[](./ior-guided-tour/03-syscalls-tab.gif) - -`4` is **Files**. The interesting key here is `d`: it rolls per-file rows up into their parent directory. Essential when you've got a process touching ten thousand files in `/usr/share/`. - -[](./ior-guided-tour/04-files-tab.gif) - -`5` is **Processes** — same idea, per-process / per-comm. - -[](./ior-guided-tour/05-processes-tab.gif) - -`6` is **Latency + Gaps**. Two histograms: how long each syscall took, and the idle-on-the-same-thread gap between syscalls. The dd loop in the demo workload spreads the latency distribution out so you can actually see it. - -[](./ior-guided-tour/06-latency-gaps-tab.gif) - -`7` is **Stream** — the live tail. This is where you spend most of your time when something's actually broken. - -[](./ior-guided-tour/07-stream-live.gif) - -## The Stream tab is the good one - -`space` pauses. In pause mode, `j`/`k` and arrow keys move the row/column cursor. Hitting `Enter` on a cell pushes a new filter onto a stack, narrowing what you see. Pile them up — comm, then syscall, then file — and `Esc` pops them off LIFO when you want to back out. - -[](./ior-guided-tour/08-stream-pause-filter.gif) - -`/` and `?` are regex search forward/backward. `n` and `N` walk matches. The search runs against every column in the ring buffer and wraps at the end. Search and filtering are different beasts: search highlights and jumps, filtering hides everything that doesn't match. - -[](./ior-guided-tour/09-stream-regex-search.gif) - -`e` exports the current filtered snapshot to a CSV in the working directory. `x` does the same for the paused stream view specifically (preserving your filter stack), `X` prompts for a filename, `E` opens the most recent export in `$EDITOR`. - -[](./ior-guided-tour/10-stream-csv-export.gif) - -## Filtering, more thoroughly - -The Enter-to-push trick isn't unique to Stream. It works the same on Files, Syscalls, and Processes: highlight a row, hit Enter, and the cell value becomes a filter against the entire dashboard. Three tabs of "I see one weird path / comm / syscall, drill in" with one keystroke. - -The filter status line gives you a one-glance summary of every active frame, written like: - -* `comm~bash` — substring match on a string column. This is what Enter-on-a-cell produces for `comm`, `syscall`, and `file`. -* `pid=1234` — exact equality. Used for `pid`, `tid`, `fd`, `ret`, `bytes`. -* `latency>=5ms` / `gap>=10us` — numeric comparison with a duration suffix. The full operator set is `>`, `<`, `=`, `>=`, `<=`, `!=`. - -Stack frames AND together, so pushing `comm~bash` and then `syscall~openat` shows you bash's openat calls, not bash OR openat. `Esc` (or `F`) pops the most recent frame. - -Two other knobs do related work: - -* `p`, `t`, `o` open the PID, TID, and probe-toggle dialogs. These are global filters: they reconfigure the BPF side, so kernel-level events for excluded PIDs/probes never even reach userspace. Cheaper than filtering a firehose, but it also means the filter applies to recordings (the parquet file only contains rows the kernel let through). -* The CLI mirrors of those dialogs let you bake the same scoping into a one-shot run: `-pid`, `-tid`, `-comm`, `-path`, plus `-tps <regex>` / `-tpsExclude <regex>` for picking which tracepoints to attach in the first place. - -[](./ior-guided-tour/11-pid-tid-probe.gif) - -## Recording - -Three persistence flows, each for a different job: - -* `R` from the dashboard starts streaming **Parquet** — every event row that survives your current TUI filter goes to disk continuously. `R` again stops. Footer shows the active file or the last error. - -[](./ior-guided-tour/12-parquet-recording.gif) - -* `sudo ./ior -flamegraph -name <n>` writes one aggregated `.ior.zst` artifact at shutdown. Aggregated counters, not per-event rows. Cheaper to write, ideal for ior's native flamegraph workflow and integration tests. - -* `sudo ./ior -parquet trace.parquet` is the headless firehose — every row, no TUI, no filtering. `sudo ./ior -plain` is even lighter: CSV to stdout, pipe it into anything. - -[](./ior-guided-tour/14-headless-modes.gif) - -## Querying a parquet trace with ClickHouse - -The schema is flat and stable: `seq, time_ns, gap_ns, latency_ns, comm, pid, tid, syscall, fd, ret, bytes, file, is_error, filter_epoch`. ClickHouse Local reads parquet directly without a server, which makes it a perfect post-mortem tool — point it at the file and run SQL: - -```sh -clickhouse local --query " - SELECT comm, syscall, count() AS n, formatReadableSize(sum(bytes)) AS total - FROM file('trace.parquet', Parquet) - GROUP BY comm, syscall - ORDER BY n DESC - LIMIT 10 -" -``` - -``` -bash read 18432 72.10 MiB -dd write 14209 1.39 GiB -fish openat 9871 0.00 B -systemd-journ… write 4112 1.62 MiB -... -``` - -The fields you actually want for performance work are `latency_ns` and `gap_ns`. P99 by syscall, only the ones that landed in error: - -```sh -clickhouse local --query " - SELECT - syscall, - count() AS n, - quantile(0.5)(latency_ns)/1000 AS p50_us, - quantile(0.99)(latency_ns)/1000 AS p99_us - FROM file('trace.parquet', Parquet) - WHERE is_error = 1 - GROUP BY syscall - ORDER BY p99_us DESC -" -``` - -Same trick works in DuckDB (`duckdb -c "SELECT ... FROM 'trace.parquet'"`), pandas, polars, anything that reads Parquet. The point of streaming Parquet rather than ior's native `.ior.zst` format is exactly this: once it's on disk, you're in the standard data-tools ecosystem. - -## Reproducing the whole demo - -```sh -mage installDemoTools # one-time: VHS via go install + ttyd from dnf -sudo -v # warm the sudo timestamp once -mage demo # ~10 minutes, fully headless, safe to background -``` - -To rebuild a single GIF after editing its tape: `TAPE=07-stream-live mage demoOne`. - -## What's still missing - -ior is pre-alpha and basically a personal tool. The headline gaps: - -* No record/replay — that was the whole point of the original I/O Riot. The new one is a tracer, not a workload simulator. I keep going back and forth on whether to put replay back in. -* No userspace symbol resolution. Stacks are at the syscall surface, not "which line of which library called read". -* No remote / cluster mode. Single host, one trace at a time. - -But the live flamegraph, the stackable stream filters, and the cheap parquet capture together cover the cases I actually hit week to week. The demo above is the easiest way to get a feel for whether it's the kind of tool you want. - -[Source on Codeberg](https://codeberg.org/snonux/ior) -[The full in-repo tutorial](https://codeberg.org/snonux/ior/src/branch/main/demo/TUTORIAL.md) |
