summaryrefslogtreecommitdiff
path: root/gemfeed/DRAFT-unveiling-ior-ng.html
blob: 634bdaf8aa85887781dfdee3c6af724bc3f96fd8 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
<html xmlns="http://www.w3.org/1999/xhtml" lang="en" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
<title>Unveiling I/O Riot NG</title>
<link rel="shortcut icon" type="image/gif" href="/favicon.ico" />
<link rel="stylesheet" href="../style.css" />
<link rel="stylesheet" href="style-override.css" />
</head>
<body>
<p class="header">
<a href="https://foo.zone">Home</a> | <a href="https://codeberg.org/snonux/foo.zone/src/branch/content-md/gemfeed/DRAFT-unveiling-ior-ng.md">Markdown</a> | <a href="gemini://foo.zone/gemfeed/DRAFT-unveiling-ior-ng.gmi">Gemini</a> | <a href="https://snonux.foo">Microblog</a> | <a href="https://irregular.ninja">Street photography</a>
</p>
<h1 style='display: inline' id='unveiling-io-riot-ng'>Unveiling I/O Riot NG</h1><br />
<br />
<span class='quote'>Draft — not in the gemfeed yet. Promote with the usual rename + index dance.</span><br />
<br />
<span>I rewrote I/O Riot. The old one was C + Systemtap and dates from 2017. The new one — call it ior — is Go + C + BPF via libbpfgo, runs on Linux, and is mostly a TUI dashboard rather than a record/replay box. Since pictures are worth more than yet another README table of key bindings, I built a demo.</span><br />
<br />
<a href='./unveiling-ior-ng/00-hero-flamegraph.png'><img alt='ior&#39;s live flamegraph: every running process, by file path, by syscall — width = event volume' title='ior&#39;s live flamegraph: every running process, by file path, by syscall — width = event volume' src='./unveiling-ior-ng/00-hero-flamegraph.png' /></a><br />
<br />
<a class='textlink' href='https://codeberg.org/snonux/ior'>I/O Riot NG on Codeberg</a><br />
<a class='textlink' href='./2018-06-01-realistic-load-testing-with-ioriot-for-linux.html'>the original I/O Riot post (2018)</a><br />
<br />
<h2 style='display: inline' id='table-of-contents'>Table of Contents</h2><br />
<br />
<ul>
<li><a href='#unveiling-io-riot-ng'>Unveiling I/O Riot NG</a></li>
<li>⇢ <a href='#what-it-does'>What it does</a></li>
<li>⇢ <a href='#a-short-detour-ebpf-and-libbpfgo'>A short detour: eBPF and libbpfgo</a></li>
<li>⇢ ⇢ <a href='#co-re--the-part-that-makes-ior-actually-portable'>CO-RE — the part that makes ior actually portable</a></li>
<li>⇢ ⇢ <a href='#if-you-want-to-go-deeper'>If you want to go deeper</a></li>
<li>⇢ <a href='#the-whole-thing-as-a-tape-pipeline'>The whole thing as a tape pipeline</a></li>
<li>⇢ <a href='#first-launch'>First launch</a></li>
<li>⇢ <a href='#the-seven-tabs-in-30-seconds-each'>The seven tabs, in 30 seconds each</a></li>
<li>⇢ ⇢ <a href='#2-overview'><span class='inlinecode'>2</span> Overview</a></li>
<li>⇢ ⇢ <a href='#3-syscalls'><span class='inlinecode'>3</span> Syscalls</a></li>
<li>⇢ ⇢ <a href='#4-files'><span class='inlinecode'>4</span> Files</a></li>
<li>⇢ ⇢ <a href='#5-processes'><span class='inlinecode'>5</span> Processes</a></li>
<li>⇢ ⇢ <a href='#6-latency--gaps'><span class='inlinecode'>6</span> Latency + Gaps</a></li>
<li>⇢ ⇢ <a href='#7-stream'><span class='inlinecode'>7</span> Stream</a></li>
<li>⇢ <a href='#the-stream-tab-is-the-good-one'>The Stream tab is the good one</a></li>
<li>⇢ <a href='#filtering-more-thoroughly'>Filtering, more thoroughly</a></li>
<li>⇢ <a href='#recording'>Recording</a></li>
<li>⇢ <a href='#querying-a-parquet-trace-with-clickhouse'>Querying a parquet trace with ClickHouse</a></li>
<li>⇢ <a href='#reproducing-the-whole-demo'>Reproducing the whole demo</a></li>
<li>⇢ <a href='#what-s-still-missing'>What&#39;s still missing</a></li>
</ul><br />
<h2 style='display: inline' id='what-it-does'>What it does</h2><br />
<br />
<span>ior attaches BPF tracepoints to a chunk of the synchronous-I/O syscall surface — open, read, write, stat, mmap, sync, link, fcntl, dup, the obvious ones. Each enter/exit pair becomes an event with a duration plus an inter-syscall gap, and the events feed a Bubble Tea dashboard with seven tabs: a live flamegraph, an overview, sortable per-syscall / per-file / per-process tables, latency histograms, and a live event stream with a stackable filter UI on top.</span><br />
<br />
<span>Same shape as the old I/O Riot in spirit: capture what the system is actually doing, not synthetic load. Different shape in execution: no replay engine, no separate record file unless you ask for one, no kernel-debug-info dance.</span><br />
<br />
<a href='./unveiling-ior-ng/00-logo.png'><img alt='I/O Riot NG logo' title='I/O Riot NG logo' src='./unveiling-ior-ng/00-logo.png' /></a><br />
<br />
<h2 style='display: inline' id='a-short-detour-ebpf-and-libbpfgo'>A short detour: eBPF and libbpfgo</h2><br />
<br />
<span>If you haven&#39;t touched eBPF before: it&#39;s a small in-kernel bytecode VM. You compile a tiny C program, the kernel verifies it can&#39;t crash or loop forever, and then it runs every time some hook fires — a syscall enter/exit, a kprobe, a tracepoint, a network packet. The program writes events into a ring buffer that userspace mmaps and drains. No kernel module, no patched kernel, no debug symbols required.</span><br />
<br />
<span>ior plugs into the syscall tracepoints — <span class='inlinecode'>sys_enter_openat</span>, <span class='inlinecode'>sys_exit_read</span>, etc. — and the BPF side does the bare minimum: timestamp the event, copy a few fields, push to a perf ring buffer. All the heavy lifting (string interning, latency math, aggregation, the dashboard) is in Go on the userspace side.</span><br />
<br />
<span>The kernel ships a C library called libbpf that handles loading the program, attaching it to hooks, managing maps, and reading the ring buffer. There are two well-known ways to drive that from Go:</span><br />
<br />
<ul>
<li>libbpfgo (Aqua Security): a thin cgo wrapper around libbpf. You ship libbpf along with your binary and call into the same C API that <span class='inlinecode'>bpftool</span> and <span class='inlinecode'>perf</span> use.</li>
<li>cilium/ebpf: a from-scratch pure-Go reimplementation of everything libbpf does — ELF parser, BTF resolver, syscall layer, the lot.</li>
</ul><br />
<span>I went with libbpfgo specifically because it&#39;s a wrapper, not a reimplementation. Whatever lands in libbpf upstream — new map types, new attach kinds, CO-RE fixes — I get for free the next kernel cycle. The pure-Go variant has to chase libbpf&#39;s feature set in parallel, and any divergence is on me to debug. For a tracer that&#39;s mostly value-add on the userspace side, "be a thin client of the kernel&#39;s own library" wins.</span><br />
<br />
<h3 style='display: inline' id='co-re--the-part-that-makes-ior-actually-portable'>CO-RE — the part that makes ior actually portable</h3><br />
<br />
<span>The old I/O Riot was Systemtap. Systemtap programs are translated into a kernel module against the running kernel&#39;s exact headers, and that module then has to be loaded with <span class='inlinecode'>insmod</span>. That meant: the user has to install a kernel-debuginfo package matching their running kernel, and a fresh build per host (or per kernel update). On the BSD-style "you only run what you compiled here" laptop crowd that was tolerable; on a fleet of distros + kernel versions it was a recurring tax. Half of the original I/O Riot&#39;s README was about kernel-debuginfo dance steps.</span><br />
<br />
<span>CO-RE — Compile Once, Run Everywhere — is the eBPF feature that throws all of that out. The idea, in one paragraph: when you write a BPF program that reads <span class='inlinecode'>task-&gt;mm-&gt;start_stack</span>, you don&#39;t bake the offsets of those fields into the compiled program. Instead, the compiler emits relocation records ("at this instruction, fetch the offset of <span class='inlinecode'>mm</span> inside <span class='inlinecode'>task_struct</span>"). At load time, libbpf looks up the actual offsets in the target kernel&#39;s BTF (BPF Type Format — a description of every kernel struct, embedded in <span class='inlinecode'>/sys/kernel/btf/vmlinux</span> on any modern kernel) and patches the program in place. The same <span class='inlinecode'>.bpf.o</span> that ran on a 5.10 Debian kernel runs on a 6.8 Fedora kernel without recompilation.</span><br />
<br />
<span>Pictorially, the contrast looks like this:</span><br />
<br />
<pre>
Old I/O Riot (Systemtap)                 New ior (libbpf + CO-RE)
─────────────────────────                ────────────────────────────
  .stp source                              .bpf.c source
       │                                        │
       │ needs THIS kernel&#39;s headers            │ build ONCE against vmlinux.h
       │ + debuginfo package installed          │ (generated from any kernel BTF)
       ▼                                        ▼
  per-host translate + compile              one portable .bpf.o
       │                                        │
       ▼                                        ▼
  per-host kernel module                    same binary on every host
       │                                        │
  insmod / modprobe                         libbpf loader:
       │                                        │  • read /sys/kernel/btf/vmlinux
       ▼                                        │  • patch field offsets
  attached, this kernel only                 │  • verify + load
                                              ▼
                                          attached, runs anywhere
</pre>
<br />
<span>What that buys ior in practice: I ship a single <span class='inlinecode'>ior</span> binary. On any Linux ≥4.18-ish with BTF available (which is almost all of them now — Debian, Ubuntu, Fedora, Arch, RHEL all ship <span class='inlinecode'>CONFIG_DEBUG_INFO_BTF=y</span> by default), it just works. No kernel-debuginfo dependency, no per-kernel build matrix, no DKMS hooks. The first time I tried <span class='inlinecode'>scp ior fedora-box:</span> and it ran without complaint after a 6-month gap I had to double-check it wasn&#39;t silently doing nothing.</span><br />
<br />
<span>The runtime shape of a trace pipeline lines up with that:</span><br />
<br />
<pre>
 kernel side                                userspace (this binary)
 ───────────                                ───────────────────────
                                            ┌──────────────────────┐
   tracepoint:                              │  Go process          │
   sys_enter_openat                         │  ┌────────────────┐  │
        │                                   │  │  aggregator    │  │
        ▼                                   │  │  (latency,     │  │
   ┌─────────┐                              │  │   stacks,      │  │
   │ BPF prog│ ─── perf ring buf ──────────&gt;│──│   filters)     │  │
   │ (verified                              │  └─────┬──────────┘  │
   │  bytecode)                             │        │             │
   └─────────┘                              │        ▼             │
                                            │  Bubble Tea TUI /    │
                                            │  parquet writer /    │
                                            │  CSV stdout          │
                                            └──────────────────────┘
</pre>
<br />
<span>The cost is cgo. Every call from Go into libbpf crosses the cgo boundary, which historically meant tens to ~hundred-ish nanoseconds of overhead per call — register save/restore, a stack switch onto g0, goroutine state bookkeeping. Cheap in absolute terms, but it adds up if you call into C inside a tight loop. ior keeps the actual hot path on the kernel side and only crosses into Go once per drained batch of events from the ring buffer, so the per-call cost is amortized over thousands of events. In practice it doesn&#39;t show up in profiles.</span><br />
<br />
<span>Go 1.26, the current release at the time of writing (late April 2026), is the one that finally took a serious bite out of cgo&#39;s per-call cost — the runtime can elide a chunk of the bookkeeping for calls that don&#39;t need it. Real-world wins depend heavily on the workload, but the rough direction is that cgo now feels closer to "an unusually expensive function call" than to "a context switch", which is the right mental model for almost everyone touching a C library from Go. The shorter version: cgo overhead used to be a real footgun for ports that called into C in the inner loop. With Go 1.26 it&#39;s a footnote unless you&#39;re doing many millions of small calls per second, in which case batching across the boundary still fixes it.</span><br />
<br />
<h3 style='display: inline' id='if-you-want-to-go-deeper'>If you want to go deeper</h3><br />
<br />
<span>If any of this sounds interesting and you want to learn how to write your own BPF programs, two books are the standard recommendations and both well worth the time:</span><br />
<br />
<ul>
<li>"Learning eBPF" by Liz Rice (O&#39;Reilly, 2023) is the friendlier on-ramp. It walks through writing your first programs end-to-end, covers CO-RE and BTF in plain English, and is the book I&#39;d hand to someone who has never touched the kernel side before. Liz also gave the canonical "what is eBPF" conference talk floating around YouTube, which makes a good 40-minute companion.</li>
<li>"BPF Performance Tools: Linux System and Application Observability" by Brendan Gregg (Addison-Wesley, 2019) is the encyclopedia. It&#39;s where you go after you&#39;ve understood the basics and now want a complete reference for tracing every subsystem in the kernel — file systems, networking, scheduler, languages, applications — with worked tools for each. The flame-graph-driven analysis style throughout is also exactly how ior&#39;s own flamegraph tab thinks about a workload.</li>
</ul><br />
<span>Between the two, Rice teaches you the moving parts and Gregg teaches you what to do with them.</span><br />
<br />
<h2 style='display: inline' id='the-whole-thing-as-a-tape-pipeline'>The whole thing as a tape pipeline</h2><br />
<br />
<span>The demo isn&#39;t a screencast I sat through. It&#39;s 14 VHS tapes that drive the TUI deterministically, with a background workload generator producing real syscall traffic for the trace to chew on. One <span class='inlinecode'>mage demo</span> and every GIF below regenerates from scratch. The boring part of "make a demo" — having to re-record everything when the UI shifts — goes away.</span><br />
<br />
<h2 style='display: inline' id='first-launch'>First launch</h2><br />
<br />
<!-- Generator: GNU source-highlight 3.1.9
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
<pre>sudo ./ior
</pre>
<br />
<span>You land on the PID picker. The default selection is "All PIDs", so Enter just dumps you straight at the dashboard.</span><br />
<br />
<a href='./unveiling-ior-ng/01-launch.gif'><img alt='Cold start: PID picker, then the dashboard' title='Cold start: PID picker, then the dashboard' src='./unveiling-ior-ng/01-launch.gif' /></a><br />
<br />
<span>The dashboard opens on the live flamegraph. Bars grow as new events arrive. Before walking through the keys, a paragraph on what you&#39;re looking at — flamegraphs are easier to read than they are to describe.</span><br />
<br />
<span>A flamegraph is a histogram of stacks. Each horizontal bar is one entry in a stack; every bar directly above it is a child of that entry, and the stack you read top-to-bottom is the same shape as a call chain. In ior, "stack" doesn&#39;t mean function-call stack (we don&#39;t have userspace symbols). It means a tuple of dimensions of the trace: by default <span class='inlinecode'>comm/path/tracepoint</span>, so the bottom row is per-process names, the middle row is per-file paths, and the top row is the syscall (<span class='inlinecode'>enter_read</span>, <span class='inlinecode'>enter_openat</span>, etc.). A wide bar means lots of events landed in that bucket, a narrow bar means few. There is no time axis — left-to-right is just sort order, not chronology. The whole chart is one "where is the I/O coming from?" picture.</span><br />
<br />
<span>One thing worth flagging because it&#39;s the unusual bit: this flamegraph is live. Most of the flamegraph tooling out there — Brendan Gregg&#39;s <span class='inlinecode'>flamegraph.pl</span>, all the <span class='inlinecode'>perf script | stackcollapse-* | flamegraph.pl</span> pipelines, every <span class='inlinecode'>pprof -web</span> invocation — produces a static SVG: capture a profile for N seconds, render once, browse the result. ior&#39;s tab is not that. Bars grow, shrink, appear, and disappear in real time as events stream in from the kernel — at full screen-refresh rate while the workload runs, with no pause. You can sit on this tab while you change something on the system (start a build, cycle a service, run a query) and watch the I/O shape mutate underneath you. That&#39;s a different mental model from the static "I have a profile, let me look at it" workflow most people are used to, and it&#39;s what makes the tab actually useful as an at-a-glance diagnostic surface rather than a post-mortem artifact.</span><br />
<br />
<span>Because it&#39;s live, there&#39;s also a way to throw away the accumulated history and start the rolling count from "now": <span class='inlinecode'>r</span> resets the baseline. Everything the flamegraph has been counting since launch (or since the last reset) is dropped, and from that moment the chart reflects only events that arrived after the reset. Useful for the "compare before vs after" workflow — change one thing on the box, hit <span class='inlinecode'>r</span> immediately, and the next thirty seconds of accumulation is a fresh picture of the new state.</span><br />
<br />
<span>That visualisation buys you two things you can&#39;t easily get from a tabular view. First, hierarchy: it&#39;s immediately obvious whether one process is responsible for ten thousand reads on a single file, or ten thousand reads spread across a hundred files — the first looks like one tall pillar, the second looks like a wide ridge. Second, scale: bar width is proportional to the metric (count or bytes), so a process that did 95% of the work towers over the others. The eye picks that up instantly; the same fact in a sorted table requires reading numbers and doing the ratio mentally.</span><br />
<br />
<span>Useful workflows you can do entirely from this tab:</span><br />
<br />
<ul>
<li>"What&#39;s pounding the disk?" — leave it on default order (<span class='inlinecode'>comm/path/tracepoint</span>) and watch which <span class='inlinecode'>comm</span> widens. Press <span class='inlinecode'>b</span> once to switch the metric to bytes if you care about throughput, not call count.</li>
<li>"Why is this one process slow?" — <span class='inlinecode'>l</span> (or <span class='inlinecode'>→</span>) until the cursor is on that process, then <span class='inlinecode'>enter</span> to zoom. The whole chart re-roots there and you see only that process&#39;s paths and syscalls.</li>
<li>"What&#39;s in /var/lib/X?" — press <span class='inlinecode'>o</span> once to flip ordering to <span class='inlinecode'>path/tracepoint/comm</span>, navigate to the path, zoom. Now the children show which syscalls hit it and which processes did them.</li>
<li>"Did the new deploy change the I/O shape?" — press <span class='inlinecode'>r</span> to reset the baseline, wait a bit, and the chart starts fresh with only events from the reset point onward. Pair the same syscall surface "before" vs "after" and the difference jumps out by shape.</li>
</ul><br />
<span>Now the keys. Movement uses vi-style <span class='inlinecode'>h</span>/<span class='inlinecode'>j</span>/<span class='inlinecode'>k</span>/<span class='inlinecode'>l</span> everywhere in ior — and the cursor keys work too if you&#39;d rather. <span class='inlinecode'>h</span>/<span class='inlinecode'>l</span> (or <span class='inlinecode'>←</span>/<span class='inlinecode'>→</span>) walk siblings at the current depth, <span class='inlinecode'>j</span>/<span class='inlinecode'>k</span> (or <span class='inlinecode'>↓</span>/<span class='inlinecode'>↑</span>) step shallower or deeper. <span class='inlinecode'>enter</span> zooms into the selected subtree (the rest of the chart greys out and the selection becomes the new root). <span class='inlinecode'>u</span> or <span class='inlinecode'>Esc</span> undoes the zoom. <span class='inlinecode'>b</span> toggles the metric driving bar width between event count and total bytes. <span class='inlinecode'>/</span> opens regex search; matching frames stay coloured while everything else greys out, so you can use it as a filter as well as a finder. And <span class='inlinecode'>o</span> cycles between five different stack-ordering modes, each with its own lens on the data.</span><br />
<br />
<span>The five orderings ship as built-in presets. The leftmost dimension is the bottom row of the chart (the root); the rightmost is the top row (the leaf). Pressing <span class='inlinecode'>o</span> rotates through them in this order:</span><br />
<br />
<ul>
<li><span class='inlinecode'>comm/tracepoint/path</span> — the default. Root rows are processes (by command name); each process&#39;s bar splits into the syscalls it issued, and each syscall splits further by file path. Best general-purpose view: "which programs are doing the I/O, and what kind?"</li>
<li><span class='inlinecode'>path/tracepoint/comm</span> — root by file path. Use this when you suspect a particular file or directory is the bottleneck — pick the path, see which syscalls hit it, and which processes did those syscalls. Pairs naturally with directory grouping in the Files tab.</li>
<li><span class='inlinecode'>tracepoint/comm/path</span> — root by syscall. When you already know "this is an <span class='inlinecode'>openat</span> problem" or "we&#39;re write-bound", this view collects all the openat (or write) traffic at the bottom and lets you drill into who&#39;s doing it and to which paths.</li>
<li><span class='inlinecode'>pid/tracepoint/path</span> — root by PID, not comm. Same shape as the default but each individual process gets its own bar instead of being lumped in with siblings sharing a comm. Useful when you have many bash or python instances and need to tell them apart.</li>
<li><span class='inlinecode'>comm/path/tracepoint</span> — root by process, then by file (skipping the syscall layer at the top). Best when you care about "what files does this program touch?" more than "what syscalls does it issue?" — the file column gets a full row of vertical real estate instead of being split per-syscall.</li>
</ul><br />
<span>In all five orderings, bar widths still mean the same thing — proportion of the active metric (events or bytes, toggled with <span class='inlinecode'>b</span>). The toolbar at the top of the chart always shows the current ordering as <span class='inlinecode'>o:order(&lt;dim1&gt;/&lt;dim2&gt;/&lt;dim3&gt;)</span>, so you never lose track of which lens you&#39;re looking through.</span><br />
<br />
<a href='./unveiling-ior-ng/13-tui-flamegraph.gif'><img alt='Live in-TUI flamegraph: navigate, zoom, undo, cycle order + metric' title='Live in-TUI flamegraph: navigate, zoom, undo, cycle order + metric' src='./unveiling-ior-ng/13-tui-flamegraph.gif' /></a><br />
<br />
<h2 style='display: inline' id='the-seven-tabs-in-30-seconds-each'>The seven tabs, in 30 seconds each</h2><br />
<br />
<span>The number keys jump between tabs. <span class='inlinecode'>tab</span> and <span class='inlinecode'>shift+tab</span> step.</span><br />
<br />
<h3 style='display: inline' id='2-overview'><span class='inlinecode'>2</span> Overview</h3><br />
<br />
<span>A sparkline plus the top syscalls and top paths — the at-a-glance view, useful as a "what&#39;s happening right now?" landing tab when you don&#39;t yet know what you&#39;re looking for.</span><br />
<br />
<a href='./unveiling-ior-ng/02-overview-tab.gif'><img alt='Overview tab' title='Overview tab' src='./unveiling-ior-ng/02-overview-tab.gif' /></a><br />
<br />
<h3 style='display: inline' id='3-syscalls'><span class='inlinecode'>3</span> Syscalls</h3><br />
<br />
<span>A sortable table of every syscall ior knows about, with rate, average latency, p95/p99, total bytes, and error count. <span class='inlinecode'>s</span> sorts by the selected column, <span class='inlinecode'>S</span> reverses. The most useful column when something&#39;s wrong is usually p99 — it&#39;s where you see the long-tail outlier syscall types.</span><br />
<br />
<a href='./unveiling-ior-ng/03-syscalls-tab.gif'><img alt='Syscalls table with sort + reverse-sort' title='Syscalls table with sort + reverse-sort' src='./unveiling-ior-ng/03-syscalls-tab.gif' /></a><br />
<br />
<h3 style='display: inline' id='4-files'><span class='inlinecode'>4</span> Files</h3><br />
<br />
<span>Same shape as Syscalls but rows are file paths. The interesting key here is <span class='inlinecode'>d</span>: it rolls per-file rows up into their parent directory. Essential when you&#39;ve got a process touching ten thousand files in <span class='inlinecode'>/usr/share/</span> — without it the table is unreadable noise.</span><br />
<br />
<a href='./unveiling-ior-ng/04-files-tab.gif'><img alt='Directory grouping toggle' title='Directory grouping toggle' src='./unveiling-ior-ng/04-files-tab.gif' /></a><br />
<br />
<h3 style='display: inline' id='5-processes'><span class='inlinecode'>5</span> Processes</h3><br />
<br />
<span>Same shape again, but rows are processes / comms. Best paired with the Stream tab — once you spot a culprit comm here, push it to the global filter with <span class='inlinecode'>Enter</span> and the rest of the dashboard is scoped to that process.</span><br />
<br />
<a href='./unveiling-ior-ng/05-processes-tab.gif'><img alt='Processes tab' title='Processes tab' src='./unveiling-ior-ng/05-processes-tab.gif' /></a><br />
<br />
<h3 style='display: inline' id='6-latency--gaps'><span class='inlinecode'>6</span> Latency + Gaps</h3><br />
<br />
<span>Two histograms side by side: how long each syscall took (latency), and the wall-clock interval between syscalls on the same thread (gap). Latency tells you "is the kernel slow"; gap tells you "what is the program doing between two kernel calls".</span><br />
<br />
<span>A subtle but important point about that gap: ior measures it from the exit of one syscall to the entry of the next on the same TID, but it does not know what the thread was doing in the meantime. A long gap doesn&#39;t mean the thread was idle — it might have been pinned on a CPU running pure userspace code (number-crunching, JSON parsing, GC, a busy loop). All "gap" tells you for sure is "this thread didn&#39;t call into the kernel for X microseconds." Whether that&#39;s because it was sleeping, blocked on a condition variable, computing, or scheduled out is something only the gap value alone cannot answer — pair it with <span class='inlinecode'>top</span>/<span class='inlinecode'>perf top</span> if you need to disambiguate. In practice this is still extremely useful: a syscall-driven workload with surprisingly long gaps is a strong hint that you&#39;re CPU-bound somewhere outside the kernel, and that&#39;s a different optimisation conversation than slow I/O.</span><br />
<br />
<span>The dd loop in the demo workload spreads the latency distribution out so you can actually see the shape.</span><br />
<br />
<a href='./unveiling-ior-ng/06-latency-gaps-tab.gif'><img alt='Latency + gap histograms' title='Latency + gap histograms' src='./unveiling-ior-ng/06-latency-gaps-tab.gif' /></a><br />
<br />
<h3 style='display: inline' id='7-stream'><span class='inlinecode'>7</span> Stream</h3><br />
<br />
<span>The live tail — every event as it happens, in a row-per-event ring buffer. This is where you spend most of your time when something&#39;s actually broken. The whole next section is about it.</span><br />
<br />
<a href='./unveiling-ior-ng/07-stream-live.gif'><img alt='Stream tab live-tailing' title='Stream tab live-tailing' src='./unveiling-ior-ng/07-stream-live.gif' /></a><br />
<br />
<h2 style='display: inline' id='the-stream-tab-is-the-good-one'>The Stream tab is the good one</h2><br />
<br />
<span><span class='inlinecode'>space</span> pauses. In pause mode, the same vi-style <span class='inlinecode'>h</span>/<span class='inlinecode'>j</span>/<span class='inlinecode'>k</span>/<span class='inlinecode'>l</span> (or arrow keys) move the row/column cursor across the table. Hitting <span class='inlinecode'>Enter</span> on a cell pushes a new filter onto a stack, narrowing what you see. Pile them up — comm, then syscall, then file — and <span class='inlinecode'>Esc</span> pops them off LIFO when you want to back out.</span><br />
<br />
<a href='./unveiling-ior-ng/08-stream-pause-filter.gif'><img alt='Pause, push two filters, undo with Esc' title='Pause, push two filters, undo with Esc' src='./unveiling-ior-ng/08-stream-pause-filter.gif' /></a><br />
<br />
<span><span class='inlinecode'>/</span> and <span class='inlinecode'>?</span> are regex search forward/backward. <span class='inlinecode'>n</span> and <span class='inlinecode'>N</span> walk matches. The search runs against every column in the ring buffer and wraps at the end. Search and filtering are different beasts: search highlights and jumps, filtering hides everything that doesn&#39;t match.</span><br />
<br />
<a href='./unveiling-ior-ng/09-stream-regex-search.gif'><img alt='Regex search' title='Regex search' src='./unveiling-ior-ng/09-stream-regex-search.gif' /></a><br />
<br />
<span><span class='inlinecode'>e</span> exports the current filtered snapshot to a CSV in the working directory. <span class='inlinecode'>x</span> does the same for the paused stream view specifically (preserving your filter stack), <span class='inlinecode'>X</span> prompts for a filename, <span class='inlinecode'>E</span> opens the most recent export in <span class='inlinecode'>$EDITOR</span>.</span><br />
<br />
<a href='./unveiling-ior-ng/10-stream-csv-export.gif'><img alt='CSV export' title='CSV export' src='./unveiling-ior-ng/10-stream-csv-export.gif' /></a><br />
<br />
<h2 style='display: inline' id='filtering-more-thoroughly'>Filtering, more thoroughly</h2><br />
<br />
<span>The Enter-to-push trick isn&#39;t unique to Stream. It works the same on Files, Syscalls, and Processes: highlight a row, hit Enter, and the cell value becomes a filter against the entire dashboard. Three tabs of "I see one weird path / comm / syscall, drill in" with one keystroke.</span><br />
<br />
<span>The filter status line gives you a one-glance summary of every active frame, written like:</span><br />
<br />
<ul>
<li><span class='inlinecode'>comm~bash</span> — substring match on a string column. This is what Enter-on-a-cell produces for <span class='inlinecode'>comm</span>, <span class='inlinecode'>syscall</span>, and <span class='inlinecode'>file</span>.</li>
<li><span class='inlinecode'>pid=1234</span> — exact equality. Used for <span class='inlinecode'>pid</span>, <span class='inlinecode'>tid</span>, <span class='inlinecode'>fd</span>, <span class='inlinecode'>ret</span>, <span class='inlinecode'>bytes</span>.</li>
<li><span class='inlinecode'>latency&gt;=5ms</span> / <span class='inlinecode'>gap&gt;=10us</span> — numeric comparison with a duration suffix. The full operator set is <span class='inlinecode'>&gt;</span>, <span class='inlinecode'>&lt;</span>, <span class='inlinecode'>=</span>, <span class='inlinecode'>&gt;=</span>, <span class='inlinecode'>&lt;=</span>, <span class='inlinecode'>!=</span>.</li>
</ul><br />
<span>Stack frames AND together, so pushing <span class='inlinecode'>comm~bash</span> and then <span class='inlinecode'>syscall~openat</span> shows you bash&#39;s openat calls, not bash OR openat.</span><br />
<br />
<span>Undoing is symmetric to pushing: <span class='inlinecode'>Esc</span> pops the most recent frame off the stack — one keystroke per layer, LIFO. Press it once to drop the <span class='inlinecode'>syscall~openat</span> filter and you&#39;re back to bash-only; press it again and the <span class='inlinecode'>comm~bash</span> filter goes too, leaving the unfiltered firehose. To clear the whole stack at once, just hold <span class='inlinecode'>Esc</span> until the status line reads <span class='inlinecode'>filter: all</span>. The <span class='inlinecode'>F</span> key is a synonym for <span class='inlinecode'>Esc</span> here and works from any tab — handy from Files/Syscalls/Processes where <span class='inlinecode'>Esc</span> might otherwise close a modal first.</span><br />
<br />
<span>Two other knobs do related work:</span><br />
<br />
<ul>
<li><span class='inlinecode'>p</span>, <span class='inlinecode'>t</span>, <span class='inlinecode'>o</span> open the PID, TID, and probe-toggle dialogs. These are global filters: they reconfigure the BPF side, so kernel-level events for excluded PIDs/probes never even reach userspace. Cheaper than filtering a firehose, but it also means the filter applies to recordings (the parquet file only contains rows the kernel let through).</li>
<li>The CLI mirrors of those dialogs let you bake the same scoping into a one-shot run: <span class='inlinecode'>-pid</span>, <span class='inlinecode'>-tid</span>, <span class='inlinecode'>-comm</span>, <span class='inlinecode'>-path</span>, plus <span class='inlinecode'>-tps &lt;regex&gt;</span> / <span class='inlinecode'>-tpsExclude &lt;regex&gt;</span> for picking which tracepoints to attach in the first place.</li>
</ul><br />
<a href='./unveiling-ior-ng/11-pid-tid-probe.gif'><img alt='PID, TID, and probe pickers' title='PID, TID, and probe pickers' src='./unveiling-ior-ng/11-pid-tid-probe.gif' /></a><br />
<br />
<h2 style='display: inline' id='recording'>Recording</h2><br />
<br />
<span>Three persistence flows, each for a different job:</span><br />
<br />
<ul>
<li><span class='inlinecode'>R</span> from the dashboard starts streaming Parquet — every event row that survives your current TUI filter goes to disk continuously. <span class='inlinecode'>R</span> again stops. Footer shows the active file or the last error.</li>
</ul><br />
<a href='./unveiling-ior-ng/12-parquet-recording.gif'><img alt='Parquet recording from the TUI' title='Parquet recording from the TUI' src='./unveiling-ior-ng/12-parquet-recording.gif' /></a><br />
<br />
<ul>
<li><span class='inlinecode'>sudo ./ior -flamegraph -name &lt;n&gt;</span> writes one aggregated <span class='inlinecode'>.ior.zst</span> artifact at shutdown. Aggregated counters, not per-event rows. Cheaper to write, ideal for ior&#39;s native flamegraph workflow and integration tests.</li>
</ul><br />
<ul>
<li><span class='inlinecode'>sudo ./ior -parquet trace.parquet</span> is the headless firehose — every row, no TUI, no filtering. <span class='inlinecode'>sudo ./ior -plain</span> is even lighter: CSV to stdout, pipe it into anything.</li>
</ul><br />
<a href='./unveiling-ior-ng/14-headless-modes.gif'><img alt='All three headless flows in one tape' title='All three headless flows in one tape' src='./unveiling-ior-ng/14-headless-modes.gif' /></a><br />
<br />
<h2 style='display: inline' id='querying-a-parquet-trace-with-clickhouse'>Querying a parquet trace with ClickHouse</h2><br />
<br />
<span>The schema is flat and stable: <span class='inlinecode'>seq, time_ns, gap_ns, latency_ns, comm, pid, tid, syscall, fd, ret, bytes, file, is_error, filter_epoch</span>. ClickHouse Local reads parquet directly without a server, which makes it a perfect post-mortem tool — point it at the file and run SQL:</span><br />
<br />
<!-- Generator: GNU source-highlight 3.1.9
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
<pre>clickhouse <b><u><font color="#000000">local</font></u></b> --query <font color="#808080">"</font>
<font color="#808080">  SELECT comm, syscall, count() AS n,</font>
<font color="#808080">         formatReadableSize(sum(bytes)) AS total</font>
<font color="#808080">  FROM file('trace.parquet', Parquet)</font>
<font color="#808080">  GROUP BY comm, syscall</font>
<font color="#808080">  ORDER BY n DESC</font>
<font color="#808080">  LIMIT 10</font>
<font color="#808080">"</font> --format PrettyCompactNoEscapes
</pre>
<br />
<pre>
    ┌─comm────────────┬─syscall─┬─────n─┬─total──────┐
 1. │ notify-rs inoti │ read    │ 42005 │ 732.31 KiB │
 2. │ cosmic-term     │ statx   │ 10898 │ 0.00 B     │
 3. │ cosmic-term     │ read    │ 10103 │ 4.02 MiB   │
 4. │ surface-eDP-1   │ ioctl   │  8452 │ 0.00 B     │
 5. │ cosmic-term     │ close   │  4918 │ 0.00 B     │
 6. │ cosmic-term     │ openat  │  4537 │ 0.00 B     │
 7. │ cosmic-term     │ ioctl   │  3556 │ 0.00 B     │
 8. │ tokio-runtime-w │ read    │  1976 │ 4.04 MiB   │
 9. │ cosmic-comp     │ read    │  1118 │ 6.63 KiB   │
10. │ systemd-oomd    │ read    │  1085 │ 111.97 KiB │
    └─────────────────┴─────────┴───────┴────────────┘
</pre>
<br />
<span>The fields you actually want for performance work are <span class='inlinecode'>latency_ns</span> and <span class='inlinecode'>gap_ns</span>. P99 by syscall, only the ones that landed in error:</span><br />
<br />
<!-- Generator: GNU source-highlight 3.1.9
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
<pre>clickhouse <b><u><font color="#000000">local</font></u></b> --query <font color="#808080">"</font>
<font color="#808080">  SELECT syscall, count() AS n,</font>
<font color="#808080">         round(quantile(0.5)(latency_ns)/1000,  1) AS p50_us,</font>
<font color="#808080">         round(quantile(0.99)(latency_ns)/1000, 1) AS p99_us</font>
<font color="#808080">  FROM file('trace.parquet', Parquet)</font>
<font color="#808080">  WHERE is_error = 1</font>
<font color="#808080">  GROUP BY syscall</font>
<font color="#808080">  ORDER BY p99_us DESC</font>
<font color="#808080">"</font> --format PrettyCompactNoEscapes
</pre>
<br />
<pre>
    ┌─syscall────┬─────n─┬─p50_us─┬─p99_us─┐
 1. │ statx      │  1216 │    2.2 │   16.4 │
 2. │ newfstatat │    69 │    1.7 │   16.4 │
 3. │ open       │     1 │   16.1 │   16.1 │
 4. │ mkdir      │   306 │    3.9 │   11.7 │
 5. │ readlink   │    11 │    1.5 │   10.4 │
 6. │ newstat    │    44 │    2.5 │    8.4 │
 7. │ unlinkat   │   347 │      1 │    6.2 │
 8. │ openat     │   380 │    2.1 │    5.8 │
 9. │ access     │     2 │      5 │    5.5 │
10. │ read       │ 23597 │    0.5 │    5.4 │
11. │ ioctl      │   901 │      1 │    5.3 │
12. │ writev     │     1 │    0.7 │    0.7 │
    └────────────┴───────┴────────┴────────┘
</pre>
<br />
<span>Real output, by the way — those rows are from a 30-second <span class='inlinecode'>ior -parquet trace.parquet</span> capture on the laptop I&#39;m typing this on. <span class='inlinecode'>notify-rs inoti…</span> is the inotify thread of some Rust app I had open; <span class='inlinecode'>cosmic-term</span> is the COSMIC desktop&#39;s terminal emulator. The slowest p99 errors are the directory-walking syscalls (statx, newfstatat, mkdir) at ~16 µs — bog standard.</span><br />
<br />
<span>Same trick works in DuckDB (<span class='inlinecode'>duckdb -c "SELECT ... FROM &#39;trace.parquet&#39;"</span>), pandas, polars, anything that reads Parquet. The point of streaming Parquet rather than ior&#39;s native <span class='inlinecode'>.ior.zst</span> format is exactly this: once it&#39;s on disk, you&#39;re in the standard data-tools ecosystem.</span><br />
<br />
<h2 style='display: inline' id='reproducing-the-whole-demo'>Reproducing the whole demo</h2><br />
<br />
<!-- Generator: GNU source-highlight 3.1.9
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
<pre>mage installDemoTools     <i><font color="silver"># one-time: VHS via go install + ttyd from dnf</font></i>
sudo -v                   <i><font color="silver"># warm the sudo timestamp once</font></i>
mage demo                 <i><font color="silver"># ~10 minutes, fully headless, safe to background</font></i>
</pre>
<br />
<span>To rebuild a single GIF after editing its tape: <span class='inlinecode'>TAPE=07-stream-live mage demoOne</span>.</span><br />
<br />
<h2 style='display: inline' id='what-s-still-missing'>What&#39;s still missing</h2><br />
<br />
<span>ior is pre-alpha and basically a personal tool. The headline gaps:</span><br />
<br />
<ul>
<li>No record/replay — that was the whole point of the original I/O Riot. The new one is a tracer, not a workload simulator. I keep going back and forth on whether to put replay back in.</li>
<li>No userspace symbol resolution. Stacks are at the syscall surface, not "which line of which library called read".</li>
<li>No remote / cluster mode. Single host, one trace at a time.</li>
</ul><br />
<span>But the live flamegraph, the stackable stream filters, and the cheap parquet capture together cover the cases I actually hit week to week. The demo above is the easiest way to get a feel for whether it&#39;s the kind of tool you want.</span><br />
<br />
<a class='textlink' href='https://codeberg.org/snonux/ior'>Source on Codeberg</a><br />
<a class='textlink' href='https://codeberg.org/snonux/ior/src/branch/main/demo/TUTORIAL.md'>The full in-repo tutorial</a><br />
<p class="footer">
	Generated with <a href="https://codeberg.org/snonux/gemtexter">Gemtexter 3.0.1-develop</a> |
	served by <a href="https://www.OpenBSD.org">OpenBSD</a>/<a href="https://man.ope