summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorPaul Buetow <paul@buetow.org>2025-01-15 00:17:16 +0200
committerPaul Buetow <paul@buetow.org>2025-01-15 00:17:16 +0200
commit308208fbc37f44b621c1d9b94190cf092f732aa2 (patch)
treeac469587b6baf266f0521e02af10fafe3dbcff20
parentdf3f7d0ff779c353422e8b6d76a64ba5ea3d36bd (diff)
Update content for gemtext
-rw-r--r--about/resources.gmi178
-rw-r--r--gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.gmi3
-rw-r--r--gemfeed/2025-01-15-working-with-an-sre-interview.gmi177
-rw-r--r--gemfeed/2025-01-15-working-with-an-sre-interview.gmi.tpl164
-rw-r--r--gemfeed/DRAFT-f3s-kubernetes-with-freebsd-bhyve.gmi91
-rw-r--r--gemfeed/atom.xml568
-rw-r--r--gemfeed/index.gmi1
-rw-r--r--index.gmi3
-rw-r--r--uptime-stats.gmi2
9 files changed, 709 insertions, 478 deletions
diff --git a/about/resources.gmi b/about/resources.gmi
index 327c51ea..f8d86f8e 100644
--- a/about/resources.gmi
+++ b/about/resources.gmi
@@ -35,100 +35,100 @@ You won't find any links on this site because, over time, the links will break.
In random order:
-* Object-Oriented Programming with ANSI-C; Axel-Tobias Schreiner
-* Modern Perl; Chromatic ; Onyx Neon Press
-* Higher Order Perl; Mark Dominus; Morgan Kaufmann
-* Data Science at the Command Line; Jeroen Janssens; O'Reilly
-* DevOps And Site Reliability Engineering Handbook; Stephen Fleming; Audible
-* Distributed Systems: Principles and Paradigms; Andrew S. Tanenbaum; Pearson
-* 97 things every SRE should know; Emil Stolarsky, Jaime Woo; O'Reilly
-* Developing Games in Java; David Brackeen and others...; New Riders
-* Learn You Some Erlang for Great Good; Fred Herbert; No Starch Press
-* The Go Programming Language; Alan A. A. Donovan; Addison-Wesley Professional
-* Learn You a Haskell for Great Good!; Miran Lipovaca; No Starch Press
-* Concurrency in Go; Katherine Cox-Buday; O'Reilly
+* Polished Ruby Programming; Jeremy Evans; Packt Publishing
* Site Reliability Engineering; How Google runs production systems; O'Reilly
-* The Pragmatic Programmer; David Thomas; Addison-Wesley
* 100 Go Mistakes and How to Avoid Them; Teiva Harsanyi; Manning Publications
-* Effective Java; Joshua Bloch; Addison-Wesley Professional
-* Leanring eBPF; Liz Rice; O'Reilly
-* Amazon Web Services in Action; Michael Wittig and Andreas Wittig; Manning Publications
-* Effective awk programming; Arnold Robbins; O'Reilly
-* Pro Puppet; James Turnbull, Jeffrey McCune; Apress
-* Java ist auch eine Insel; Christian Ullenboom;
-* C++ Programming Language; Bjarne Stroustrup;
+* Terraform Cookbook; Mikael Krief; Packt Publishing
+* 97 things every SRE should know; Emil Stolarsky, Jaime Woo; O'Reilly
* Tmux 2: Productive Mouse-free Development; Brain P. Hogan; The Pragmatic Programmers
-* Perl New Features; Joshua McAdams, brian d foy; Perl School
+* Learn You a Haskell for Great Good!; Miran Lipovaca; No Starch Press
+* Java ist auch eine Insel; Christian Ullenboom;
+* 21st Century C: C Tips from the New School; Ben Klemens; O'Reilly
* The Docker Book; James Turnbull; Kindle
-* Polished Ruby Programming; Jeremy Evans; Packt Publishing
-* The Kubernetes Book; Nigel Poulton; Unabridged Audiobook
-* Terraform Cookbook; Mikael Krief; Packt Publishing
-* Systemprogrammierung in Go; Frank Müller; dpunkt
-* Programming Perl aka "The Camel Book"; Tom Christiansen, brian d foy, Larry Wall & Jon Orwant; O'Reilly
+* Clusterbau mit Linux-HA; Michael Schwartzkopff; O'Reilly
+* Raku Recipes; J.J. Merelo; Apress
+* Object-Oriented Programming with ANSI-C; Axel-Tobias Schreiner
+* Funktionale Programmierung; Peter Pepper; Springer
* The Practise of System and Network Administration; Thomas A. Limoncelli, Christina J. Hogan, Strata R. Chalup; Addison-Wesley Professional Pro Git; Scott Chacon, Ben Straub; Apress
-* DNS and BIND; Cricket Liu; O'Reilly
-* Go Brain Teasers - Exercise Your Mind; Miki Tebeka; The Pragmatic Programmers
+* The Go Programming Language; Alan A. A. Donovan; Addison-Wesley Professional
+* Learn You Some Erlang for Great Good; Fred Herbert; No Starch Press
+* C++ Programming Language; Bjarne Stroustrup;
+* Systemprogrammierung in Go; Frank Müller; dpunkt
+* The Pragmatic Programmer; David Thomas; Addison-Wesley
+* Effective Java; Joshua Bloch; Addison-Wesley Professional
+* DevOps And Site Reliability Engineering Handbook; Stephen Fleming; Audible
* Hands-on Infrastructure Monitoring with Prometheus; Joel Bastos, Pedro Araujo; Packt
-* Raku Recipes; J.J. Merelo; Apress
+* Go Brain Teasers - Exercise Your Mind; Miki Tebeka; The Pragmatic Programmers
* The KCNA (Kubernetes and Cloud Native Associate) Book; Nigel Poulton
-* Clusterbau mit Linux-HA; Michael Schwartzkopff; O'Reilly
-* 21st Century C: C Tips from the New School; Ben Klemens; O'Reilly
-* Systems Performance Tuning; Gian-Paolo D. Musumeci and others...; O'Reilly
+* Developing Games in Java; David Brackeen and others...; New Riders
+* Data Science at the Command Line; Jeroen Janssens; O'Reilly
+* Pro Puppet; James Turnbull, Jeffrey McCune; Apress
+* Programming Perl aka "The Camel Book"; Tom Christiansen, brian d foy, Larry Wall & Jon Orwant; O'Reilly
+* Effective awk programming; Arnold Robbins; O'Reilly
+* DNS and BIND; Cricket Liu; O'Reilly
+* Leanring eBPF; Liz Rice; O'Reilly
+* Higher Order Perl; Mark Dominus; Morgan Kaufmann
+* Concurrency in Go; Katherine Cox-Buday; O'Reilly
* Raku Fundamentals; Moritz Lenz; Apress
-* Funktionale Programmierung; Peter Pepper; Springer
+* Distributed Systems: Principles and Paradigms; Andrew S. Tanenbaum; Pearson
+* Think Raku (aka Think Perl 6); Laurent Rosenfeld, Allen B. Downey; O'Reilly
* Ultimate Go Notebook; Bill Kennedy
* Kubernetes Cookbook; Sameer Naik, Sébastien Goasguen, Jonathan Michaux; O'Reilly
+* The Kubernetes Book; Nigel Poulton; Unabridged Audiobook
+* Amazon Web Services in Action; Michael Wittig and Andreas Wittig; Manning Publications
+* Modern Perl; Chromatic ; Onyx Neon Press
+* Perl New Features; Joshua McAdams, brian d foy; Perl School
* The DevOps Handbook; Gene Kim, Jez Humble, Patrick Debois, John Willis; Audible
-* Think Raku (aka Think Perl 6); Laurent Rosenfeld, Allen B. Downey; O'Reilly
+* Systems Performance Tuning; Gian-Paolo D. Musumeci and others...; O'Reilly
## Technical references
I didn't read them from the beginning to the end, but I am using them to look up things. The books are in random order:
+* Understanding the Linux Kernel; Daniel P. Bovet, Marco Cesati; O'Reilly
+* Groovy Kurz & Gut; Joerg Staudemeier; O'Reilly
* Implementing Service Level Objectives; Alex Hidalgo; O'Reilly
+* The Linux Programming Interface; Michael Kerrisk; No Starch Press
* BPF Performance Tools - Linux System and Application Observability, Brendan Gregg; Addison Wesley
* Relayd and Httpd Mastery; Michael W Lucas
-* Groovy Kurz & Gut; Joerg Staudemeier; O'Reilly
-* Understanding the Linux Kernel; Daniel P. Bovet, Marco Cesati; O'Reilly
* Algorithms; Robert Sedgewick, Kevin Wayne; Addison Wesley
-* The Linux Programming Interface; Michael Kerrisk; No Starch Press
## Self-development and soft-skills books
In random order:
-* Solve for Happy; Mo Gawdat
-* Stop starting, start finishing; Arne Roock; Lean-Kanban University
-* The 7 Habits Of Highly Effective People; Stephen R. Covey; Simon & Schuster UK
-* Digital Minimalism; Cal Newport; Portofolio Penguin
-* The Off Switch; Mark Cropley; Virgin Books
+* Who Moved My Cheese?; Dr. Spencer Johnson; Vermilion
* Ultralearning; Anna Laurent; Self-published via Amazon
-* Influence without Authority; A. Cohen, D. Bradford; Wiley
+* Buddah and Einstein walk into a Bar; Guy Joseph Ale, Claire Bloom; Blackstone Publishing
* The Phoenix Project - A Novel About IT, DevOps, and Helping your Business Win; Gene Kim and Kevin Behr; Trade Select
-* Time Management for System Administrators; Thomas A. Limoncelli; O'Reilly
+* Influence without Authority; A. Cohen, D. Bradford; Wiley
+* Stop starting, start finishing; Arne Roock; Lean-Kanban University
+* The Joy of Missing Out; Christina Crook; New Society Publishers
* The Power of Now; Eckhard Tolle; Yellow Kite
-* Eat That Frog!; Brian Tracy; Hodder Paperbacks
-* The Good Enough Job; Simone Stolzoff; Ebury Edge
-* 101 Essays that change the way you think; Brianna Wiest; Audible
-* Staff Engineer: Leadership beyond the management track; Will Larson; Audible
+* Time Management for System Administrators; Thomas A. Limoncelli; O'Reilly
+* The Complete Software Developer's Career Guide; John Sonmez; Unabridged Audiobook
+* Consciousness: A Very Short Introduction; Susan Blackmore; Oxford Uiversity Press
* Soft Skills; John Sommez; Manning Publications
+* The Good Enough Job; Simone Stolzoff; Ebury Edge
* Eat That Frog; Brian Tracy
-* The Obstacle Is The Way; Ryan Holiday; Profile Books Ltd
-* The Joy of Missing Out; Christina Crook; New Society Publishers
-* Ultralearning; Scott Young; Thorsons
-* The Complete Software Developer's Career Guide; John Sonmez; Unabridged Audiobook
-* Search Inside Yourself - The Unexpected path to Achieving Success, Happiness (and World Peace); Chade-Meng Tan, Daniel Goleman, Jon Kabat-Zinn; HarperOne
* The Daily Stoic; Ryan Holiday, Stephen Hanselman; Profile Books
+* Slow Productivity; Cal Newport; Penguin Random House
* So Good They Can't Ignore You; Cal Newport; Business Plus
-* Deep Work; Cal Newport; Piatkus
+* 101 Essays that change the way you think; Brianna Wiest; Audible
+* Solve for Happy; Mo Gawdat
* The Bullet Journal Method; Ryder Carroll; Fourth Estate
+* Digital Minimalism; Cal Newport; Portofolio Penguin
+* Eat That Frog!; Brian Tracy; Hodder Paperbacks
+* The Off Switch; Mark Cropley; Virgin Books
* Atomic Habits; James Clear; Random House Business
+* Search Inside Yourself - The Unexpected path to Achieving Success, Happiness (and World Peace); Chade-Meng Tan, Daniel Goleman, Jon Kabat-Zinn; HarperOne
+* Deep Work; Cal Newport; Piatkus
+* Ultralearning; Scott Young; Thorsons
+* The Obstacle Is The Way; Ryan Holiday; Profile Books Ltd
* Psycho-Cybernetics; Maxwell Maltz; Perigee Books
-* Buddah and Einstein walk into a Bar; Guy Joseph Ale, Claire Bloom; Blackstone Publishing
-* Slow Productivity; Cal Newport; Penguin Random House
+* The 7 Habits Of Highly Effective People; Stephen R. Covey; Simon & Schuster UK
+* Staff Engineer: Leadership beyond the management track; Will Larson; Audible
* Never Split the Difference; Chris Voss, Tahl Raz; Random House Business
-* Consciousness: A Very Short Introduction; Susan Blackmore; Oxford Uiversity Press
-* Who Moved My Cheese?; Dr. Spencer Johnson; Vermilion
=> ../notes/index.gmi Here are notes of mine for some of the books
@@ -136,30 +136,30 @@ In random order:
Some of these were in-person with exams; others were online learning lectures only. In random order:
-* F5 Loadbalancers Training; 2-day on-site training; F5, Inc.
-* Structure and Interpretation of Computer Programs; Harold Abelson and more...;
-* Cloud Operations on AWS - Learn how to configure, deploy, maintain, and troubleshoot your AWS environments; 3-day online live training with labs; Amazon
-* Protocol buffers; O'Reilly Online
-* MySQL Deep Dive Workshop; 2-day on-site training
-* The Ultimate Kubernetes Bootcamp; School of Devops; O'Reilly Online
-* Apache Tomcat Best Practises; 3-day on-site training
-* Functional programming lecture; Remote University of Hagen
-* Scripting Vim; Damian Conway; O'Reilly Online
* Linux Security and Isolation APIs Training; Michael Kerrisk; 3-day on-site training
-* The Well-Grounded Rubyist Video Edition; David. A. Black; O'Reilly Online
-* AWS Immersion Day; Amazon; 1-day interactive online training
+* Structure and Interpretation of Computer Programs; Harold Abelson and more...;
* Developing IaC with Terraform (with Live Lessons); O'Reilly Online
+* Ultimate Go Programming; Bill Kennedy; O'Reilly Online
+* Scripting Vim; Damian Conway; O'Reilly Online
+* F5 Loadbalancers Training; 2-day on-site training; F5, Inc.
* Algorithms Video Lectures; Robert Sedgewick; O'Reilly Online
+* AWS Immersion Day; Amazon; 1-day interactive online training
+* Protocol buffers; O'Reilly Online
+* Cloud Operations on AWS - Learn how to configure, deploy, maintain, and troubleshoot your AWS environments; 3-day online live training with labs; Amazon
+* The Well-Grounded Rubyist Video Edition; David. A. Black; O'Reilly Online
+* Functional programming lecture; Remote University of Hagen
* Red Hat Certified System Administrator; Course + certification (Although I had the option, I decided not to take the next course as it is more effective to self learn what I need)
-* Ultimate Go Programming; Bill Kennedy; O'Reilly Online
+* Apache Tomcat Best Practises; 3-day on-site training
+* MySQL Deep Dive Workshop; 2-day on-site training
+* The Ultimate Kubernetes Bootcamp; School of Devops; O'Reilly Online
## Technical guides
These are not whole books, but guides (smaller or larger) which I found very useful. in random order:
+* How CPUs work at https://cpu.land
* Raku Guide at https://raku.guide
* Advanced Bash-Scripting Guide
-* How CPUs work at https://cpu.land
## Podcasts
@@ -167,45 +167,45 @@ These are not whole books, but guides (smaller or larger) which I found very use
In random order:
+* The Pragmatic Engineer Podcast
+* Cup o' Go [Golang]
+* Backend Banter
+* Deep Questions with Cal Newport
+* Dev Interrupted
* Fallthrough [Golang]
* The ProdCast (Google SRE Podcast)
-* Dev Interrupted
-* Deep Questions with Cal Newport
+* Fork Around And Find Out
+* Hidden Brain
* Maintainable
-* The Pragmatic Engineer Podcast
-* Backend Banter
* The Changelog Podcast(s)
-* Hidden Brain
-* Cup o' Go [Golang]
-* Fork Around And Find Out
### Podcasts I liked
I liked them but am not listening to them anymore. The podcasts have either "finished" (no more episodes) or I stopped listening to them due to time constraints or a shift in my interests.
-* Modern Mentor
-* FLOSS weekly
-* Ship It (predecessor of Fork Around And Find Out)
* Go Time (predecessor of fallthrough)
-* Java Pub House
* CRE: Chaosradio Express [german]
+* Ship It (predecessor of Fork Around And Find Out)
+* Modern Mentor
+* Java Pub House
+* FLOSS weekly
## Newsletters I like
This is a mix of tech and non-tech newsletters I am subscribed to. In random order:
+* Golang Weekly
* Applied Go Weekly Newsletter
-* Register Spill
-* Monospace Mentor
+* Ruby Weekly
+* Andreas Brandhorst Newsletter (Sci-Fi author)
* Changelog News
-* Golang Weekly
-* The Prgagmatic Engineer
+* The Pragmatic Engineer
+* VK Newsletter
* The Imperfectionist
-* Andreas Brandhorst Newsletter (Sci-Fi author)
+* Register Spill
* byteSizeGo
-* VK Newsletter
-* Ruby Weekly
* The Valuable Dev
+* Monospace Mentor
# Formal education
diff --git a/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.gmi b/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.gmi
index 79c8cd60..d4d01e6d 100644
--- a/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.gmi
+++ b/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.gmi
@@ -131,6 +131,7 @@ root@f0:~ # freebsd-update reboot
```
I also added the following entries for the three FreeBSD boxes to the `/etc/hosts` file:
+
```sh
root@f0:~ # cat <<END >>/etc/hosts
192.168.1.130 f0 f0.lan f0.lan.buetow.org
@@ -139,6 +140,8 @@ root@f0:~ # cat <<END >>/etc/hosts
END
```
+You might wonder why bother using the hosts file? Why not use DNS properly? The reason is simplicity. I don't manage 100 hosts, only a few here and there. Having an OpenWRT router in my home, I could also configure everything there, but maybe I'll do that later. For now, keep it simple and straightforward.
+
## After install
After that, I installed the following additional packages:
diff --git a/gemfeed/2025-01-15-working-with-an-sre-interview.gmi b/gemfeed/2025-01-15-working-with-an-sre-interview.gmi
new file mode 100644
index 00000000..9285075f
--- /dev/null
+++ b/gemfeed/2025-01-15-working-with-an-sre-interview.gmi
@@ -0,0 +1,177 @@
+# Working with an SRE Interview
+
+> Published at 2025-01-15T00:16:04+02:00
+
+I have been interviewed by Florian Buetow about what it's like working with a Site Reliability Engineer from the point of view of a Software Engineer, Data Scientist, and AI Engineer.
+
+=> https://www.cracking-ai-engineering.com/writing/2025/01/12/working-with-an-sre-interview/ See original interview here
+
+Below, I am posting the interview here on my blog as well.
+
+## Table of Contents
+
+* ⇢ Working with an SRE Interview
+* ⇢ ⇢ Preamble
+* ⇢ ⇢ Introducing Paul
+* ⇢ ⇢ How did you get started?
+* ⇢ ⇢ Roles and Career Progression
+* ⇢ ⇢ Anecdotes and Best Practices
+* ⇢ ⇢ Working with Different Teams
+* ⇢ ⇢ Using AI Tools
+* ⇢ ⇢ SRE Learning Resources
+* ⇢ ⇢ Blogging
+* ⇢ ⇢ Wrap-up
+* ⇢ ⇢ Closing comments
+
+## Preamble
+
+In this insightful interview, Paul Bütow, a Principal Site Reliability Engineer at Mimecast, shares over a decade of experience in the field. Paul highlights the role of an Embedded SRE, emphasizing the importance of automation, observability, and effective incident management. We also focused on the key question of how you can work effectively with an SRE weather you are an individual contributor or a manager, a software engineer or data scientist. And how you can learn more about site reliability engineering.
+
+## Introducing Paul
+
+Hi Paul, please introduce yourself briefly to the audience. Who are you, what do you do for a living, and where do you work?
+
+> My name is Paul Bütow, I work at Mimecast, and I’m a Principal Site Reliability Engineer there. I’ve been with Mimecast for almost ten years now. The company specializes in email security, including things like archiving, phishing detection, malware protection, and spam filtering.
+
+You mentioned that you’re an ‘Embedded SRE.’ What does that mean exactly?
+
+> It means that I’m directly part of the software engineering team, not in a separate Ops department. I ensure that nothing is deployed manually, and everything runs through automation. I also set up monitoring and observability. These are two distinct aspects: monitoring alerts us when something breaks, while observability helps us identify trends. I also create runbooks so we know what to do when specific incidents occur frequently.
+
+> Infrastructure SREs on the other hand handle the foundational setup, like providing the Kubernetes cluster itself or ensuring the operating systems are installed. They don't work on the application directly but ensure the base infrastructure is there for others to use. This works well when a company has multiple teams that need shared infrastructure.
+
+## How did you get started?
+
+How did your interest in Linux or FreeBSD start?
+
+> It began during my school days. We had a PC with DOS at home, and I eventually bought Suse Linux 5.3. Shortly after, I discovered FreeBSD because I liked its handbook so much. I wanted to understand exactly how everything worked, so I also tried Linux from Scratch. That involves installing every package manually to gain a better understanding of operating systems.
+
+=> https://www.FreeBSD.org
+=> https://linuxfromscratch.org/
+
+And after school, you pursued computer science, correct?
+
+> Exactly. I wasn’t sure at first whether I wanted to be a software developer or a system administrator. I applied for both and eventually accepted an offer as a Linux system administrator. This was before 'SRE' became a buzzword, but much of what I did back then-automation, infrastructure as code, monitoring-is now considered part of the typical SRE role.
+
+## Roles and Career Progression
+
+Tell us about how you joined Mimecast. When did you fully embrace the SRE role?
+
+> I started as a Linux sysadmin at 1&1. I managed an ad server farm with hundreds of systems and later handled load balancers. Together with an architect, we managed F5 load balancers distributing around 2,000 services, including for portals like web.de and GMX. I also led the operations team technically for a while before moving to London to join Mimecast.
+
+> At Mimecast, the job title was explicitly 'Site Reliability Engineer.' The biggest difference was that I was no longer in a separate Ops department but embedded directly within the storage and search backend team. I loved that because we could plan features together-from automation to measurability and observability. Mimecast also operates thousands of physical servers for email archiving, which was fascinating since I already had experience with large distributed systems at 1&1. It was the right step for me because it allowed me to work close to the code while remaining hands-on with infrastructure.
+
+What are the differences between SRE, DevOps, SysAdmin, and Architects?
+
+> SREs are like the next step after SysAdmins. A SysAdmin might manually install servers, replace disks, or use simple scripts for automation, while SREs use infrastructure as code and focus on reliability through SLIs, SLOs, and automation. DevOps isn’t really a job-it’s more of a way of working, where developers are involved in operations tasks like setting up CI/CD pipelines or on-call shifts. Architects focus on designing systems and infrastructures, such as load balancers or distributed systems, working alongside SREs to ensure the systems meet the reliability and scalability requirements. The specific responsibilities of each role depend on the company, and there is often overlap.
+
+What are the most important reliability lessons you’ve learned so far?
+
+* Don’t leave SRE aspects as an afterthought. It’s much better to discuss automation, monitoring, SLIs, and SLOs early on. Traditional sysadmins often installed systems manually, but today, we do everything via infrastructure as code-using tools like Terraform or Puppet.
+* I also distinguish between monitoring and observability. Monitoring tells us, 'The server is down, alarm!' Observability dives deeper, showing trends like increasing latency so we can act proactively.
+* SLI, SLO, and SLA are core elements. We focus on what users actually experience-for example, how quickly an email is sent-and set our goals accordingly.
+* Runbooks are also crucial. When something goes wrong at night, you don’t want to start from scratch. A runbook outlines how to debug and resolve specific problems, saving time and reducing downtime.
+
+## Anecdotes and Best Practices
+
+Runbooks sound very practical. Can you explain how they’re used day-to-day?
+
+> Runbooks are essentially guides for handling specific incidents. For instance, if a service won’t start, the runbook will specify where the logs are and which commands to use. Observability takes it a step further, helping us spot changes early-like rising error rates or latency-so we can address issues before they escalate.
+
+When should you decide to put something into a runbook, and when is it unnecessary?
+
+> If an issue happens frequently, it should be documented in a runbook so that anyone, even someone new, can follow the steps to fix it. The idea is that 90% of the common incidents should be covered. For example, if a service is down, the runbook would specify where to find logs, which commands to check, and what actions to take. On the other hand, rare or complex issues, where the resolution depends heavily on context or varies each time, don’t make sense to include in detail. For those, it’s better to focus on general troubleshooting steps.
+
+How do you search for and find the correct runbooks?
+
+> Runbooks should be linked directly in the alert you receive. For example, if you get an alert about a service not running, the alert will have a link to the runbook that tells you what to check, like logs or commands to run. Runbooks are best stored in an internal wiki, so if you don’t find the link in the alert, you know where to search. The important thing is that runbooks are easy to find and up to date because that’s what makes them useful during incidents.
+
+Do you have an interesting war story you can share with us?
+
+> Sure. At 1&1, we had a proprietary ad server software that ran a SQL query during startup. The query got slower over time, eventually timing out and preventing the server from starting. Since we couldn’t access the source code, we searched the binary for the SQL and patched it. By pinpointing the issue, a developer was able to adjust the SQL. This collaboration between sysadmin and developer perspectives highlights the value of SRE work.
+
+## Working with Different Teams
+
+You’re embedded in a team-how does collaboration with developers work practically?
+
+> We plan everything together from the start. If there’s a new feature, we discuss infrastructure, automated deployments, and monitoring right away. Developers are experts in the code, and I bring the infrastructure expertise. This avoids unpleasant surprises before going live.
+
+How about working with data scientists or ML engineers? Are there differences?
+
+> The principles are the same. ML models also need to be deployed and monitored. You deal with monitoring, resource allocation, and identifying performance drops. Whether it’s a microservice or an ML job, at the end of the day, it’s all running on servers or clusters that must remain stable.
+
+What about working with managers or the FinOps team?
+
+> We often discuss costs, especially in the cloud, where scaling up resources is easy. It’s crucial to know our metrics: do we have enough capacity? Do we need all instances? Or is the CPU only at 5% utilization? This data helps managers decide whether the budget is sufficient or if optimizations are needed.
+
+Do you have practical tips for working with SREs?
+
+> Yes, I have a few:
+
+* Early involvement: Include SREs from the beginning in your project.
+* Runbooks & documentation: Document recurring errors.
+* Try first: Try to understand the issue yourself before immediately asking the SRE.
+* Basic infra knowledge: Kubernetes and Terraform aren’t magic. Some basic understanding helps every developer.
+
+## Using AI Tools
+
+Let’s talk about AI. How do you use it in your daily work?
+
+> For boilerplate code, like Terraform snippets, I often use ChatGPT. It saves time, although I always review and adjust the output. Log analysis is another exciting application. Instead of manually going through millions of lines, AI can summarize key outliers or errors.
+
+Do you think AI could largely replace SREs or significantly change the role?
+
+> I see AI as an additional tool. SRE requires a deep understanding of how distributed systems work internally. While AI can assist with routine tasks or quickly detect anomalies, human expertise is indispensable for complex issues.
+
+## SRE Learning Resources
+
+What resources would you recommend for learning about SRE?
+
+> The Google SRE book is a classic, though a bit dry. I really like 'Seeking SRE,' as it offers various perspectives on SRE, with many practical stories from different companies.
+
+=> https://sre.google/books/
+=> https://www.oreilly.com/library/view/seeking-sre/9781491978856 Seeking SRE
+
+Do you have a podcast recommendation?
+
+> The Google SRE prodcast is quite interesting. It offers insights into how Google approaches SRE, along with perspectives from external guests.
+
+=> https://sre.google/prodcast/
+
+## Blogging
+
+You also have a blog. What motivates you to write regularly?
+
+> Writing helps me learn the most. It also serves as a personal reference. Sometimes I look up how I solved a problem a year ago. And of course, others tackling similar projects might find inspiration in my posts.
+
+What do you blog about?
+
+> Mostly technical topics I find exciting, like homelab projects, Kubernetes, or book summaries on IT and productivity. It’s a personal blog, so I write about what I enjoy.
+
+## Wrap-up
+
+To wrap up, what are three things every team should keep in mind for stability?
+
+> First, maintain runbooks and documentation to avoid chaos at night. Second, automate everything-manual installs in production are risky. Third, define SLIs, SLOs, and SLAs early so everyone knows what we’re monitoring and guaranteeing.
+
+Is there a motto or mindset that particularly inspires you as an SRE?
+
+> "Keep it simple and stupid"-KISS. Not everything has to be overly complex. And always stay curious. I’m still fascinated by how systems work under the hood.
+
+Where can people find you online?
+
+> You can find links to my socials on my website paul.buetow.org
+> I regularly post articles and link to everything else I’m working on outside of work.
+
+=> https://paul.buetow.org
+
+Thank you very much for your time and this insightful interview into the world of site reliability engineering
+
+> My pleasure, this was fun.
+
+## Closing comments
+
+Dear reader, I hope this conversation with Paul Bütow provided an exciting peak into the world of Site Reliability Engineering. Whether you’re a software developer, data scientist, ML engineer, or manager, reliable systems are always a team effort. Hopefully, you’ve taken some insights or tips from Paul’s experiences for your own team or next project. Thanks for joining us, and best of luck refining your own SRE practices!
+
+E-Mail your comments to `paul@nospam.buetow.org` :-)
+
+=> ../ Back to the main site
diff --git a/gemfeed/2025-01-15-working-with-an-sre-interview.gmi.tpl b/gemfeed/2025-01-15-working-with-an-sre-interview.gmi.tpl
new file mode 100644
index 00000000..8399cc58
--- /dev/null
+++ b/gemfeed/2025-01-15-working-with-an-sre-interview.gmi.tpl
@@ -0,0 +1,164 @@
+# Working with an SRE Interview
+
+> Published at 2025-01-15T00:16:04+02:00
+
+I have been interviewed by Florian Buetow about what it's like working with a Site Reliability Engineer from the point of view of a Software Engineer, Data Scientist, and AI Engineer.
+
+=> https://www.cracking-ai-engineering.com/writing/2025/01/12/working-with-an-sre-interview/ See original interview here
+
+Below, I am posting the interview here on my blog as well.
+
+<< template::inline::toc
+
+## Preamble
+
+In this insightful interview, Paul Bütow, a Principal Site Reliability Engineer at Mimecast, shares over a decade of experience in the field. Paul highlights the role of an Embedded SRE, emphasizing the importance of automation, observability, and effective incident management. We also focused on the key question of how you can work effectively with an SRE weather you are an individual contributor or a manager, a software engineer or data scientist. And how you can learn more about site reliability engineering.
+
+## Introducing Paul
+
+Hi Paul, please introduce yourself briefly to the audience. Who are you, what do you do for a living, and where do you work?
+
+> My name is Paul Bütow, I work at Mimecast, and I’m a Principal Site Reliability Engineer there. I’ve been with Mimecast for almost ten years now. The company specializes in email security, including things like archiving, phishing detection, malware protection, and spam filtering.
+
+You mentioned that you’re an ‘Embedded SRE.’ What does that mean exactly?
+
+> It means that I’m directly part of the software engineering team, not in a separate Ops department. I ensure that nothing is deployed manually, and everything runs through automation. I also set up monitoring and observability. These are two distinct aspects: monitoring alerts us when something breaks, while observability helps us identify trends. I also create runbooks so we know what to do when specific incidents occur frequently.
+
+> Infrastructure SREs on the other hand handle the foundational setup, like providing the Kubernetes cluster itself or ensuring the operating systems are installed. They don't work on the application directly but ensure the base infrastructure is there for others to use. This works well when a company has multiple teams that need shared infrastructure.
+
+## How did you get started?
+
+How did your interest in Linux or FreeBSD start?
+
+> It began during my school days. We had a PC with DOS at home, and I eventually bought Suse Linux 5.3. Shortly after, I discovered FreeBSD because I liked its handbook so much. I wanted to understand exactly how everything worked, so I also tried Linux from Scratch. That involves installing every package manually to gain a better understanding of operating systems.
+
+=> https://www.FreeBSD.org
+=> https://linuxfromscratch.org/
+
+And after school, you pursued computer science, correct?
+
+> Exactly. I wasn’t sure at first whether I wanted to be a software developer or a system administrator. I applied for both and eventually accepted an offer as a Linux system administrator. This was before 'SRE' became a buzzword, but much of what I did back then-automation, infrastructure as code, monitoring-is now considered part of the typical SRE role.
+
+## Roles and Career Progression
+
+Tell us about how you joined Mimecast. When did you fully embrace the SRE role?
+
+> I started as a Linux sysadmin at 1&1. I managed an ad server farm with hundreds of systems and later handled load balancers. Together with an architect, we managed F5 load balancers distributing around 2,000 services, including for portals like web.de and GMX. I also led the operations team technically for a while before moving to London to join Mimecast.
+
+> At Mimecast, the job title was explicitly 'Site Reliability Engineer.' The biggest difference was that I was no longer in a separate Ops department but embedded directly within the storage and search backend team. I loved that because we could plan features together-from automation to measurability and observability. Mimecast also operates thousands of physical servers for email archiving, which was fascinating since I already had experience with large distributed systems at 1&1. It was the right step for me because it allowed me to work close to the code while remaining hands-on with infrastructure.
+
+What are the differences between SRE, DevOps, SysAdmin, and Architects?
+
+> SREs are like the next step after SysAdmins. A SysAdmin might manually install servers, replace disks, or use simple scripts for automation, while SREs use infrastructure as code and focus on reliability through SLIs, SLOs, and automation. DevOps isn’t really a job-it’s more of a way of working, where developers are involved in operations tasks like setting up CI/CD pipelines or on-call shifts. Architects focus on designing systems and infrastructures, such as load balancers or distributed systems, working alongside SREs to ensure the systems meet the reliability and scalability requirements. The specific responsibilities of each role depend on the company, and there is often overlap.
+
+What are the most important reliability lessons you’ve learned so far?
+
+* Don’t leave SRE aspects as an afterthought. It’s much better to discuss automation, monitoring, SLIs, and SLOs early on. Traditional sysadmins often installed systems manually, but today, we do everything via infrastructure as code-using tools like Terraform or Puppet.
+* I also distinguish between monitoring and observability. Monitoring tells us, 'The server is down, alarm!' Observability dives deeper, showing trends like increasing latency so we can act proactively.
+* SLI, SLO, and SLA are core elements. We focus on what users actually experience-for example, how quickly an email is sent-and set our goals accordingly.
+* Runbooks are also crucial. When something goes wrong at night, you don’t want to start from scratch. A runbook outlines how to debug and resolve specific problems, saving time and reducing downtime.
+
+## Anecdotes and Best Practices
+
+Runbooks sound very practical. Can you explain how they’re used day-to-day?
+
+> Runbooks are essentially guides for handling specific incidents. For instance, if a service won’t start, the runbook will specify where the logs are and which commands to use. Observability takes it a step further, helping us spot changes early-like rising error rates or latency-so we can address issues before they escalate.
+
+When should you decide to put something into a runbook, and when is it unnecessary?
+
+> If an issue happens frequently, it should be documented in a runbook so that anyone, even someone new, can follow the steps to fix it. The idea is that 90% of the common incidents should be covered. For example, if a service is down, the runbook would specify where to find logs, which commands to check, and what actions to take. On the other hand, rare or complex issues, where the resolution depends heavily on context or varies each time, don’t make sense to include in detail. For those, it’s better to focus on general troubleshooting steps.
+
+How do you search for and find the correct runbooks?
+
+> Runbooks should be linked directly in the alert you receive. For example, if you get an alert about a service not running, the alert will have a link to the runbook that tells you what to check, like logs or commands to run. Runbooks are best stored in an internal wiki, so if you don’t find the link in the alert, you know where to search. The important thing is that runbooks are easy to find and up to date because that’s what makes them useful during incidents.
+
+Do you have an interesting war story you can share with us?
+
+> Sure. At 1&1, we had a proprietary ad server software that ran a SQL query during startup. The query got slower over time, eventually timing out and preventing the server from starting. Since we couldn’t access the source code, we searched the binary for the SQL and patched it. By pinpointing the issue, a developer was able to adjust the SQL. This collaboration between sysadmin and developer perspectives highlights the value of SRE work.
+
+## Working with Different Teams
+
+You’re embedded in a team-how does collaboration with developers work practically?
+
+> We plan everything together from the start. If there’s a new feature, we discuss infrastructure, automated deployments, and monitoring right away. Developers are experts in the code, and I bring the infrastructure expertise. This avoids unpleasant surprises before going live.
+
+How about working with data scientists or ML engineers? Are there differences?
+
+> The principles are the same. ML models also need to be deployed and monitored. You deal with monitoring, resource allocation, and identifying performance drops. Whether it’s a microservice or an ML job, at the end of the day, it’s all running on servers or clusters that must remain stable.
+
+What about working with managers or the FinOps team?
+
+> We often discuss costs, especially in the cloud, where scaling up resources is easy. It’s crucial to know our metrics: do we have enough capacity? Do we need all instances? Or is the CPU only at 5% utilization? This data helps managers decide whether the budget is sufficient or if optimizations are needed.
+
+Do you have practical tips for working with SREs?
+
+> Yes, I have a few:
+
+* Early involvement: Include SREs from the beginning in your project.
+* Runbooks & documentation: Document recurring errors.
+* Try first: Try to understand the issue yourself before immediately asking the SRE.
+* Basic infra knowledge: Kubernetes and Terraform aren’t magic. Some basic understanding helps every developer.
+
+## Using AI Tools
+
+Let’s talk about AI. How do you use it in your daily work?
+
+> For boilerplate code, like Terraform snippets, I often use ChatGPT. It saves time, although I always review and adjust the output. Log analysis is another exciting application. Instead of manually going through millions of lines, AI can summarize key outliers or errors.
+
+Do you think AI could largely replace SREs or significantly change the role?
+
+> I see AI as an additional tool. SRE requires a deep understanding of how distributed systems work internally. While AI can assist with routine tasks or quickly detect anomalies, human expertise is indispensable for complex issues.
+
+## SRE Learning Resources
+
+What resources would you recommend for learning about SRE?
+
+> The Google SRE book is a classic, though a bit dry. I really like 'Seeking SRE,' as it offers various perspectives on SRE, with many practical stories from different companies.
+
+=> https://sre.google/books/
+=> https://www.oreilly.com/library/view/seeking-sre/9781491978856 Seeking SRE
+
+Do you have a podcast recommendation?
+
+> The Google SRE prodcast is quite interesting. It offers insights into how Google approaches SRE, along with perspectives from external guests.
+
+=> https://sre.google/prodcast/
+
+## Blogging
+
+You also have a blog. What motivates you to write regularly?
+
+> Writing helps me learn the most. It also serves as a personal reference. Sometimes I look up how I solved a problem a year ago. And of course, others tackling similar projects might find inspiration in my posts.
+
+What do you blog about?
+
+> Mostly technical topics I find exciting, like homelab projects, Kubernetes, or book summaries on IT and productivity. It’s a personal blog, so I write about what I enjoy.
+
+## Wrap-up
+
+To wrap up, what are three things every team should keep in mind for stability?
+
+> First, maintain runbooks and documentation to avoid chaos at night. Second, automate everything-manual installs in production are risky. Third, define SLIs, SLOs, and SLAs early so everyone knows what we’re monitoring and guaranteeing.
+
+Is there a motto or mindset that particularly inspires you as an SRE?
+
+> "Keep it simple and stupid"-KISS. Not everything has to be overly complex. And always stay curious. I’m still fascinated by how systems work under the hood.
+
+Where can people find you online?
+
+> You can find links to my socials on my website paul.buetow.org
+> I regularly post articles and link