summaryrefslogtreecommitdiff
path: root/gemfeed
diff options
context:
space:
mode:
Diffstat (limited to 'gemfeed')
-rw-r--r--gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.html3
-rw-r--r--gemfeed/2025-01-15-working-with-an-sre-interview.html200
-rw-r--r--gemfeed/DRAFT-f3s-kubernetes-with-freebsd-bhyve.html100
-rw-r--r--gemfeed/atom.xml568
-rw-r--r--gemfeed/index.html1
5 files changed, 485 insertions, 387 deletions
diff --git a/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.html b/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.html
index 6ea80fc3..cfa8ac68 100644
--- a/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.html
+++ b/gemfeed/2024-12-03-f3s-kubernetes-with-freebsd-part-2.html
@@ -154,6 +154,7 @@ root@f0:~ <i><font color="silver"># freebsd-update reboot</font></i>
</pre>
<br />
<span>I also added the following entries for the three FreeBSD boxes to the <span class='inlinecode'>/etc/hosts</span> file:</span><br />
+<br />
<!-- Generator: GNU source-highlight 3.1.9
by Lorenzo Bettini
http://www.lorenzobettini.it
@@ -165,6 +166,8 @@ http://www.gnu.org/software/src-highlite -->
END
</pre>
<br />
+<span>You might wonder why bother using the hosts file? Why not use DNS properly? The reason is simplicity. I don&#39;t manage 100 hosts, only a few here and there. Having an OpenWRT router in my home, I could also configure everything there, but maybe I&#39;ll do that later. For now, keep it simple and straightforward.</span><br />
+<br />
<h2 style='display: inline' id='after-install'>After install</h2><br />
<br />
<span>After that, I installed the following additional packages:</span><br />
diff --git a/gemfeed/2025-01-15-working-with-an-sre-interview.html b/gemfeed/2025-01-15-working-with-an-sre-interview.html
new file mode 100644
index 00000000..f4e85097
--- /dev/null
+++ b/gemfeed/2025-01-15-working-with-an-sre-interview.html
@@ -0,0 +1,200 @@
+<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
+<html xmlns="http://www.w3.org/1999/xhtml" lang="en" xml:lang="en">
+<head>
+<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
+<title>Working with an SRE Interview</title>
+<link rel="shortcut icon" type="image/gif" href="/favicon.ico" />
+<link rel="stylesheet" href="../style.css" />
+<link rel="stylesheet" href="style-override.css" />
+</head>
+<body>
+<p class="header">
+View this page as <a href="https://codeberg.org/snonux/foo.zone/src/branch/content-md/gemfeed/2025-01-15-working-with-an-sre-interview.md">Markdown</a> | <a href="gemini://foo.zone/gemfeed/2025-01-15-working-with-an-sre-interview.gmi">Gemtext</a>
+</p>
+<h1 style='display: inline' id='working-with-an-sre-interview'>Working with an SRE Interview</h1><br />
+<br />
+<span class='quote'>Published at 2025-01-15T00:16:04+02:00</span><br />
+<br />
+<span>I have been interviewed by Florian Buetow about what it&#39;s like working with a Site Reliability Engineer from the point of view of a Software Engineer, Data Scientist, and AI Engineer. </span><br />
+<br />
+<a class='textlink' href='https://www.cracking-ai-engineering.com/writing/2025/01/12/working-with-an-sre-interview/'>See original interview here</a><br />
+<br />
+<span>Below, I am posting the interview here on my blog as well.</span><br />
+<br />
+<h2 style='display: inline' id='table-of-contents'>Table of Contents</h2><br />
+<br />
+<ul>
+<li><a href='#working-with-an-sre-interview'>Working with an SRE Interview</a></li>
+<li>⇢ <a href='#preamble-'>Preamble </a></li>
+<li>⇢ <a href='#introducing-paul'>Introducing Paul</a></li>
+<li>⇢ <a href='#how-did-you-get-started'>How did you get started?</a></li>
+<li>⇢ <a href='#roles-and-career-progression'>Roles and Career Progression</a></li>
+<li>⇢ <a href='#anecdotes-and-best-practices'>Anecdotes and Best Practices</a></li>
+<li>⇢ <a href='#working-with-different-teams'>Working with Different Teams</a></li>
+<li>⇢ <a href='#using-ai-tools'>Using AI Tools</a></li>
+<li>⇢ <a href='#sre-learning-resources'>SRE Learning Resources</a></li>
+<li>⇢ <a href='#blogging'>Blogging</a></li>
+<li>⇢ <a href='#wrap-up'>Wrap-up</a></li>
+<li>⇢ <a href='#closing-comments'>Closing comments</a></li>
+</ul><br />
+<h2 style='display: inline' id='preamble-'>Preamble </h2><br />
+<br />
+<span>In this insightful interview, Paul Bütow, a Principal Site Reliability Engineer at Mimecast, shares over a decade of experience in the field. Paul highlights the role of an Embedded SRE, emphasizing the importance of automation, observability, and effective incident management. We also focused on the key question of how you can work effectively with an SRE weather you are an individual contributor or a manager, a software engineer or data scientist. And how you can learn more about site reliability engineering.</span><br />
+<br />
+<h2 style='display: inline' id='introducing-paul'>Introducing Paul</h2><br />
+<br />
+<span>Hi Paul, please introduce yourself briefly to the audience. Who are you, what do you do for a living, and where do you work?</span><br />
+<br />
+<span class='quote'>My name is Paul Bütow, I work at Mimecast, and I’m a Principal Site Reliability Engineer there. I’ve been with Mimecast for almost ten years now. The company specializes in email security, including things like archiving, phishing detection, malware protection, and spam filtering.</span><br />
+<br />
+<span>You mentioned that you’re an ‘Embedded SRE.’ What does that mean exactly?</span><br />
+<br />
+<span class='quote'>It means that I’m directly part of the software engineering team, not in a separate Ops department. I ensure that nothing is deployed manually, and everything runs through automation. I also set up monitoring and observability. These are two distinct aspects: monitoring alerts us when something breaks, while observability helps us identify trends. I also create runbooks so we know what to do when specific incidents occur frequently.</span><br />
+<br />
+<span class='quote'>Infrastructure SREs on the other hand handle the foundational setup, like providing the Kubernetes cluster itself or ensuring the operating systems are installed. They don&#39;t work on the application directly but ensure the base infrastructure is there for others to use. This works well when a company has multiple teams that need shared infrastructure.</span><br />
+<br />
+<h2 style='display: inline' id='how-did-you-get-started'>How did you get started?</h2><br />
+<br />
+<span>How did your interest in Linux or FreeBSD start?</span><br />
+<br />
+<span class='quote'>It began during my school days. We had a PC with DOS at home, and I eventually bought Suse Linux 5.3. Shortly after, I discovered FreeBSD because I liked its handbook so much. I wanted to understand exactly how everything worked, so I also tried Linux from Scratch. That involves installing every package manually to gain a better understanding of operating systems.</span><br />
+<br />
+<a class='textlink' href='https://www.FreeBSD.org'>https://www.FreeBSD.org</a><br />
+<a class='textlink' href='https://linuxfromscratch.org/'>https://linuxfromscratch.org/</a><br />
+<br />
+<span>And after school, you pursued computer science, correct?</span><br />
+<br />
+<span class='quote'>Exactly. I wasn’t sure at first whether I wanted to be a software developer or a system administrator. I applied for both and eventually accepted an offer as a Linux system administrator. This was before &#39;SRE&#39; became a buzzword, but much of what I did back then-automation, infrastructure as code, monitoring-is now considered part of the typical SRE role.</span><br />
+<br />
+<h2 style='display: inline' id='roles-and-career-progression'>Roles and Career Progression</h2><br />
+<br />
+<span>Tell us about how you joined Mimecast. When did you fully embrace the SRE role?</span><br />
+<br />
+<span class='quote'>I started as a Linux sysadmin at 1&amp;1. I managed an ad server farm with hundreds of systems and later handled load balancers. Together with an architect, we managed F5 load balancers distributing around 2,000 services, including for portals like web.de and GMX. I also led the operations team technically for a while before moving to London to join Mimecast.</span><br />
+<br />
+<span class='quote'>At Mimecast, the job title was explicitly &#39;Site Reliability Engineer.&#39; The biggest difference was that I was no longer in a separate Ops department but embedded directly within the storage and search backend team. I loved that because we could plan features together-from automation to measurability and observability. Mimecast also operates thousands of physical servers for email archiving, which was fascinating since I already had experience with large distributed systems at 1&amp;1. It was the right step for me because it allowed me to work close to the code while remaining hands-on with infrastructure.</span><br />
+<br />
+<span>What are the differences between SRE, DevOps, SysAdmin, and Architects?</span><br />
+<br />
+<span class='quote'>SREs are like the next step after SysAdmins. A SysAdmin might manually install servers, replace disks, or use simple scripts for automation, while SREs use infrastructure as code and focus on reliability through SLIs, SLOs, and automation. DevOps isn’t really a job-it’s more of a way of working, where developers are involved in operations tasks like setting up CI/CD pipelines or on-call shifts. Architects focus on designing systems and infrastructures, such as load balancers or distributed systems, working alongside SREs to ensure the systems meet the reliability and scalability requirements. The specific responsibilities of each role depend on the company, and there is often overlap. </span><br />
+<br />
+<span>What are the most important reliability lessons you’ve learned so far?</span><br />
+<br />
+<ul>
+<li>Don’t leave SRE aspects as an afterthought. It’s much better to discuss automation, monitoring, SLIs, and SLOs early on. Traditional sysadmins often installed systems manually, but today, we do everything via infrastructure as code-using tools like Terraform or Puppet.</li>
+<li>I also distinguish between monitoring and observability. Monitoring tells us, &#39;The server is down, alarm!&#39; Observability dives deeper, showing trends like increasing latency so we can act proactively.</li>
+<li>SLI, SLO, and SLA are core elements. We focus on what users actually experience-for example, how quickly an email is sent-and set our goals accordingly.</li>
+<li>Runbooks are also crucial. When something goes wrong at night, you don’t want to start from scratch. A runbook outlines how to debug and resolve specific problems, saving time and reducing downtime.</li>
+</ul><br />
+<h2 style='display: inline' id='anecdotes-and-best-practices'>Anecdotes and Best Practices</h2><br />
+<br />
+<span>Runbooks sound very practical. Can you explain how they’re used day-to-day?</span><br />
+<br />
+<span class='quote'>Runbooks are essentially guides for handling specific incidents. For instance, if a service won’t start, the runbook will specify where the logs are and which commands to use. Observability takes it a step further, helping us spot changes early-like rising error rates or latency-so we can address issues before they escalate.</span><br />
+<br />
+<span>When should you decide to put something into a runbook, and when is it unnecessary?</span><br />
+<br />
+<span class='quote'>If an issue happens frequently, it should be documented in a runbook so that anyone, even someone new, can follow the steps to fix it. The idea is that 90% of the common incidents should be covered. For example, if a service is down, the runbook would specify where to find logs, which commands to check, and what actions to take. On the other hand, rare or complex issues, where the resolution depends heavily on context or varies each time, don’t make sense to include in detail. For those, it’s better to focus on general troubleshooting steps. </span><br />
+<br />
+<span>How do you search for and find the correct runbooks?</span><br />
+<br />
+<span class='quote'>Runbooks should be linked directly in the alert you receive. For example, if you get an alert about a service not running, the alert will have a link to the runbook that tells you what to check, like logs or commands to run. Runbooks are best stored in an internal wiki, so if you don’t find the link in the alert, you know where to search. The important thing is that runbooks are easy to find and up to date because that’s what makes them useful during incidents. </span><br />
+<br />
+<span>Do you have an interesting war story you can share with us?</span><br />
+<br />
+<span class='quote'>Sure. At 1&amp;1, we had a proprietary ad server software that ran a SQL query during startup. The query got slower over time, eventually timing out and preventing the server from starting. Since we couldn’t access the source code, we searched the binary for the SQL and patched it. By pinpointing the issue, a developer was able to adjust the SQL. This collaboration between sysadmin and developer perspectives highlights the value of SRE work.</span><br />
+<br />
+<h2 style='display: inline' id='working-with-different-teams'>Working with Different Teams</h2><br />
+<br />
+<span>You’re embedded in a team-how does collaboration with developers work practically?</span><br />
+<br />
+<span class='quote'>We plan everything together from the start. If there’s a new feature, we discuss infrastructure, automated deployments, and monitoring right away. Developers are experts in the code, and I bring the infrastructure expertise. This avoids unpleasant surprises before going live.</span><br />
+<br />
+<span>How about working with data scientists or ML engineers? Are there differences?</span><br />
+<br />
+<span class='quote'>The principles are the same. ML models also need to be deployed and monitored. You deal with monitoring, resource allocation, and identifying performance drops. Whether it’s a microservice or an ML job, at the end of the day, it’s all running on servers or clusters that must remain stable.</span><br />
+<br />
+<span>What about working with managers or the FinOps team?</span><br />
+<br />
+<span class='quote'>We often discuss costs, especially in the cloud, where scaling up resources is easy. It’s crucial to know our metrics: do we have enough capacity? Do we need all instances? Or is the CPU only at 5% utilization? This data helps managers decide whether the budget is sufficient or if optimizations are needed.</span><br />
+<br />
+<span>Do you have practical tips for working with SREs?</span><br />
+<br />
+<span class='quote'>Yes, I have a few:</span><br />
+<br />
+<ul>
+<li>Early involvement: Include SREs from the beginning in your project.</li>
+<li>Runbooks &amp; documentation: Document recurring errors.</li>
+<li>Try first: Try to understand the issue yourself before immediately asking the SRE.</li>
+<li>Basic infra knowledge: Kubernetes and Terraform aren’t magic. Some basic understanding helps every developer.</li>
+</ul><br />
+<h2 style='display: inline' id='using-ai-tools'>Using AI Tools</h2><br />
+<br />
+<span>Let’s talk about AI. How do you use it in your daily work?</span><br />
+<br />
+<span class='quote'>For boilerplate code, like Terraform snippets, I often use ChatGPT. It saves time, although I always review and adjust the output. Log analysis is another exciting application. Instead of manually going through millions of lines, AI can summarize key outliers or errors.</span><br />
+<br />
+<span>Do you think AI could largely replace SREs or significantly change the role?</span><br />
+<br />
+<span class='quote'>I see AI as an additional tool. SRE requires a deep understanding of how distributed systems work internally. While AI can assist with routine tasks or quickly detect anomalies, human expertise is indispensable for complex issues.</span><br />
+<br />
+<h2 style='display: inline' id='sre-learning-resources'>SRE Learning Resources</h2><br />
+<br />
+<span>What resources would you recommend for learning about SRE?</span><br />
+<br />
+<span class='quote'>The Google SRE book is a classic, though a bit dry. I really like &#39;Seeking SRE,&#39; as it offers various perspectives on SRE, with many practical stories from different companies.</span><br />
+<br />
+<a class='textlink' href='https://sre.google/books/'>https://sre.google/books/</a><br />
+<a class='textlink' href='https://www.oreilly.com/library/view/seeking-sre/9781491978856'>Seeking SRE</a><br />
+<br />
+<span>Do you have a podcast recommendation?</span><br />
+<br />
+<span class='quote'>The Google SRE prodcast is quite interesting. It offers insights into how Google approaches SRE, along with perspectives from external guests.</span><br />
+<br />
+<a class='textlink' href='https://sre.google/prodcast/'>https://sre.google/prodcast/</a><br />
+<br />
+<h2 style='display: inline' id='blogging'>Blogging</h2><br />
+<br />
+<span>You also have a blog. What motivates you to write regularly?</span><br />
+<br />
+<span class='quote'>Writing helps me learn the most. It also serves as a personal reference. Sometimes I look up how I solved a problem a year ago. And of course, others tackling similar projects might find inspiration in my posts.</span><br />
+<br />
+<span>What do you blog about?</span><br />
+<br />
+<span class='quote'>Mostly technical topics I find exciting, like homelab projects, Kubernetes, or book summaries on IT and productivity. It’s a personal blog, so I write about what I enjoy.</span><br />
+<br />
+<h2 style='display: inline' id='wrap-up'>Wrap-up</h2><br />
+<br />
+<span>To wrap up, what are three things every team should keep in mind for stability?</span><br />
+<br />
+<span class='quote'>First, maintain runbooks and documentation to avoid chaos at night. Second, automate everything-manual installs in production are risky. Third, define SLIs, SLOs, and SLAs early so everyone knows what we’re monitoring and guaranteeing.</span><br />
+<br />
+<span>Is there a motto or mindset that particularly inspires you as an SRE?</span><br />
+<br />
+<span class='quote'>"Keep it simple and stupid"-KISS. Not everything has to be overly complex. And always stay curious. I’m still fascinated by how systems work under the hood.</span><br />
+<br />
+<span>Where can people find you online?</span><br />
+<br />
+<span class='quote'>You can find links to my socials on my website paul.buetow.org</span><br />
+<span class='quote'>I regularly post articles and link to everything else I’m working on outside of work.</span><br />
+<br />
+<a class='textlink' href='https://paul.buetow.org'>https://paul.buetow.org</a><br />
+<br />
+<span>Thank you very much for your time and this insightful interview into the world of site reliability engineering</span><br />
+<br />
+<span class='quote'>My pleasure, this was fun.</span><br />
+<br />
+<h2 style='display: inline' id='closing-comments'>Closing comments</h2><br />
+<br />
+<span>Dear reader, I hope this conversation with Paul Bütow provided an exciting peak into the world of Site Reliability Engineering. Whether you’re a software developer, data scientist, ML engineer, or manager, reliable systems are always a team effort. Hopefully, you’ve taken some insights or tips from Paul’s experiences for your own team or next project. Thanks for joining us, and best of luck refining your own SRE practices!</span><br />
+<br />
+<span>E-Mail your comments to <span class='inlinecode'>paul@nospam.buetow.org</span> :-)</span><br />
+<br />
+<a class='textlink' href='../'>Back to the main site</a><br />
+<p class="footer">
+Generated with <a href="https://codeberg.org/snonux/gemtexter">Gemtexter 3.0.1-develop</a> |
+served by <a href="https://www.OpenBSD.org">OpenBSD</a>/<a href="https://man.openbsd.org/relayd.8">relayd(8)</a>+<a href="https://man.openbsd.org/httpd.8">httpd(8)</a> |
+<a href="https://foo.zone/site-mirrors.html">Site Mirrors</a>
+</p>
+</body>
+</html>
diff --git a/gemfeed/DRAFT-f3s-kubernetes-with-freebsd-bhyve.html b/gemfeed/DRAFT-f3s-kubernetes-with-freebsd-bhyve.html
index 2db3175d..39d9cf28 100644
--- a/gemfeed/DRAFT-f3s-kubernetes-with-freebsd-bhyve.html
+++ b/gemfeed/DRAFT-f3s-kubernetes-with-freebsd-bhyve.html
@@ -30,12 +30,15 @@ View this page as <a href="https://codeberg.org/snonux/foo.zone/src/branch/conte
<li>⇢ ⇢ <a href='#iso-download'>ISO download</a></li>
<li>⇢ ⇢ <a href='#vm-configuration'>VM configuration</a></li>
<li>⇢ ⇢ <a href='#vm-installation'>VM installation</a></li>
+<li>⇢ ⇢ <a href='#increase-of-the-disk-image'>Increase of the disk image</a></li>
+<li>⇢ ⇢ <a href='#connect-to-vpn'>Connect to VPN</a></li>
+<li>⇢ <a href='#after-install'>After install</a></li>
</ul><br />
<h2 style='display: inline' id='introduction'>Introduction</h2><br />
<br />
<span>In this blog post, we are going to install the Bhyve hypervisor.</span><br />
<br />
-<span>The FreeBSD Bhyve hypervisor is a lightweight, modern hypervisor that enables virtualization on FreeBSD systems. Bhyve&#39;s strengths include its minimal overhead, which allows it to achieve near-native performance for virtual machines. It is designed to be efficient and lightweight, leveraging the capabilities of the FreeBSD operating system for performance and network management. </span><br />
+<span>The FreeBSD Bhyve hypervisor is a lightweight, modern hypervisor that enables virtualization on FreeBSD systems. Bhyve&#39;s strengths include its minimal overhead, which allows it to achieve near-native performance for virtual machines. It is designed to be efficient and lightweight, leveraging the capabilities of the FreeBSD operating system for performance and network management.</span><br />
<br />
<span>Bhyve supports running a variety of guest operating systems, including FreeBSD, Linux, and Windows, on hardware platforms that support hardware virtualization extensions (such as Intel VT-x or AMD-V). In our case, we are going to virtualize Rocky Linux, which later on in this series will be used to run k3s.</span><br />
<br />
@@ -51,15 +54,15 @@ View this page as <a href="https://codeberg.org/snonux/foo.zone/src/branch/conte
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
-<pre>paul@f2:~ % doas pkg install vm-bhyve bhyve-firmware
-paul@f2:~ % doas sysrc vm_enable=YES
+<pre>paul@f0:~ % doas pkg install vm-bhyve bhyve-firmware
+paul@f0:~ % doas sysrc vm_enable=YES
vm_enable: -&gt; YES
-paul@f2:~ % doas sysrc vm_dir=zfs:zroot/bhyve
+paul@f0:~ % doas sysrc vm_dir=zfs:zroot/bhyve
vm_dir: -&gt; zfs:zroot/bhyve
-paul@f2:~ % doas zfs create zroot/bhyve
-paul@f2:~ % doas vm init
-paul@f2:~ % doas vm create public
-paul@f2:~ % doas vm switch add public re0
+paul@f0:~ % doas zfs create zroot/bhyve
+paul@f0:~ % doas vm init
+paul@f0:~ % doas vm switch create public
+paul@f0:~ % doas vm switch add public re0
</pre>
<br />
<span>Bhyve stores all it&#39;s data in the <span class='inlinecode'>/bhyve</span> of the <span class='inlinecode'>zroot</span> ZFS pool:</span><br />
@@ -68,7 +71,7 @@ paul@f2:~ % doas vm switch add public re0
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
-<pre>paul@f2:~ % zfs list | grep bhyve
+<pre>paul@f0:~ % zfs list | grep bhyve
zroot/bhyve <font color="#000000">1</font>.74M 453G <font color="#000000">1</font>.74M /zroot/bhyve
</pre>
<br />
@@ -78,7 +81,7 @@ zroot/bhyve <font color="#000000">1</font>.74M
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
-<pre>paul@f2:~ % doas ln -s /zroot/bhyve/ /bhyve
+<pre>paul@f0:~ % doas ln -s /zroot/bhyve/ /bhyve
</pre>
<br />
@@ -88,7 +91,7 @@ http://www.gnu.org/software/src-highlite -->
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
-<pre>paul@f2:~ % doas vm list
+<pre>paul@f0:~ % doas vm list
NAME DATASTORE LOADER CPU MEMORY VNC AUTO STATE
</pre>
<br />
@@ -102,10 +105,10 @@ NAME DATASTORE LOADER CPU MEMORY VNC AUTO STATE
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
-<pre>paul@f2:~ % doas vm iso \
+<pre>paul@f0:~ % doas vm iso \
https://download.rockylinux.org/pub/rocky/<font color="#000000">9</font>/isos/x86_64/Rocky-<font color="#000000">9.5</font>-x86_64-minimal.iso
/zroot/bhyve/.iso/Rocky-<font color="#000000">9.5</font>-x86_64-minimal.iso <font color="#000000">1808</font> MB <font color="#000000">4780</font> kBps 06m28s
-paul@f2:/bhyve % doas vm create rocky
+paul@f0:/bhyve % doas vm create rocky
</pre>
<h3 style='display: inline' id='vm-configuration'>VM configuration</h3><br />
<br />
@@ -115,7 +118,7 @@ paul@f2:/bhyve % doas vm create rocky
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
-<pre>paul@f2:/bhyve/rocky % cat rocky.conf
+<pre>paul@f0:/bhyve/rocky % cat rocky.conf
loader=<font color="#808080">"bhyveload"</font>
cpu=<font color="#000000">1</font>
memory=256M
@@ -127,7 +130,28 @@ uuid=<font color="#808080">"1c4655ac-c828-11ef-a920-e8ff1ed71ca0"</font>
network0_mac=<font color="#808080">"58:9c:fc:0d:13:3f"</font>
</pre>
<br />
-<span>but in order to make Rocky Linux boot, it...</span><br />
+<span>Whereas the <span class='inlinecode'>uuid</span> and the <span class='inlinecode'>network0_mac</span> differ on each of the 3 hosts.</span><br />
+<br />
+<span>but in order to make Rocky Linux boot it (plus some other adjustments, e.g. as I am intending to run the majority of the workload in the k3s cluster running on those linux VMs, I give them beefy specs like 4 CPU cores and 14GB RAM), I modified it to:</span><br />
+<br />
+<!-- Generator: GNU source-highlight 3.1.9
+by Lorenzo Bettini
+http://www.lorenzobettini.it
+http://www.gnu.org/software/src-highlite -->
+<pre>guest=<font color="#808080">"linux"</font>
+loader=<font color="#808080">"uefi"</font>
+uefi_vars=<font color="#808080">"yes"</font>
+cpu=<font color="#000000">4</font>
+memory=14G
+network0_type=<font color="#808080">"virtio-net"</font>
+network0_switch=<font color="#808080">"public"</font>
+disk0_type=<font color="#808080">"virtio-blk"</font>
+disk0_name=<font color="#808080">"disk0.img"</font>
+graphics=<font color="#808080">"yes"</font>
+graphics_vga=io
+uuid=<font color="#808080">"1c45400b-c828-11ef-8871-e8ff1ed71cac"</font>
+network0_mac=<font color="#808080">"58:9c:fc:0d:13:3f"</font>
+</pre>
<br />
<h3 style='display: inline' id='vm-installation'>VM installation</h3><br />
<br />
@@ -135,7 +159,7 @@ network0_mac=<font color="#808080">"58:9c:fc:0d:13:3f"</font>
by Lorenzo Bettini
http://www.lorenzobettini.it
http://www.gnu.org/software/src-highlite -->
-<pre>paul@f2:~ % doas vm install rocky Rocky-<font color="#000000">9.5</font>-x86_64-minimal.iso
+<pre>paul@f0:~ % doas vm install rocky Rocky-<font color="#000000">9.5</font>-x86_64-minimal.iso
Starting rocky
* found guest <b><u><font color="#000000">in</font></u></b> /zroot/bhyve/rocky
* booting...
@@ -150,6 +174,50 @@ root bhyve <font color="#000000">6079</font> <font color="#000000">8</
<br />
<span>Port 5900 is now also open for VNC connections, so we connect to it with a VNC client and run through the installation dialogs. I&#39;m sure this could be done unattended or more automated, but we have only 3 VMs to install, and the automation doesn&#39;t seem worth it as we are doing it only once.</span><br />
<br />
+<h3 style='display: inline' id='increase-of-the-disk-image'>Increase of the disk image</h3><br />
+<br />
+<span>By default the VMs disk image is only 20G, which is a bit small for my purposes, so I stopped the VMs again and run <span class='inlinecode'>truncate</span> on the image file to enlarge them to 100G, and re-started the installation:</span><br />
+<br />
+<!-- Generator: GNU source-highlight 3.1.9
+by Lorenzo Bettini
+http://www.lorenzobettini.it
+http://www.gnu.org/software/src-highlite -->
+<pre>paul@f0:/bhyve/rocky % doas vm stop rocky
+paul@f0:/bhyve/rocky % doas truncate -s 100G disk0.img
+paul@f0:/bhyve/rocky % doas vm install rocky Rocky-<font color="#000000">9.5</font>-x86_64-minimal.iso
+</pre>
+<br />
+<h3 style='display: inline' id='connect-to-vpn'>Connect to VPN</h3><br />
+<br />
+<span>For the installation, I opened the VPN client on my Fedora laptop (GNOME comes with a simple VPN client) and ran through the base installation for each of the VMs manually. I am sure this could have been automated a bit more, but there were just 3 VMs, and it wasn&#39;t worth the effort. The three VNC addresses of the VMs were: <span class='inlinecode'>vnc://f0:5900</span>, <span class='inlinecode'>vnc://f1:5900</span>, and <span class='inlinecode'>vnc://f0:5900</span>.</span><br />
+<br />
+<span>I mostly selected the default settings (auto partitioning on the 100GB drive and a root user password). After the installation, the VMs were rebooted.</span><br />
+<br />
+<h2 style='display: inline' id='after-install'>After install</h2><br />
+<br />
+<span>After that, I changed the network configuration to be static here as well.</span><br />
+<br />
+<span>As per previous post of this series, the 3 FreeBSD hosts were already in my <span class='inlinecode'>/etc/hosts</span> file:</span><br />
+<br />
+<pre>
+192.168.1.130 f0 f0.lan f0.lan.buetow.org
+192.168.1.131 f1 f1.lan f1.lan.buetow.org
+192.168.1.132 f2 f2.lan f2.lan.buetow.org
+</pre>
+<br />
+<span>For the Rocky VMs I added those:</span><br />
+<br />
+<!-- Generator: GNU source-highlight 3.1.9
+by Lorenzo Bettini
+http://www.lorenzobettini.it
+http://www.gnu.org/software/src-highlite -->
+<pre>cat &lt;&lt;END &gt;&gt;/etc/hosts
+<font color="#000000">192.168</font>.<font color="#000000">1.120</font> r0 r0.lan r0.lan.buetow.org
+<font color="#000000">192.168</font>.<font color="#000000">1.121</font> r1 r1.lan r1.lan.buetow.org
+<font color="#000000">192.168</font>.<font color="#000000">1.122</font> r2 r2.lan r2.lan.buetow.org
+END
+</pre>
+<span>and configured the IPs accordingly on the VMs themselves.</span><br />
<br />
<br />
<span>Other *BSD-related posts:</span><br />
diff --git a/gemfeed/atom.xml b/gemfeed/atom.xml
index 2bded8e7..792c34c9 100644
--- a/gemfeed/atom.xml
+++ b/gemfeed/atom.xml
@@ -1,12 +1,205 @@
<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
- <updated>2024-12-31T19:00:17+02:00</updated>
+ <updated>2025-01-15T00:16:04+02:00</updated>
<title>foo.zone feed</title>
<subtitle>To be in the .zone!</subtitle>
<link href="https://foo.zone/gemfeed/atom.xml" rel="self" />
<link href="https://foo.zone/" />
<id>https://foo.zone/</id>
<entry>
+ <title>Working with an SRE Interview</title>
+ <link href="https://foo.zone/gemfeed/2025-01-15-working-with-an-sre-interview.html" />
+ <id>https://foo.zone/gemfeed/2025-01-15-working-with-an-sre-interview.html</id>
+ <updated>2025-01-15T00:16:04+02:00</updated>
+ <author>
+ <name>Paul Buetow aka snonux</name>
+ <email>paul@dev.buetow.org</email>
+ </author>
+ <summary>I have been interviewed by Florian Buetow about what it's like working with a Site Reliability Engineer from the point of view of a Software Engineer, Data Scientist, and AI Engineer. </summary>
+ <content type="xhtml">
+ <div xmlns="http://www.w3.org/1999/xhtml">
+ <h1 style='display: inline' id='working-with-an-sre-interview'>Working with an SRE Interview</h1><br />
+<br />
+<span>I have been interviewed by Florian Buetow about what it&#39;s like working with a Site Reliability Engineer from the point of view of a Software Engineer, Data Scientist, and AI Engineer. </span><br />
+<br />
+<a class='textlink' href='https://www.cracking-ai-engineering.com/writing/2025/01/12/working-with-an-sre-interview/'>See original interview here</a><br />
+<br />
+<span>Below, I am posting the interview here on my blog as well.</span><br />
+<br />
+<h2 style='display: inline' id='table-of-contents'>Table of Contents</h2><br />
+<br />
+<ul>
+<li><a href='#working-with-an-sre-interview'>Working with an SRE Interview</a></li>
+<li>⇢ <a href='#preamble-'>Preamble </a></li>
+<li>⇢ <a href='#introducing-paul'>Introducing Paul</a></li>
+<li>⇢ <a href='#how-did-you-get-started'>How did you get started?</a></li>
+<li>⇢ <a href='#roles-and-career-progression'>Roles and Career Progression</a></li>
+<li>⇢ <a href='#anecdotes-and-best-practices'>Anecdotes and Best Practices</a></li>
+<li>⇢ <a href='#working-with-different-teams'>Working with Different Teams</a></li>
+<li>⇢ <a href='#using-ai-tools'>Using AI Tools</a></li>
+<li>⇢ <a href='#sre-learning-resources'>SRE Learning Resources</a></li>
+<li>⇢ <a href='#blogging'>Blogging</a></li>
+<li>⇢ <a href='#wrap-up'>Wrap-up</a></li>
+<li>⇢ <a href='#closing-comments'>Closing comments</a></li>
+</ul><br />
+<h2 style='display: inline' id='preamble-'>Preamble </h2><br />
+<br />
+<span>In this insightful interview, Paul Bütow, a Principal Site Reliability Engineer at Mimecast, shares over a decade of experience in the field. Paul highlights the role of an Embedded SRE, emphasizing the importance of automation, observability, and effective incident management. We also focused on the key question of how you can work effectively with an SRE weather you are an individual contributor or a manager, a software engineer or data scientist. And how you can learn more about site reliability engineering.</span><br />
+<br />
+<h2 style='display: inline' id='introducing-paul'>Introducing Paul</h2><br />
+<br />
+<span>Hi Paul, please introduce yourself briefly to the audience. Who are you, what do you do for a living, and where do you work?</span><br />
+<br />
+<span class='quote'>My name is Paul Bütow, I work at Mimecast, and I’m a Principal Site Reliability Engineer there. I’ve been with Mimecast for almost ten years now. The company specializes in email security, including things like archiving, phishing detection, malware protection, and spam filtering.</span><br />
+<br />
+<span>You mentioned that you’re an ‘Embedded SRE.’ What does that mean exactly?</span><br />
+<br />
+<span class='quote'>It means that I’m directly part of the software engineering team, not in a separate Ops department. I ensure that nothing is deployed manually, and everything runs through automation. I also set up monitoring and observability. These are two distinct aspects: monitoring alerts us when something breaks, while observability helps us identify trends. I also create runbooks so we know what to do when specific incidents occur frequently.</span><br />
+<br />
+<span class='quote'>Infrastructure SREs on the other hand handle the foundational setup, like providing the Kubernetes cluster itself or ensuring the operating systems are installed. They don&#39;t work on the application directly but ensure the base infrastructure is there for others to use. This works well when a company has multiple teams that need shared infrastructure.</span><br />
+<br />
+<h2 style='display: inline' id='how-did-you-get-started'>How did you get started?</h2><br />
+<br />
+<span>How did your interest in Linux or FreeBSD start?</span><br />
+<br />
+<span class='quote'>It began during my school days. We had a PC with DOS at home, and I eventually bought Suse Linux 5.3. Shortly after, I discovered FreeBSD because I liked its handbook so much. I wanted to understand exactly how everything worked, so I also tried Linux from Scratch. That involves installing every package manually to gain a better understanding of operating systems.</span><br />
+<br />
+<a class='textlink' href='https://www.FreeBSD.org'>https://www.FreeBSD.org</a><br />
+<a class='textlink' href='https://linuxfromscratch.org/'>https://linuxfromscratch.org/</a><br />
+<br />
+<span>And after school, you pursued computer science, correct?</span><br />
+<br />
+<span class='quote'>Exactly. I wasn’t sure at first whether I wanted to be a software developer or a system administrator. I applied for both and eventually accepted an offer as a Linux system administrator. This was before &#39;SRE&#39; became a buzzword, but much of what I did back then-automation, infrastructure as code, monitoring-is now considered part of the typical SRE role.</span><br />
+<br />
+<h2 style='display: inline' id='roles-and-career-progression'>Roles and Career Progression</h2><br />
+<br />
+<span>Tell us about how you joined Mimecast. When did you fully embrace the SRE role?</span><br />
+<br />
+<span class='quote'>I started as a Linux sysadmin at 1&amp;1. I managed an ad server farm with hundreds of systems and later handled load balancers. Together with an architect, we managed F5 load balancers distributing around 2,000 services, including for portals like web.de and GMX. I also led the operations team technically for a while before moving to London to join Mimecast.</span><br />
+<br />
+<span class='quote'>At Mimecast, the job title was explicitly &#39;Site Reliability Engineer.&#39; The biggest difference was that I was no longer in a separate Ops department but embedded directly within the storage and search backend team. I loved that because we could plan features together-from automation to measurability and observability. Mimecast also operates thousands of physical servers for email archiving, which was fascinating since I already had experience with large distributed systems at 1&amp;1. It was the right step for me because it allowed me to work close to the code while remaining hands-on with infrastructure.</span><br />
+<br />
+<span>What are the differences between SRE, DevOps, SysAdmin, and Architects?</span><br />
+<br />
+<span class='quote'>SREs are like the next step after SysAdmins. A SysAdmin might manually install servers, replace disks, or use simple scripts for automation, while SREs use infrastructure as code and focus on reliability through SLIs, SLOs, and automation. DevOps isn’t really a job-it’s more of a way of working, where developers are involved in operations tasks like setting up CI/CD pipelines or on-call shifts. Architects focus on designing systems and infrastructures, such as load balancers or distributed systems, working alongside SREs to ensure the systems meet the reliability and scalability requirements. The specific responsibilities of each role depend on the company, and there is often overlap. </span><br />
+<br />
+<span>What are the most important reliability lessons you’ve learned so far?</span><br />
+<br />
+<ul>
+<li>Don’t leave SRE aspects as an afterthought. It’s much better to discuss automation, monitoring, SLIs, and SLOs early on. Traditional sysadmins often installed systems manually, but today, we do everything via infrastructure as code-using tools like Terraform or Puppet.</li>
+<li>I also distinguish between monitoring and observability. Monitoring tells us, &#39;The server is down, alarm!&#39; Observability dives deeper, showing trends like increasing latency so we can act proactively.</li>
+<li>SLI, SLO, and SLA are core elements. We focus on what users actually experience-for example, how quickly an email is sent-and set our goals accordingly.</li>
+<li>Runbooks are also crucial. When something goes wrong at night, you don’t want to start from scratch. A runbook outlines how to debug and resolve specific problems, saving time and reducing downtime.</li>
+</ul><br />
+<h2 style='display: inline' id='anecdotes-and-best-practices'>Anecdotes and Best Practices</h2><br />
+<br />
+<span>Runbooks sound very practical. Can you explain how they’re used day-to-day?</span><br />
+<br />
+<span class='quote'>Runbooks are essentially guides for handling specific incidents. For instance, if a service won’t start, the runbook will specify where the logs are and which commands to use. Observability takes it a step further, helping us spot changes early-like rising error rates or latency-so we can address issues before they escalate.</span><br />
+<br />
+<span>When should you decide to put something into a runbook, and when is it unnecessary?</span><br />
+<br />
+<span class='quote'>If an issue happens frequently, it should be documented in a runbook so that anyone, even someone new, can follow the steps to fix it. The idea is that 90% of the common incidents should be covered. For example, if a service is down, the runbook would specify where to find logs, which commands to check, and what actions to take. On the other hand, rare or complex issues, where the resolution depends heavily on context or varies each time, don’t make sense to include in detail. For those, it’s better to focus on general troubleshooting steps. </span><br />
+<br />
+<span>How do you search for and find the correct run