summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
-rw-r--r--gemfeed/2023-08-18-site-reliability-engineering-part-1.html7
-rw-r--r--gemfeed/2023-11-19-site-reliability-engineering-part-2.html13
-rw-r--r--gemfeed/2024-01-09-site-reliability-engineering-part-3.html11
-rw-r--r--gemfeed/2024-09-07-site-reliability-engineering-part-4.html88
-rw-r--r--gemfeed/atom.xml643
-rw-r--r--gemfeed/index.html5
-rw-r--r--index.html7
-rw-r--r--uptime-stats.html2
8 files changed, 217 insertions, 559 deletions
diff --git a/gemfeed/2023-08-18-site-reliability-engineering-part-1.html b/gemfeed/2023-08-18-site-reliability-engineering-part-1.html
index e240e898..3eb827e6 100644
--- a/gemfeed/2023-08-18-site-reliability-engineering-part-1.html
+++ b/gemfeed/2023-08-18-site-reliability-engineering-part-1.html
@@ -15,8 +15,9 @@
<span>Being a Site Reliability Engineer (SRE) is like stepping into a lively, ever-evolving universe. The world of SRE mixes together different tech, a unique culture, and a whole lot of determination. It’s one of the toughest but most exciting jobs out there. There&#39;s zero chance of getting bored because there&#39;s always a fresh challenge to tackle and new technology to play around with. It&#39;s not just about the tech side of things either; it&#39;s heavily rooted in communication, collaboration, and teamwork. As someone currently working as an SRE, I’m here to break it all down for you in this blog series. Let&#39;s dive into what SRE is really all about!</span><br />
<br />
<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture (You are currently reading this)</a><br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE</a><br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</a><br />
<br />
<pre>
▓▓▓▓░░
@@ -62,7 +63,7 @@ DC on fire:
<br />
<span>Continue with the second part of this series:</span><br />
<br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
<br />
<span>E-Mail your comments to <span class='inlinecode'>paul@nospam.buetow.org</span> :-)</span><br />
<br />
diff --git a/gemfeed/2023-11-19-site-reliability-engineering-part-2.html b/gemfeed/2023-11-19-site-reliability-engineering-part-2.html
index 9747d38d..4011f03e 100644
--- a/gemfeed/2023-11-19-site-reliability-engineering-part-2.html
+++ b/gemfeed/2023-11-19-site-reliability-engineering-part-2.html
@@ -2,21 +2,22 @@
<html xmlns="http://www.w3.org/1999/xhtml" lang="en" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
-<title>Site Reliability Engineering - Part 2: Operational Balance in SRE</title>
+<title>Site Reliability Engineering - Part 2: Operational Balance</title>
<link rel="shortcut icon" type="image/gif" href="/favicon.ico" />
<link rel="stylesheet" href="../style.css" />
<link rel="stylesheet" href="style-override.css" />
</head>
<body>
-<h1 style='display: inline' id='site-reliability-engineering---part-2-operational-balance-in-sre'>Site Reliability Engineering - Part 2: Operational Balance in SRE</h1><br />
+<h1 style='display: inline' id='site-reliability-engineering---part-2-operational-balance'>Site Reliability Engineering - Part 2: Operational Balance</h1><br />
<br />
<span class='quote'>Published at 2023-11-19T00:18:18+03:00</span><br />
<br />
<span>This is the second part of my Site Reliability Engineering (SRE) series. I am currently employed as a Site Reliability Engineer and will try to share what SRE is about in this blog series.</span><br />
<br />
<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture</a><br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE (You are currently reading this)</a><br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance (You are currently reading this)</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</a><br />
<br />
<pre>
⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣠⣾⣷⣄⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
@@ -33,7 +34,7 @@
⠀⠀⠀⠀⠀⠀⠴⠶⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠶⠦⠀⠀
</pre>
<br />
-<h2 style='display: inline' id='operational-balance-in-sre-striking-the-right-balance-between-reliability-and-speed'>Operational Balance in SRE: Striking the Right Balance Between Reliability and Speed</h2><br />
+<h2 style='display: inline' id='striking-the-right-balance-between-reliability-and-speed'>Striking the Right Balance Between Reliability and Speed</h2><br />
<br />
<span>Site Reliability Engineering is more than just a bunch of best practices or methods. It&#39;s a guiding light for engineering teams, helping them navigate the tricky waters of modern software development and system management.</span><br />
<span>In the world of software production, there are two big forces that often clash: the push for fast feature releases (velocity) and the need for reliable systems. Traditionally, moving faster meant more risk. SRE helps balance these opposing goals with things like error budgets and SLIs/SLOs. These tools give teams a clear way to measure how much they can push changes without hurting system health. So, the error budget becomes a balancing act, helping teams trade off between innovation and reliability.</span><br />
@@ -52,7 +53,7 @@
<br />
<span>Continue with the third part of this series:</span><br />
<br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
<br />
<span>E-Mail your comments to <span class='inlinecode'>paul@nospam.buetow.org</span> :-)</span><br />
<br />
diff --git a/gemfeed/2024-01-09-site-reliability-engineering-part-3.html b/gemfeed/2024-01-09-site-reliability-engineering-part-3.html
index 87f12c86..a4a50770 100644
--- a/gemfeed/2024-01-09-site-reliability-engineering-part-3.html
+++ b/gemfeed/2024-01-09-site-reliability-engineering-part-3.html
@@ -2,21 +2,22 @@
<html xmlns="http://www.w3.org/1999/xhtml" lang="en" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
-<title>Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</title>
+<title>Site Reliability Engineering - Part 3: On-Call Culture</title>
<link rel="shortcut icon" type="image/gif" href="/favicon.ico" />
<link rel="stylesheet" href="../style.css" />
<link rel="stylesheet" href="style-override.css" />
</head>
<body>
-<h1 style='display: inline' id='site-reliability-engineering---part-3-on-call-culture-and-the-human-side'>Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</h1><br />
+<h1 style='display: inline' id='site-reliability-engineering---part-3-on-call-culture'>Site Reliability Engineering - Part 3: On-Call Culture</h1><br />
<br />
<span class='quote'>Published at 2024-01-09T18:35:48+02:00</span><br />
<br />
<span>Welcome to Part 3 of my Site Reliability Engineering (SRE) series. I&#39;m currently working as a Site Reliability Engineer, and I’m here to share what SRE is all about in this blog series.</span><br />
<br />
<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture</a><br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE</a><br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side (You are currently reading this)</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture (You are currently reading this)</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</a><br />
<br />
<pre>
..--""""----..
@@ -44,7 +45,7 @@
</pre>
<br />
-<h2 style='display: inline' id='on-call-culture-and-the-human-side-putting-well-being-first-in-the-world-of-reliability'>On-Call Culture and the Human Side: Putting Well-being First in the World of Reliability</h2><br />
+<h2 style='display: inline' id='putting-well-being-first'>Putting Well-being First</h2><br />
<br />
<span>Site Reliability Engineering is all about keeping systems reliable, but we often forget how important the human side is. A healthy on-call culture is just as crucial as any technical fix. The well-being of the engineers really matters.</span><br />
<br />
diff --git a/gemfeed/2024-09-07-site-reliability-engineering-part-4.html b/gemfeed/2024-09-07-site-reliability-engineering-part-4.html
new file mode 100644
index 00000000..d3ff4e72
--- /dev/null
+++ b/gemfeed/2024-09-07-site-reliability-engineering-part-4.html
@@ -0,0 +1,88 @@
+<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
+<html xmlns="http://www.w3.org/1999/xhtml" lang="en" xml:lang="en">
+<head>
+<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
+<title>Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</title>
+<link rel="shortcut icon" type="image/gif" href="/favicon.ico" />
+<link rel="stylesheet" href="../style.css" />
+<link rel="stylesheet" href="style-override.css" />
+</head>
+<body>
+<h1 style='display: inline' id='site-reliability-engineering---part-4-onboarding-for-on-call-engineers'>Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</h1><br />
+<br />
+<span class='quote'>Published at 2024-09-07T16:27:58+03:00</span><br />
+<br />
+<span>Welcome to Part 4 of my Site Reliability Engineering (SRE) series. I&#39;m currently working as a Site Reliability Engineer, and I’m here to share what SRE is all about in this blog series.</span><br />
+<br />
+<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers (You are currently reading this)</a><br />
+<br />
+<pre>
+ __..._ _...__
+ _..-" `Y` "-._
+ \ Once upon | /
+ \\ a time..| //
+ \\\ | ///
+ \\\ _..---.|.---.._ ///
+jgs \\`_..---.Y.---.._`//
+</pre>
+<br />
+<span>This time, I want to share some tips on how to onboard software engineers, QA engineers, and Site Reliability Engineers (SREs) to the primary on-call rotation. Traditionally, onboarding might take half a year (depending on the complexity of the infrastructure), but with a bit of strategy and structured sessions, we&#39;ve managed to reduce it to just six weeks per person. Let&#39;s dive in!</span><br />
+<br />
+<h2 style='display: inline' id='setting-the-scene-tier-1-on-call-rotation'>Setting the Scene: Tier-1 On-Call Rotation</h2><br />
+<br />
+<span>First things first, let&#39;s talk about Tier-1. This is where the magic begins. Tier-1 covers over 80% of the common on-call cases and is the perfect breeding ground for new on-call engineers to get their feet wet. It&#39;s designed to be manageable training ground.</span><br />
+<br />
+<h3 style='display: inline' id='why-tier-1'>Why Tier-1?</h3><br />
+<br />
+<ul>
+<li>Easy to Understand: Every on-call engineer should be familiar with Tier-1 tasks. </li>
+<li>Training Ground: This is where engineers start their on-call career. It&#39;s purposefully kept simple so that it&#39;s not overwhelming right off the bat.</li>
+<li>Runbook/recipe driven: Every alert is attached to a comprehensive runbook, making it easy for every engineer to follow.</li>
+</ul><br />
+<h2 style='display: inline' id='onboarding-process-from-6-months-to-6-weeks'>Onboarding Process: From 6 Months to 6 Weeks</h2><br />
+<br />
+<span>So how did we cut down the onboarding time so drastically? Here’s the breakdown of our process:</span><br />
+<br />
+<span>Knowledge Transfer (KT) Sessions: We kicked things off with more than 10 KT sessions, complete with video recordings. These sessions are comprehensive and cover everything from the basics to some more advanced topics. The recorded sessions mean that new engineers can revisit them anytime they need a refresher.</span><br />
+<br />
+<span>Shadowing Sessions: Each new engineer undergoes two on-call week shadowing sessions. This hands-on experience is invaluable. They get to see real-time incident handling and resolution, gaining practical knowledge that&#39;s hard to get from just reading docs.</span><br />
+<br />
+<span>Comprehensive Runbooks: We created 64 runbooks (by the time writing this probably more than 100) that are composable like Lego bricks. Each runbook covers a specific scenario and guides the engineer step-by-step to resolution. Pairing these with monitoring alerts linked directly to Confluence docs, and from there to the respective runbooks, ensures every alert can be navigated with ease (well, there are always exceptions to the rule...).</span><br />
+<br />
+<span>Self-Sufficiency &amp; Confidence Building: With all these resources at their fingertips, our on-call engineers become self-sufficient for most of the common issues they&#39;ll face (new starters can now handle around 80% of the most common issue after 6 weeks they had joined the company). This boosts their confidence and ensures they can handle Tier-1 incidents independently.</span><br />
+<br />
+<span>Documentation and Feedback Loop: Continuous improvement is key. We regularly update our documentation based on feedback from the engineers. This makes our process even more robust and user-friendly.</span><br />
+<br />
+<h2 style='display: inline' id='it-s-all-about-the-tiers'>It&#39;s All About the Tiers</h2><br />
+<br />
+<span>Let’s briefly touch on the Tier levels:</span><br />
+<br />
+<ul>
+<li>Tier 1: Easy and foundational tasks. Perfect for getting new engineers started. This covers around 80% of all on-call cases we face. This is what we trained on.</li>
+<li>Tier 2: Slightly more complex, requiring more background knowledge. We trained on some of the topics but not all.</li>
+<li>Tier 3: Requires a good understanding of the platform/architecture. Likely needs KT sessions with domain experts.</li>
+<li>Tier DE (Domain Expert): The heavy hitters. Domain experts are required for these tasks. </li>
+</ul><br />
+<h3 style='display: inline' id='growing-into-higher-tiers'>Growing into Higher Tiers</h3><br />
+<br />
+<span>From Tier-1, engineers naturally grow into Tier-2 and beyond. The structured training and gradual increase in complexity help ensure a smooth transition as they gain experience and confidence. The key here is that engineers stay curous and engaged in the on-call, so that they always keep learning.</span><br />
+<br />
+<h2 style='display: inline' id='keeping-runbooks-up-to-date'>Keeping Runbooks Up to Date</h2><br />
+<br />
+<span>It is important that runbooks are not a "project to be finished"; runbooks have to be maintained and updated over time. Sections may change, new runbooks need to be added, and old ones can be deleted. So the acceptance criteria of an on-call shift would not just be reacting to alerts and incidents, but also reviewing and updating the current runbooks.</span><br />
+<br />
+<h2 style='display: inline' id='conclusion'>Conclusion</h2><br />
+<br />
+<span>By structuring the onboarding process with KT sessions, shadowing, comprehensive runbooks, and a feedback loop, we&#39;ve been able to fast-track the process from six months to just six weeks. This not only prepares our engineers for the on-call rotation quicker but also ensures they&#39;re confident and capable when handling incidents.</span><br />
+<br />
+<span>If you&#39;re looking to optimize your on-call onboarding process, these strategies could be your ticket to a more efficient and effective transition. Happy on-calling!</span><br />
+<p class="footer">
+Generated by <a href="https://codeberg.org/snonux/gemtexter">Gemtexter 3.0.0-develop</a> |
+served by <a href="https://www.OpenBSD.org">OpenBSD</a>/<a href="https://man.openbsd.org/httpd.8">httpd(8)</a> |
+<a href="https://foo.zone/site-mirrors.html">Site Mirrors</a>
+</p>
+</body>
+</html>
diff --git a/gemfeed/atom.xml b/gemfeed/atom.xml
index 6b7c59a3..154c409c 100644
--- a/gemfeed/atom.xml
+++ b/gemfeed/atom.xml
@@ -1,12 +1,98 @@
<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
- <updated>2024-09-07T16:11:00+03:00</updated>
+ <updated>2024-09-07T16:32:50+03:00</updated>
<title>foo.zone feed</title>
<subtitle>To be in the .zone!</subtitle>
<link href="https://foo.zone/gemfeed/atom.xml" rel="self" />
<link href="https://foo.zone/" />
<id>https://foo.zone/</id>
<entry>
+ <title>Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</title>
+ <link href="https://foo.zone/gemfeed/2024-09-07-site-reliability-engineering-part-4.html" />
+ <id>https://foo.zone/gemfeed/2024-09-07-site-reliability-engineering-part-4.html</id>
+ <updated>2024-09-07T16:27:58+03:00</updated>
+ <author>
+ <name>Paul Buetow aka snonux</name>
+ <email>paul@dev.buetow.org</email>
+ </author>
+ <summary>Welcome to Part 4 of my Site Reliability Engineering (SRE) series. I'm currently working as a Site Reliability Engineer, and I’m here to share what SRE is all about in this blog series.</summary>
+ <content type="xhtml">
+ <div xmlns="http://www.w3.org/1999/xhtml">
+ <h1 style='display: inline' id='site-reliability-engineering---part-4-onboarding-for-on-call-engineers'>Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</h1><br />
+<br />
+<span class='quote'>Published at 2024-09-07T16:27:58+03:00</span><br />
+<br />
+<span>Welcome to Part 4 of my Site Reliability Engineering (SRE) series. I&#39;m currently working as a Site Reliability Engineer, and I’m here to share what SRE is all about in this blog series.</span><br />
+<br />
+<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers (You are currently reading this)</a><br />
+<br />
+<pre>
+ __..._ _...__
+ _..-" `Y` "-._
+ \ Once upon | /
+ \\ a time..| //
+ \\\ | ///
+ \\\ _..---.|.---.._ ///
+jgs \\`_..---.Y.---.._`//
+</pre>
+<br />
+<span>This time, I want to share some tips on how to onboard software engineers, QA engineers, and Site Reliability Engineers (SREs) to the primary on-call rotation. Traditionally, onboarding might take half a year (depending on the complexity of the infrastructure), but with a bit of strategy and structured sessions, we&#39;ve managed to reduce it to just six weeks per person. Let&#39;s dive in!</span><br />
+<br />
+<h2 style='display: inline' id='setting-the-scene-tier-1-on-call-rotation'>Setting the Scene: Tier-1 On-Call Rotation</h2><br />
+<br />
+<span>First things first, let&#39;s talk about Tier-1. This is where the magic begins. Tier-1 covers over 80% of the common on-call cases and is the perfect breeding ground for new on-call engineers to get their feet wet. It&#39;s designed to be manageable training ground.</span><br />
+<br />
+<h3 style='display: inline' id='why-tier-1'>Why Tier-1?</h3><br />
+<br />
+<ul>
+<li>Easy to Understand: Every on-call engineer should be familiar with Tier-1 tasks. </li>
+<li>Training Ground: This is where engineers start their on-call career. It&#39;s purposefully kept simple so that it&#39;s not overwhelming right off the bat.</li>
+<li>Runbook/recipe driven: Every alert is attached to a comprehensive runbook, making it easy for every engineer to follow.</li>
+</ul><br />
+<h2 style='display: inline' id='onboarding-process-from-6-months-to-6-weeks'>Onboarding Process: From 6 Months to 6 Weeks</h2><br />
+<br />
+<span>So how did we cut down the onboarding time so drastically? Here’s the breakdown of our process:</span><br />
+<br />
+<span>Knowledge Transfer (KT) Sessions: We kicked things off with more than 10 KT sessions, complete with video recordings. These sessions are comprehensive and cover everything from the basics to some more advanced topics. The recorded sessions mean that new engineers can revisit them anytime they need a refresher.</span><br />
+<br />
+<span>Shadowing Sessions: Each new engineer undergoes two on-call week shadowing sessions. This hands-on experience is invaluable. They get to see real-time incident handling and resolution, gaining practical knowledge that&#39;s hard to get from just reading docs.</span><br />
+<br />
+<span>Comprehensive Runbooks: We created 64 runbooks (by the time writing this probably more than 100) that are composable like Lego bricks. Each runbook covers a specific scenario and guides the engineer step-by-step to resolution. Pairing these with monitoring alerts linked directly to Confluence docs, and from there to the respective runbooks, ensures every alert can be navigated with ease (well, there are always exceptions to the rule...).</span><br />
+<br />
+<span>Self-Sufficiency &amp; Confidence Building: With all these resources at their fingertips, our on-call engineers become self-sufficient for most of the common issues they&#39;ll face (new starters can now handle around 80% of the most common issue after 6 weeks they had joined the company). This boosts their confidence and ensures they can handle Tier-1 incidents independently.</span><br />
+<br />
+<span>Documentation and Feedback Loop: Continuous improvement is key. We regularly update our documentation based on feedback from the engineers. This makes our process even more robust and user-friendly.</span><br />
+<br />
+<h2 style='display: inline' id='it-s-all-about-the-tiers'>It&#39;s All About the Tiers</h2><br />
+<br />
+<span>Let’s briefly touch on the Tier levels:</span><br />
+<br />
+<ul>
+<li>Tier 1: Easy and foundational tasks. Perfect for getting new engineers started. This covers around 80% of all on-call cases we face. This is what we trained on.</li>
+<li>Tier 2: Slightly more complex, requiring more background knowledge. We trained on some of the topics but not all.</li>
+<li>Tier 3: Requires a good understanding of the platform/architecture. Likely needs KT sessions with domain experts.</li>
+<li>Tier DE (Domain Expert): The heavy hitters. Domain experts are required for these tasks. </li>
+</ul><br />
+<h3 style='display: inline' id='growing-into-higher-tiers'>Growing into Higher Tiers</h3><br />
+<br />
+<span>From Tier-1, engineers naturally grow into Tier-2 and beyond. The structured training and gradual increase in complexity help ensure a smooth transition as they gain experience and confidence. The key here is that engineers stay curous and engaged in the on-call, so that they always keep learning.</span><br />
+<br />
+<h2 style='display: inline' id='keeping-runbooks-up-to-date'>Keeping Runbooks Up to Date</h2><br />
+<br />
+<span>It is important that runbooks are not a "project to be finished"; runbooks have to be maintained and updated over time. Sections may change, new runbooks need to be added, and old ones can be deleted. So the acceptance criteria of an on-call shift would not just be reacting to alerts and incidents, but also reviewing and updating the current runbooks.</span><br />
+<br />
+<h2 style='display: inline' id='conclusion'>Conclusion</h2><br />
+<br />
+<span>By structuring the onboarding process with KT sessions, shadowing, comprehensive runbooks, and a feedback loop, we&#39;ve been able to fast-track the process from six months to just six weeks. This not only prepares our engineers for the on-call rotation quicker but also ensures they&#39;re confident and capable when handling incidents.</span><br />
+<br />
+<span>If you&#39;re looking to optimize your on-call onboarding process, these strategies could be your ticket to a more efficient and effective transition. Happy on-calling!</span><br />
+ </div>
+ </content>
+ </entry>
+ <entry>
<title>Projects I support</title>
<link href="https://foo.zone/gemfeed/2024-09-07-projects-i-support.html" />
<id>https://foo.zone/gemfeed/2024-09-07-projects-i-support.html</id>
@@ -2473,7 +2559,7 @@ http://www.gnu.org/software/src-highlite -->
</content>
</entry>
<entry>
- <title>Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</title>
+ <title>Site Reliability Engineering - Part 3: On-Call Culture</title>
<link href="https://foo.zone/gemfeed/2024-01-09-site-reliability-engineering-part-3.html" />
<id>https://foo.zone/gemfeed/2024-01-09-site-reliability-engineering-part-3.html</id>
<updated>2024-01-09T18:35:48+02:00</updated>
@@ -2484,15 +2570,16 @@ http://www.gnu.org/software/src-highlite -->
<summary>Welcome to Part 3 of my Site Reliability Engineering (SRE) series. I'm currently working as a Site Reliability Engineer, and I’m here to share what SRE is all about in this blog series.</summary>
<content type="xhtml">
<div xmlns="http://www.w3.org/1999/xhtml">
- <h1 style='display: inline' id='site-reliability-engineering---part-3-on-call-culture-and-the-human-side'>Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</h1><br />
+ <h1 style='display: inline' id='site-reliability-engineering---part-3-on-call-culture'>Site Reliability Engineering - Part 3: On-Call Culture</h1><br />
<br />
<span class='quote'>Published at 2024-01-09T18:35:48+02:00</span><br />
<br />
<span>Welcome to Part 3 of my Site Reliability Engineering (SRE) series. I&#39;m currently working as a Site Reliability Engineer, and I’m here to share what SRE is all about in this blog series.</span><br />
<br />
<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture</a><br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE</a><br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side (You are currently reading this)</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture (You are currently reading this)</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</a><br />
<br />
<pre>
..--""""----..
@@ -2520,7 +2607,7 @@ http://www.gnu.org/software/src-highlite -->
</pre>
<br />
-<h2 style='display: inline' id='on-call-culture-and-the-human-side-putting-well-being-first-in-the-world-of-reliability'>On-Call Culture and the Human Side: Putting Well-being First in the World of Reliability</h2><br />
+<h2 style='display: inline' id='putting-well-being-first'>Putting Well-being First</h2><br />
<br />
<span>Site Reliability Engineering is all about keeping systems reliable, but we often forget how important the human side is. A healthy on-call culture is just as crucial as any technical fix. The well-being of the engineers really matters.</span><br />
<br />
@@ -2977,7 +3064,7 @@ echo baz
</content>
</entry>
<entry>
- <title>Site Reliability Engineering - Part 2: Operational Balance in SRE</title>
+ <title>Site Reliability Engineering - Part 2: Operational Balance</title>
<link href="https://foo.zone/gemfeed/2023-11-19-site-reliability-engineering-part-2.html" />
<id>https://foo.zone/gemfeed/2023-11-19-site-reliability-engineering-part-2.html</id>
<updated>2023-11-19T00:18:18+03:00</updated>
@@ -2988,15 +3075,16 @@ echo baz
<summary>This is the second part of my Site Reliability Engineering (SRE) series. I am currently employed as a Site Reliability Engineer and will try to share what SRE is about in this blog series.</summary>
<content type="xhtml">
<div xmlns="http://www.w3.org/1999/xhtml">
- <h1 style='display: inline' id='site-reliability-engineering---part-2-operational-balance-in-sre'>Site Reliability Engineering - Part 2: Operational Balance in SRE</h1><br />
+ <h1 style='display: inline' id='site-reliability-engineering---part-2-operational-balance'>Site Reliability Engineering - Part 2: Operational Balance</h1><br />
<br />
<span class='quote'>Published at 2023-11-19T00:18:18+03:00</span><br />
<br />
<span>This is the second part of my Site Reliability Engineering (SRE) series. I am currently employed as a Site Reliability Engineer and will try to share what SRE is about in this blog series.</span><br />
<br />
<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture</a><br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE (You are currently reading this)</a><br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance (You are currently reading this)</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</a><br />
<br />
<pre>
⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣠⣾⣷⣄⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
@@ -3013,7 +3101,7 @@ echo baz
⠀⠀⠀⠀⠀⠀⠴⠶⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠶⠦⠀⠀
</pre>
<br />
-<h2 style='display: inline' id='operational-balance-in-sre-striking-the-right-balance-between-reliability-and-speed'>Operational Balance in SRE: Striking the Right Balance Between Reliability and Speed</h2><br />
+<h2 style='display: inline' id='striking-the-right-balance-between-reliability-and-speed'>Striking the Right Balance Between Reliability and Speed</h2><br />
<br />
<span>Site Reliability Engineering is more than just a bunch of best practices or methods. It&#39;s a guiding light for engineering teams, helping them navigate the tricky waters of modern software development and system management.</span><br />
<span>In the world of software production, there are two big forces that often clash: the push for fast feature releases (velocity) and the need for reliable systems. Traditionally, moving faster meant more risk. SRE helps balance these opposing goals with things like error budgets and SLIs/SLOs. These tools give teams a clear way to measure how much they can push changes without hurting system health. So, the error budget becomes a balancing act, helping teams trade off between innovation and reliability.</span><br />
@@ -3032,7 +3120,7 @@ echo baz
<br />
<span>Continue with the third part of this series:</span><br />
<br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
<br />
<span>E-Mail your comments to <span class='inlinecode'>paul@nospam.buetow.org</span> :-)</span><br />
<br />
@@ -3849,8 +3937,9 @@ http://www.gnu.org/software/src-highlite -->
<span>Being a Site Reliability Engineer (SRE) is like stepping into a lively, ever-evolving universe. The world of SRE mixes together different tech, a unique culture, and a whole lot of determination. It’s one of the toughest but most exciting jobs out there. There&#39;s zero chance of getting bored because there&#39;s always a fresh challenge to tackle and new technology to play around with. It&#39;s not just about the tech side of things either; it&#39;s heavily rooted in communication, collaboration, and teamwork. As someone currently working as an SRE, I’m here to break it all down for you in this blog series. Let&#39;s dive into what SRE is really all about!</span><br />
<br />
<a class='textlink' href='./2023-08-18-site-reliability-engineering-part-1.html'>2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture (You are currently reading this)</a><br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE</a><br />
-<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture and the Human Side</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
+<a class='textlink' href='./2024-01-09-site-reliability-engineering-part-3.html'>2024-01-09 Site Reliability Engineering - Part 3: On-Call Culture</a><br />
+<a class='textlink' href='./2024-09-07-site-reliability-engineering-part-4.html'>2024-09-07 Site Reliability Engineering - Part 4: Onboarding for On-Call Engineers</a><br />
<br />
<pre>
▓▓▓▓░░
@@ -3896,7 +3985,7 @@ DC on fire:
<br />
<span>Continue with the second part of this series:</span><br />
<br />
-<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance in SRE</a><br />
+<a class='textlink' href='./2023-11-19-site-reliability-engineering-part-2.html'>2023-11-19 Site Reliability Engineering - Part 2: Operational Balance</a><br />
<br />
<span>E-Mail your comments to <span class='inlinecode'>paul@nospam.buetow.org</span> :-)</span><br />
<br />
@@ -9081,528 +9170,4 @@ GNU/kFreeBSD rhea.buetow.org 8.0-RELEASE-p5 FreeBSD 8.0-RELEASE-p5 #2: Sat Nov 2
</div>
</content>
</entry>
- <entry>
- <title>Bash Golf Part 2</title>
- <link href="https://foo.zone/gemfeed/2022-01-01-bash-golf-part-2.html" />
- <id>https://foo.zone/gemfeed/2022-01-01-bash-golf-part-2.html</id>
- <updated>2022-01-01T23:36:15+00:00</updated>
- <author>
- <name>Paul Buetow aka snonux</name>
- <email>paul@dev.buetow.org</email>
- </author>
- <summary>This is the second blog post about my Bash Golf series. This series is random Bash tips, tricks and weirdnesses I came across. It's a collection of smaller articles I wrote in an older (in German language) blog, which I translated and refreshed with some new content.</summary>
- <content type="xhtml">
- <div xmlns="http://www.w3.org/1999/xhtml">
- <h1 style='display: inline' id='bash-golf-part-2'>Bash Golf Part 2</h1><br />
-<br />
-<span class='quote'>Published at 2022-01-01T23:36:15+00:00; Updated at 2022-01-05</span><br />
-<br />
-<span>This is the second blog post about my Bash Golf series. This series is random Bash tips, tricks and weirdnesses I came across. It&#39;s a collection of smaller articles I wrote in an older (in German language) blog, which I translated and refreshed with some new content.</span><br />
-<br />
-<a class='textlink' href='./2021-11-29-bash-golf-part-1.html'>2021-11-29 Bash Golf Part 1</a><br />
-<a class='textlink' href='./2022-01-01-bash-golf-part-2.html'>2022-01-01 Bash Golf Part 2 (You are currently reading this)</a><br />
-<a class='textlink' href='./2023-12-10-bash-golf-part-3.html'>2023-12-10 Bash Golf Part 3</a><br />
-<br />
-<pre>
- &#39;\ &#39;\ . . |&gt;18&gt;&gt;
- \ \ . &#39; . |
- O&gt;&gt; O&gt;&gt; . &#39;o |
- \ .\. .. . |
- /\ . /\ . . |
- / / . / / .&#39; . |
-jgs^^^^^^^`^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
- Art by Joan Stark, mod. by Paul Buetow
-</pre>
-<br />
-<h2 style='display: inline' id='table-of-contents'>Table of Contents</h2><br />
-<br />
-<ul>
-<li><a href='#bash-golf-part-2'>Bash Golf Part 2</a></li>
-<li>⇢ <a href='#redirection'>Redirection</a></li>
-<li>⇢ <a href='#here'>HERE</a></li>
-<li>⇢ <a href='#random'>RANDOM</a></li>
-<li>⇢ <a href='#set--x-and-set--e-and-pipefile'>set -x and set -e and pipefile</a></li>
-<li>⇢ ⇢ <a href='#-x'>-x</a></li>
-<li>⇢ ⇢ <a href='#-e'>-e</a></li>
-<li>⇢ ⇢ <a href='#pipefail'>pipefail</a></li>
-</ul><br />
-<h2 style='display: inline' id='redirection'>Redirection</h2><br />
-<br />
-<span>Let&#39;s have a closer look at Bash redirection. As you might already know that there are 3 standard file descriptors:</span><br />
-<br />
-<ul>
-<li>0 aka stdin (standard input)</li>
-<li>1 aka stdout (standard output)</li>
-<li>2 aka stderr (standard error output)</li>
-</ul><br />
-<span>These are most certainly the ones you are using on regular basis. "/proc/self/fd" lists all file descriptors which are open by the current process (in this case: the current Bash shell itself):</span>&