From 5162ceb2cd04cc644015c79a75591b4d3818cd38 Mon Sep 17 00:00:00 2001 From: Paul Buetow Date: Sat, 19 Aug 2023 11:28:07 +0300 Subject: Update content for html --- ...17-career-guide-and-soft-skills-book-notes.html | 4 +- ...-08-18-site-reliability-engineering-part-1.html | 2 +- gemfeed/DRAFT-site-reliability-engineering.html | 2 +- gemfeed/atom.xml | 12973 +++++++++---------- gemfeed/index.html | 2 +- 5 files changed, 6197 insertions(+), 6786 deletions(-) (limited to 'gemfeed') diff --git a/gemfeed/2023-07-17-career-guide-and-soft-skills-book-notes.html b/gemfeed/2023-07-17-career-guide-and-soft-skills-book-notes.html index 501c5b15..e6468978 100644 --- a/gemfeed/2023-07-17-career-guide-and-soft-skills-book-notes.html +++ b/gemfeed/2023-07-17-career-guide-and-soft-skills-book-notes.html @@ -2,13 +2,13 @@ -'Software Developmers Career Guide %%TITLE%% Soft Skills' book notes +'Software Developmers Career Guide and Soft Skills' book notes -

"Software Developmers Career Guide & Soft Skills" book notes


+

"Software Developmers Career Guide and Soft Skills" book notes



Published at 2023-07-17T04:56:20+03:00

diff --git a/gemfeed/2023-08-18-site-reliability-engineering-part-1.html b/gemfeed/2023-08-18-site-reliability-engineering-part-1.html index 03c3f334..05ef4ba2 100644 --- a/gemfeed/2023-08-18-site-reliability-engineering-part-1.html +++ b/gemfeed/2023-08-18-site-reliability-engineering-part-1.html @@ -41,7 +41,7 @@ DC on fire:

SRE and Organizational Culture: Navigating the Nexus



-At the heart of SRE lies the proactive mindset of 'prevention over cure'. Traditional IT models focused predominantly on reactive solutions, but SRE mandates a shift towards foresight. By adopting Service Level Indicators (SLIs) and Service Level Objectives (SLOs), teams are equipped with clear metrics and goals that guide them toward ensuring reliability and user satisfaction. However, these aren't mere numbers. They reflect an organisational culture prioritising user experience and constant system alignment with user needs.
+At the heart of SRE lies the proactive mindset of "prevention over cure". Traditional IT models focused predominantly on reactive solutions, but SRE mandates a shift towards foresight. By adopting Service Level Indicators (SLIs) and Service Level Objectives (SLOs), teams are equipped with clear metrics and goals that guide them toward ensuring reliability and user satisfaction. However, these aren't mere numbers. They reflect an organisational culture prioritising user experience and constant system alignment with user needs.

Another defining SRE concept is the "error budget". This ingenious framework accepts that no system is flawless. Failures are inevitable. However, instead of being punitive, the culture here is to accept, learn, and iterate. By providing teams with a "budget" for errors, organisations foster an environment where innovation is encouraged, and failures are viewed as learning opportunities.

diff --git a/gemfeed/DRAFT-site-reliability-engineering.html b/gemfeed/DRAFT-site-reliability-engineering.html index c8fba6ad..f7c9949c 100644 --- a/gemfeed/DRAFT-site-reliability-engineering.html +++ b/gemfeed/DRAFT-site-reliability-engineering.html @@ -10,7 +10,7 @@

On-Call Culture and the Human Aspect: Prioritising Well-being in the Realm of Reliability



-Site Reliability Engineering is synonymous with ensuring system reliability, but the human factor is an often-underestimated component of this discipline. It is evident that fostering a healthy on-call culture is as critical as any technical solution. In the world of constant alerts, pages, and incident management, the well-being of the engineers becomes paramount.
+Site Reliability Engineering is synonymous with ensuring system reliability, but the human factor is an often-underestimated component of this discipline. Fostering a healthy on-call culture is as critical as any technical solution. In the world of constant alerts, pages, and incident management, the well-being of the engineers becomes paramount.

Firstly, a healthy on-call rotation is about more than just managing and responding to incidents. It's about the entire ecosystem that supports this practice. Establishing happy and healthy on-call rotations is akin to possessing a superpower. This involves reducing pain points, offering mentorship, rapid iteration, and ensuring that engineers have the right tools and processes. It acknowledges that while systems are crucial, the engineers who maintain them are invaluable.

diff --git a/gemfeed/atom.xml b/gemfeed/atom.xml index ffb78621..db0678e4 100644 --- a/gemfeed/atom.xml +++ b/gemfeed/atom.xml @@ -1,281 +1,267 @@ - 2023-07-12T07:37:47+03:00 + 2023-08-19T11:27:51+03:00 foo.zone feed To be in the .zone! https://foo.zone/ - KISS server monitoring with Gogios - - https://foo.zone/gemfeed/2023-06-01-kiss-server-monitoring-with-gogios.html - 2023-06-01T21:10:17+03:00 + Site Reliability Engineering - Part 2: Operational Balance in SRE + + https://foo.zone/gemfeed/2023-08-19-site-reliability-engineering-part-2.html + 2023-08-19T00:18:18+03:00 - Paul Buetow + Paul Buetow aka snonux paul@dev.buetow.org - Gogios is a minimalistic and easy-to-use monitoring tool I programmed in Google Go designed specifically for small-scale self-hosted servers and virtual machines. The primary purpose of Gogios is to monitor my personal server infrastructure for `foo.zone`, my MTAs, my authoritative DNS servers, my NextCloud, Wallabag and Anki sync server installations, etc. + This is the second part of my Site Reliability Engineering (SRE) series. I am currently employed as a Principal Site Reliability Engineer and will attempt to share what SRE is about in this blog series.
-

KISS server monitoring with Gogios


+

Site Reliability Engineering - Part 2: Operational Balance in SRE



-Published at 2023-06-01T21:10:17+03:00
+Published at 2023-08-19T00:18:18+03:00

-Gogios logo
+This is the second part of my Site Reliability Engineering (SRE) series. I am currently employed as a Principal Site Reliability Engineer and will attempt to share what SRE is about in this blog series.

-

Introduction


+2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture
+2023-08-19 Site Reliability Engineering - Part 2: Operational Balance in SRE (You are currently reading this)

-Gogios is a minimalistic and easy-to-use monitoring tool I programmed in Google Go designed specifically for small-scale self-hosted servers and virtual machines. The primary purpose of Gogios is to monitor my personal server infrastructure for foo.zone, my MTAs, my authoritative DNS servers, my NextCloud, Wallabag and Anki sync server installations, etc.
+
+⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣠⣾⣷⣄⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
+⠀⠀⠀⠀⣾⠿⠿⠿⠶⠾⠿⠿⣿⣿⣿⣿⣿⣿⠿⠿⠶⠶⠿⠿⠿⣷⠀⠀⠀⠀
+⠀⠀⠀⣸⢿⣆⠀⠀⠀⠀⠀⠀⠀⠙⢿⡿⠉⠀⠀⠀⠀⠀⠀⠀⣸⣿⡆⠀⠀⠀
+⠀⠀⢠⡟⠀⢻⣆⠀⠀⠀⠀⠀⠀⠀⣾⣧⠀⠀⠀⠀⠀⠀⠀⣰⡟⠀⢻⡄⠀⠀
+⠀⢀⣾⠃⠀⠀⢿⡄⠀⠀⠀⠀⠀⢠⣿⣿⡀⠀⠀⠀⠀⠀⢠⡿⠀⠀⠘⣷⡀⠀
+⠀⣼⣏⣀⣀⣀⣈⣿⡀⠀⠀⠀⠀⣸⣿⣿⡇⠀⠀⠀⠀⢀⣿⣃⣀⣀⣀⣸⣧⠀
+⠀⢻⣿⣿⣿⣿⣿⣿⠃⠀⠀⠀⠀⣿⣿⣿⣿⠀⠀⠀⠀⠈⢿⣿⣿⣿⣿⣿⡿⠀
+⠀⠀⠉⠛⠛⠛⠋⠁⠀⠀⠀⠀⢸⣿⣿⣿⣿⡆⠀⠀⠀⠀⠈⠙⠛⠛⠛⠉⠀⠀
+⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠸⣿⣿⣿⣿⠇⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
+⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣠⣾⣿⣿⣷⣄⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
+⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣸⣿⣿⣿⣿⣿⣿⣆⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
+⠀⠀⠀⠀⠀⠀⠴⠶⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠿⠶⠦⠀⠀
+

-With compatibility with the Nagios Check API, Gogios offers a simple yet effective solution to monitor a limited number of resources. In theory, Gogios scales to a couple of thousand checks, though. You can clone it from Codeberg here:
+

Operational Balance in SRE: Finding the Equilibrium in Reliability and Velocity



-https://codeberg.org/snonux/gogios
+Site Reliability Engineering has established itself as more than just a set of best practices or methodologies. Instead, it stands as a beacon of operational excellence, which guides engineering teams through the turbulent waters of modern software development and system management.

-
-    _____________________________    ____________________________
-   /                             \  /                            \
-  |    _______________________    ||    ______________________    |
-  |   /                       \   ||   /                      \   |
-  |   | # Alerts with status c|   ||   | # Unhandled alerts:  |   |
-  |   | hanged:               |   ||   |                      |   |
-  |   |                       |   ||   | CRITICAL: Check Pizza|   |
-  |   | OK->CRITICAL: Check Pi|   ||   | : Late delivery      |   |
-  |   | zza: Late delivery    |   ||   |                      |   |
-  |   |                       |   ||   | WARNING: Check Thirst|   |
-  |   |                       |   ||   | : OutofKombuchaExcept|   |
-  |   \_______________________/   ||   \______________________/   |
-  |  /|\ GOGIOS MONITOR 1    _    ||  /|\ GOGIOS MONITOR 2   _    |
-   \_____________________________/  \____________________________/
-     !_________________________!      !________________________!
-
-------------------------------------------------
-ASCII art was modified by Paul Buetow
-The original can be found at
-https://asciiart.website/index.php?art=objects/computers
-
+In the universe of software production, two fundamental forces are often at odds: The drive for rapid feature release (velocity) and the need for system reliability. Traditionally, the faster teams moved, the more risk was introduced into systems. SRE offers a approach to mitigate these conflicting drives through concepts like error budgets and SLIs/SLOs. These mechanisms provide a tangible metric, allowing teams to quantify how much they can push changes while ensuring they don't compromise system health. Thus, the error budget becomes a balancing act, where teams weigh the trade-offs between innovation and reliability.

-

Motivation


+A quintessential component of this balance is the dichotomy between operations and coding. According to SRE principles, an engineer should ideally spend an equal amount of time on operations work and coding - 50% on each. This isn't just a random metric; it's a reflection of the value SRE places on both maintaining operational excellence and progressing forward with innovations. This balance ensures that while SREs are solving today's problems, they are also preparing for tomorrow's challenges.

-With experience in monitoring solutions like Nagios, Icinga, Prometheus and OpsGenie, these tools often came with many features that I didn't necessarily need for personal use. Contact groups, host groups, check clustering, and the requirement of operating a DBMS and a WebUI added complexity and bloat to my monitoring setup.
+However, not all operational tasks are equal. SRE differentiates between "ops work" and "toil". While ops work is integral to system maintenance and can provide value, toil represents repetitive, mundane tasks which offer little value in the long run. Recognising and minimising toil is crucial. A culture that allows engineers to drown in toil stifles innovation and growth. Hence, an organisation's approach to toil indicates its operational health and commitment to balance.

-My primary goal was to have a single email address for notifications and a simple mechanism to periodically execute standard Nagios check scripts and notify me of any state changes. I wanted the most minimalistic monitoring solution possible but wasn't satisfied with the available options.
+A cornerstone of achieving operational balance lies in the tools and processes SREs use. Effective monitoring, observability tools, and ensuring that tools can handle high cardinality data are foundational. These aren't just technical requisites but reflective of an organisational culture prioritising proactive problem-solving. By having systems that effectively flag potential issues before they escalate, SREs can maintain the delicate balance between system stability and forward momentum.

-This led me to create Gogios, a lightweight monitoring tool tailored to my specific needs. I chose the Go programming language for this project as it comes, in my opinion, with the best balance of ease to use and performance.
+Moreover, operational balance isn't just a technological or process challenge; it's a human one. The health of on-call engineers is as crucial as the health of the services they manage. On-call postmortems, continuous feedback loops, and recognising gaps (be it tooling, operational expertise, or resources) ensure that the human elements of operations are noticed.

-

Features


+In conclusion, operational balance in SRE is not a static thing but an ongoing journey. It requires organisations to constantly evaluate their practices, tools, and, most importantly, their culture. By achieving this balance, organisations can ensure that they have time for innovation while maintaining the robustness and reliability of their systems, resulting in sustainable long-term success.

-
    -
  • Compatible with Nagios Check scripts: Gogios leverages the widely-used Nagios Check API, allowing to use existing Nagios plugins.
  • -
  • Lightweight and Minimalistic: Gogios is designed to be simple and fairly easy to set up.
  • -
  • Configurable Check Timeout and Concurrency: Gogios allows you to set a timeout for checks and configure the number of concurrent checks, offering flexibility in monitoring your resources.
  • -
  • Configurable check dependency: A check can depend on another check, which enables scenarios like not executing an HTTP check when the server isn't pingable.
  • -
  • Retries: Check retry and retry intervals are configurable per check.
  • -
  • Email Notifications: Gogios can send email notifications regarding the status of monitored services, ensuring you stay informed about potential issues.
  • -
  • CRON-based Execution: Gogios can be quickly scheduled to run periodically via CRON, allowing you to automate monitoring without needing a complex setup.
  • -

-

Example alert


+That all sounds very romantic. The truth is, it is brutal to archive the perfect balance. No system will ever be perfect. But at least we should aim for it!

-This is an example alert report received via E-Mail. Whereas, [C:2 W:0 U:0 OK:51] means that we've got two alerts in status critical, 0 warnings, 0 unknowns and 51 OKs.
+The third part of this blog series will be published soon :-)
+
+E-Mail your comments to paul at buetow.org :-)
+
+Back to the main site
+
+
+
+ + Site Reliability Engineering - Part 1: SRE and Organizational Culture + + https://foo.zone/gemfeed/2023-08-18-site-reliability-engineering-part-1.html + 2023-08-18T22:43:47+03:00 + + Paul Buetow aka snonux + paul@dev.buetow.org + + The universe of Site Reliability Engineering (SRE) is like an intricate tapestry woven with diverse technology, culture, and personal grit threads. Site Reliability Engineering is one of the most demanding jobs. With all the facets, it is impossible to get bored. There is always a new challenge to master, and there is always a new technology to tinker with. It's not just technical; it's also about communication, collaboration and teamwork. I am currently employed as a Principal Site Reliability Engineer and will attempt to share what SRE is about in this blog series. + +
+

Site Reliability Engineering - Part 1: SRE and Organizational Culture


+
+Published at 2023-08-18T22:43:47+03:00
+
+The universe of Site Reliability Engineering (SRE) is like an intricate tapestry woven with diverse technology, culture, and personal grit threads. Site Reliability Engineering is one of the most demanding jobs. With all the facets, it is impossible to get bored. There is always a new challenge to master, and there is always a new technology to tinker with. It's not just technical; it's also about communication, collaboration and teamwork. I am currently employed as a Principal Site Reliability Engineer and will attempt to share what SRE is about in this blog series.
+
+2023-08-18 Site Reliability Engineering - Part 1: SRE and Organizational Culture (You are currently reading this)
+2023-08-19 Site Reliability Engineering - Part 2: Operational Balance in SRE

-Subject: GOGIOS Report [C:2 W:0 U:0 OK:51]
-
-This is the recent Gogios report!
-
-# Alerts with status changed:
-
-OK->CRITICAL: Check ICMP4 vulcan.buetow.org: Check command timed out
-OK->CRITICAL: Check ICMP6 vulcan.buetow.org: Check command timed out
-
-# Unhandled alerts:
-
-CRITICAL: Check ICMP4 vulcan.buetow.org: Check command timed out
-CRITICAL: Check ICMP6 vulcan.buetow.org: Check command timed out
-
-Have a nice day!
+▓▓▓▓░░                                                                                  
+                                                                                          
+DC on fire:
+                                                                                          
+                ▓▓                                    ▓▓                ▓▓                
+      ░░  ░░    ▓▓▓▓                  ██                  ░░            ▓▓▓▓        ▓▓    
+    ▓▓░░░░  ░░  ▓▓▓▓                              ▓▓░░                  ▓▓▓▓              
+    ░░░░      ▓▓▓▓▓▓        ▓▓      ▓▓            ▓▓                  ▓▓▓▓▓▓      ▓▓      
+    ▓▓░░    ▓▓▒▒▒▒▓▓▓▓    ▓▓        ▓▓▓▓        ▓▓▓▓▓▓              ▓▓▒▒▒▒▓▓▓▓    ▓▓▓▓    
+  ██▓▓      ▓▓▒▒░░▒▒▓▓  ▓▓██      ▓▓▓▓▓▓        ▓▓▒▒▓▓              ▓▓▒▒░░▒▒▓▓  ██▓▓▓▓    
+  ▓▓▓▓██  ▓▓▒▒░░░░▒▒▓▓  ▓▓▓▓      ▓▓▒▒▒▒▓▓    ▓▓▒▒░░▒▒▓▓██▓▓      ▓▓▒▒░░░░▒▒▓▓  ▓▓▒▒▒▒▓▓  
+  ▓▓▒▒▒▒▓▓▓▓▒▒░░▒▒▓▓▓▓▓▓▒▒▒▒▓▓  ▓▓▓▓░░▒▒▓▓    ▓▓▒▒░░▒▒▓▓▒▒▒▒▓▓    ▓▓▒▒░░▒▒▓▓▓▓▓▓▓▓░░▒▒▓▓  
+  ▒▒░░▒▒▓▓▓▓▒▒░░▒▒▓▓▓▓▒▒░░▒▒▓▓  ▓▓▒▒░░▒▒▓▓    ▓▓░░░░▒▒▒▒░░░░▒▒██████▒▒░░▒▒██▓▓▓▓▒▒░░▒▒▓▓██
+  ░░░░▒▒▓▓▒▒░░▒▒▓▓▓▓▓▓▒▒░░▒▒▓▓██▒▒░░░░▒▒▓▓  ▓▓▒▒░░▒▒▓▓▒▒▒▒░░▒▒▓▓▓▓▒▒░░▒▒▓▓▓▓▓▓▒▒░░░░▒▒▓▓▓▓
+  ░░░░▒▒▓▓▒▒░░░░▓▓██▒▒░░░░▒▒▓▓██▒▒░░░░▒▒██▓▓▓▓▒▒░░▒▒▓▓▓▓▒▒░░░░▒▒▓▓▒▒░░░░██▓▓▓▓▒▒░░░░▒▒████
+  ▒▒░░▒▒▓▓▓▓░░░░▒▒▓▓▒▒▒▒░░░░▒▒▓▓▓▓▒▒░░░░▒▒▓▓▓▓▒▒░░░░▒▒▓▓▒▒░░▒▒▓▓▓▓▓▓░░░░▒▒▓▓▓▓▓▓▒▒░░░░▒▒▓▓
+  ▒▒░░▒▒▓▓▒▒▒▒░░▒▒██▒▒▒▒░░▒▒▒▒██▒▒▒▒░░░░░░▒▒▓▓▒▒░░░░▒▒▒▒░░░░▒▒████▒▒▒▒░░▒▒██▓▓▒▒▒▒░░░░░░▒▒
+  ░░░░░░▒▒░░░░░░░░▒▒▒▒▒▒░░░░▒▒▒▒▒▒░░░░░░░░▒▒▒▒░░░░░░▒▒▒▒░░░░░░▒▒▒▒░░░░░░░░▒▒▒▒▒▒░░░░░░░░▒▒
+  ░░░░░░░░░░▒▒░░░░░░░░░░░░░░░░░░░░░░░░▒▒░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░▒▒░░░░░░░░░░░░░░░░░░
 

-

Installation


+

SRE and Organizational Culture: Navigating the Nexus



-

Compiling and installing Gogios


+At the heart of SRE lies the proactive mindset of "prevention over cure". Traditional IT models focused predominantly on reactive solutions, but SRE mandates a shift towards foresight. By adopting Service Level Indicators (SLIs) and Service Level Objectives (SLOs), teams are equipped with clear metrics and goals that guide them toward ensuring reliability and user satisfaction. However, these aren't mere numbers. They reflect an organisational culture prioritising user experience and constant system alignment with user needs.

-This document is primarily written for OpenBSD, but applying the corresponding steps to any Unix-like (e.g. Linux-based) operating system should be easy. On systems other than OpenBSD, you may always have to replace does with the sudo command and replace the /usr/local/bin path with /usr/bin.
+Another defining SRE concept is the "error budget". This ingenious framework accepts that no system is flawless. Failures are inevitable. However, instead of being punitive, the culture here is to accept, learn, and iterate. By providing teams with a "budget" for errors, organisations foster an environment where innovation is encouraged, and failures are viewed as learning opportunities.

-To compile and install Gogios on OpenBSD, follow these steps:
+But SRE isn't just about technology and metrics; it's deeply human. It challenges the "hero culture" that plagues many IT teams. While individual heroics might occasionally save the day, a sustainable model requires collective expertise. An SRE culture recognises that heroes achieve their best within teams, negating the need for a hero-centric environment. This philosophy promotes a balanced on-call experience, emphasising the importance of trust, ownership, effective communication, and collaboration as cornerstones of team success. I personally have fallen into the hero trap, and I know it is unsustainable to be the only go-to person for every problem.

- -
git clone https://codeberg.org/snonux/gogios.git
-cd gogios
-go build -o gogios cmd/gogios/main.go
-doas cp gogios /usr/local/bin/gogios
-doas chmod 755 /usr/local/bin/gogios
-
+Additionally, the SRE model requires good documentation. However, it's essential to ensure that this documentation undergoes the same quality checks as code, reinforcing effective onboarding, training and communication.

-You can use cross-compilation if you want to compile Gogios for OpenBSD on a Linux system without installing the Go compiler on OpenBSD. Follow these steps:
+Organisations might face a significant challenge when adopting SRE. It is convincing various teams and leadership of its merits. Some might feel SRE principles counter their goals. They might prioritise feature rollouts over reliability or view SRE practices as cumbersome. Hence, fostering an SRE culture often demands patient explanations and showcasing tangible benefits, such as increased release velocity and improved user experience.

- -
export GOOS=openbsd
-export GOARCH=amd64
-go build -o gogios cmd/gogios/main.go
-
+Monitoring and observability form another SRE pillar, emphasising the need for high-quality tools to query and analyse data. This ties back to the cultural emphasis on continuous learning and adaptability. SREs, by nature, need to be curious, ready to delve into anomalies, and keen on adopting new tools and practices.

-On your OpenBSD system, copy the binary to /usr/local/bin/gogios and set the correct permissions as described in the previous section. All steps described here you could automate with your configuration management system of choice. I use Rexify, the friendly configuration management system, to automate the installation, but that is out of the scope of this document.
+Ultimately, the success of SRE within any organisation depends on the broader acceptance of its principles. It demands a move away from siloed operations, where SRE acts as a bandage on flawed systems, to a model where reliability is everyone's responsibility. It calls for cultural transformation from the on-call engineers to the boardroom.

-https://www.rexify.org
+In essence, the integration of SRE principles transcends technical practices. It paves the way for a shift in organisational culture that values proactive prevention, continuous learning, collaboration, and transparent communication. The successful melding of SRE and corporate culture promises not just reliable systems but also a robust, resilient, and progressive work environment.

-

Setting up user, group and directories


+Organisations with the implementation of SLIs, SLOs and error budgets are already advanced in their SRE journey. It takes a lot of communication, convincing, and patience until that point is reached.

-It is best to create a dedicated system user and group for Gogios to ensure proper isolation and security. Here are the steps to create the _gogios user and group under OpenBSD:
+Continue with the second part of this series:

- -
doas adduser -group _gogios -batch _gogios
-doas usermod -d /var/run/gogios _gogios
-doas mkdir -p /var/run/gogios
-doas chown _gogios:_gogios /var/run/gogios
-doas chmod 750 /var/run/gogios
-
+2023-08-19 Site Reliability Engineering - Part 2: Operational Balance in SRE

-Please note that creating a user and group might differ depending on your operating system. For other operating systems, consult their documentation for creating system users and groups.
+E-Mail your comments to paul at buetow.org :-)

-

Installing monitoring plugins


+Back to the main site
+
+
+
+ + Gemtexter 2.1.0 - Let's Gemtext again³ + + https://foo.zone/gemfeed/2023-07-21-gemtexter-2.1.0-lets-gemtext-again-3.html + 2023-07-21T10:19:31+03:00 + + Paul Buetow aka snonux + paul@dev.buetow.org + + I proudly announce that I've released Gemtexter version `2.1.0`. What is Gemtexter? It's my minimalist static site generator for Gemini Gemtext, HTML and Markdown, written in GNU Bash. + +
+

Gemtexter 2.1.0 - Let's Gemtext again³



-Gogios relies on external Nagios or Icinga monitoring plugin scripts. On OpenBSD, you can install the monitoring-plugins package with Gogios. The monitoring-plugins package is a collection of monitoring plugins, similar to Nagios plugins, that can be used to monitor various services and resources:
+Published at 2023-07-21T10:19:31+03:00

- -
doas pkg_add monitoring-plugins
-doas pkg_add nrpe # If you want to execute checks remotely via NRPE.
+
+-=[ typewriters ]=-  1/98
+                                        .-------.
+       .-------.                       _|~~ ~~  |_
+      _|~~ ~~  |_       .-------.    =(_|_______|_)
+    =(_|_______|_)=    _|~~ ~~  |_     |:::::::::|
+      |:::::::::|    =(_|_______|_)    |:::::::[]|
+      |:::::::[]|      |:::::::::|     |o=======.|
+      |o=======.|      |:::::::[]|     `"""""""""`
+ jgs  `"""""""""`      |o=======.|
+  mod. by Paul Buetow  `"""""""""`
 

-Once the installation is complete, you can find the monitoring plugins in the /usr/local/libexec/nagios directory, which then can be configured to be used in gogios.json.
+I proudly announce that I've released Gemtexter version 2.1.0. What is Gemtexter? It's my minimalist static site generator for Gemini Gemtext, HTML and Markdown, written in GNU Bash.

-

Configuration


+https://codeberg.org/snonux/gemtexter

-

MTA


+

Why Bash?



-Gogios requires a local Mail Transfer Agent (MTA) such as Postfix or OpenBSD SMTPD running on the same server where the CRON job (see about the CRON job further below) is executed. The local MTA handles email delivery, allowing Gogios to send email notifications to monitor status changes. Before using Gogios, ensure that you have a properly configured MTA installed and running on your server to facilitate the sending of emails. Once the MTA is set up and functioning correctly, Gogios can leverage it to send email notifications.
+This project is too complex for a Bash script. Writing it in Bash was to try out how maintainable a "larger" Bash script could be. It's still pretty maintainable and helps me try new Bash tricks here and then!

-You can use the mail command to send an email via the command line on OpenBSD. Here's an example of how to send a test email to ensure that your email server is working correctly:
+Let's list what's new!

-
-echo 'This is a test email from OpenBSD.' | mail -s 'Test Email' your-email@example.com
-
+

Switch to GPL3 license



-Check the recipient's inbox to confirm the delivery of the test email. If the email is delivered successfully, it indicates that your email server is configured correctly and functioning. Please check your MTA logs in case of issues.
+Many (almost all) of the tools and commands (GNU Bash, GMU Sed, GNU Date, GNU Grep, GNU Source Highlight) used by Gemtexter are licensed under the GPL anyway. So why not use the same? This was an easy switch, as I was the only code contributor so far!

-

Configuring Gogios


+

Source code highlighting support



-To configure Gogios, create a JSON configuration file (e.g., /etc/gogios.json). Here's an example configuration:
+The HTML output now supports source code highlighting, which is pretty neat if your site is about programming. The requirement is to have the source-highlight command, which is GNU Source Highlight, to be installed. Once done, you can annotate a bare block with the language to be highlighted. E.g.:
+
+
+ ```bash
+ if [ -n "$foo" ]; then
+   echo "$foo"
+ fi
+ ```
+
+
+The result will look like this (you can see the code highlighting only in the Web version, not in the Geminispace version of this site):

-
{
-  "EmailTo": "paul@dev.buetow.org",
-  "EmailFrom": "gogios@buetow.org",
-  "CheckTimeoutS": 10,
-  "CheckConcurrency": 2,
-  "StateDir": "/var/run/gogios",
-  "Checks": {
-    "Check ICMP4 www.foo.zone": {
-      "Plugin": "/usr/local/libexec/nagios/check_ping",
-      "Args": [ "-H", "www.foo.zone", "-4", "-w", "50,10%", "-c", "100,15%" ],
-      "Retries": 3,
-      "RetryInterval": 10
-    },
-    "Check ICMP6 www.foo.zone": {
-      "Plugin": "/usr/local/libexec/nagios/check_ping",
-      "Args": [ "-H", "www.foo.zone", "-6", "-w", "50,10%", "-c", "100,15%" ],
-      "Retries": 3,
-      "RetryInterval": 10
-    },
-    "www.foo.zone HTTP IPv4": {
-      "Plugin": "/usr/local/libexec/nagios/check_http",
-      "Args": ["www.foo.zone", "-4"],
-      "DependsOn": ["Check ICMP4 www.foo.zone"]
-    },
-    "www.foo.zone HTTP IPv6": {
-      "Plugin": "/usr/local/libexec/nagios/check_http",
-      "Args": ["www.foo.zone", "-6"],
-      "DependsOn": ["Check ICMP6 www.foo.zone"]
-    }
-    "Check NRPE Disk Usage foo.zone": {
-      "Plugin": "/usr/local/libexec/nagios/check_nrpe",
-      "Args": ["-H", "foo.zone", "-c", "check_disk", "-p", "5666", "-4"]
-    }
-  }
-}
+
if [ -n "$foo" ]; then
+  echo "$foo"
+fi
 

-
    -
  • EmailTo: Specifies the recipient of the email notifications.
  • -
  • EmailFrom: Indicates the sender's email address for email notifications.
  • -
  • CheckTimeoutS: Sets the timeout for checks in seconds.
  • -
  • CheckConcurrency: Determines the number of concurrent checks that can run simultaneously.
  • -
  • StateDir: Specifies the directory where Gogios stores its persistent state in a state.json file.
  • -
  • Checks: Defines a list of checks to be performed, each with a unique name, plugin path, and arguments.
  • -

-Adjust the configuration file according to your needs, specifying the checks you want Gogios to perform.
+Please run source-highlight --lang-list for a list of all supported languages.

-If you want to execute checks only when another check succeeded (status OK), use DependsOn. In the example above, the HTTP checks won't run when the hosts aren't pingable. They will show up as UNKNOWN in the report.
+

HTML exact variant



-Retries and RetryInterval are optional check configuration parameters. In case of failure, Gogios will retry Retries times each RetryInterval seconds.
+Gemtexter is there to convert your Gemini Capsule into other formats, such as HTML and Markdown. An HTML exact variant can now be enabled in the gemtexter.conf by adding the line declare -rx HTML_VARIANT=exact. The HTML/CSS output changed to reflect a more exact Gemtext appearance and to respect the same spacing as you would see in the Geminispace.

-For remote checks, use the check_nrpe plugin. You also need to have the NRPE server set up correctly on the target host (out of scope for this document).
+

Use of Hack webfont by default



-The state.json file mentioned above keeps track of the monitoring state and check results between Gogios runs, enabling Gogios only to send email notifications when there are changes in the check status.
+The Hack web font is a typeface designed explicitly for source code. It's a derivative of the Bitstream Vera and DejaVu Mono lineage, but it features many improvements and refinements that make it better suited to reading and writing code.

-

Running Gogios


+The font has distinctive glyphs for every character, which helps to reduce confusion between similar-looking characters. For example, the characters "0" (zero), "O" (capital o), and "o" (lowercase o), or "1" (one), "l" (lowercase L), and "I" (capital i) all have distinct looks in Hack, making it easier to read and understand code at a glance.

-Now it is time to give it a first run. On OpenBSD, do:
+Hack is open-source and freely available for use and modification under the MIT License.
+
+

HTML Mastodon verification support


+
+The following link explains how URL verification works in Mastodon:
+
+https://joinmastodon.org/verification
+
+So we have to hyperlink to the Mastodon profile to be verified and also to include a rel='me' into the tag. In order to do that add this to the gemtexter.conf (replace the URI to your Mastodon profile accordingly):

-
doas -u _gogios /usr/local/bin/gogios -cfg /etc/gogios.json
+
declare -xr MASTODON_URI='https://fosstodon.org/@snonux'
 

-To run Gogios via CRON on OpenBSD as the gogios user and check all services once per minute, follow these steps:
-
-Type doas crontab -e -u _gogios and press Enter to open the crontab file for the _gogios user for editing and add the following lines to the crontab file:
+and add the following into your index.gmi:

-*/5 8-22 * * * /usr/local/bin/gogios -cfg /etc/gogios.json
-0 7 * * * /usr/local/bin/gogios -renotify -cfg /etc/gogios.json
+=> https://fosstodon.org/@snonux Me at Mastodon
 

-Gogios is now configured to run every five minutes from 8 am to 10 pm via CRON as the _gogios user. It will execute the checks and send monitoring status whenever a check status changes via email according to your configuration. Also, Gogios will run once at 7 am every morning and re-notify all unhandled alerts as a reminder.
+The resulting line in the HTML output will be something as follows:

-

High-availability


+ +
<a href='https://fosstodon.org/@snonux' rel='me'>Me at Mastodon</a>
+

-To create a high-availability Gogios setup, you can install Gogios on two servers that will monitor each other using the NRPE (Nagios Remote Plugin Executor) plugin. By running Gogios in alternate CRON intervals on both servers, you can ensure that even if one server goes down, the other will continue monitoring your infrastructure and sending notifications.
+

More



-
    -
  • Install Gogios on both servers following the compilation and installation instructions provided earlier.
  • -
  • Install the NRPE server (out of scope for this document) and plugin on both servers. This plugin allows you to execute Nagios check scripts on remote hosts.
  • -
  • Configure Gogios on both servers to monitor each other using the NRPE plugin. Add a check to the Gogios configuration file (/etc/gogios.json) on both servers that uses the NRPE plugin to execute a check script on the other server. For example, if you have Server A and Server B, the configuration on Server A should include a check for Server B, and vice versa.
  • -
  • Set up alternate CRON intervals on both servers. Configure the CRON job on Server A to run Gogios at minutes 0, 10, 20, ..., and on Server B to run at minutes 5, 15, 25, ... This will ensure that if one server goes down, the other server will continue monitoring and sending notifications.
  • -
  • Gogios doesn't support clustering. So it means when both servers are up, unhandled alerts will be notified via E-Mail twice; from each server once. That's the trade-off for simplicity.
  • -

-There are plans to make it possible to execute certain checks only on certain nodes (e.g. on elected leader or master nodes). This is still in progress (check out my Gorum Git project).
+Additionally, there were a couple of bug fixes, refactorings and overall improvements in the documentation made.

-

Conclusion:


+Other related posts are:

-Gogios is a lightweight and straightforward monitoring tool that is perfect for small-scale environments. With its compatibility with the Nagios Check API, email notifications, and CRON-based scheduling, Gogios offers an easy-to-use solution for those looking to monitor a limited number of resources. I personally use it to execute around 500 checks on my personal server infrastructure. I am very happy with this solution.
+2021-04-24 Welcome to the Geminispace
+2021-06-05 Gemtexter - One Bash script to rule it all
+2022-08-27 Gemtexter 1.1.0 - Let's Gemtext again
+2023-03-25 Gemtexter 2.0.0 - Let's Gemtext again²
+2023-07-21 Gemtexter 2.1.0 - Let's Gemtext again³ (You are currently reading this)

E-Mail your comments to paul at buetow.org :-)

@@ -284,23 +270,22 @@ http://www.gnu.org/software/src-highlite --> - 'The Obstacle is the Way' book notes - - https://foo.zone/gemfeed/2023-05-06-the-obstacle-is-the-way-book-notes.html - 2023-05-06T17:23:16+03:00 + 'Software Developmers Career Guide and Soft Skills' book notes + + https://foo.zone/gemfeed/2023-07-17-career-guide-and-soft-skills-book-notes.html + 2023-07-17T04:56:20+03:00 - Paul Buetow + Paul Buetow aka snonux paul@dev.buetow.org - These are my personal takeaways after reading 'The Obstacle Is the Way' by Ryan Holiday. This is mainly for my own use, but you might find it helpful too. + These notes are of two books by 'John Sommez' I found helpful. I also added some of my own keypoints to it. These notes are mainly for my own use, but you might find them helpful, too.
-

"The Obstacle is the Way" book notes


+

"Software Developmers Career Guide and Soft Skills" book notes



-Published at 2023-05-06T17:23:16+03:00
-
-These are my personal takeaways after reading "The Obstacle Is the Way" by Ryan Holiday. This is mainly for my own use, but you might find it helpful too.
+Published at 2023-07-17T04:56:20+03:00

+These notes are of two books by "John Sommez" I found helpful. I also added some of my own keypoints to it. These notes are mainly for my own use, but you might find them helpful, too.

          ,..........   ..........,
@@ -314,532 +299,570 @@ http://www.gnu.org/software/src-highlite -->
                     '''
 

-"The obstacle is the way" is a powerful statement that encapsulates the wisdom of turning challenges into opportunities for growth and success. We will explore using obstacles as fuel, transforming weaknesses into strengths, and adopting a mindset that allows us to be creative and persistent in the face of adversity.
+

Improve



-

Reframe your perspective


-
-The obstacle in your path can become your path to success. Instead of being paralyzed by challenges, see them as opportunities to learn and grow. Remember, the things that hurt us often instruct us.
-
-We spend a lot of time trying to get things perfect and look at the rules, but what matters is that it works; it doesn't need to be after the book. Focus on results rather than on beautiful methods. In Jujitsu, it does matter that you bring your opponent down, but not how. There are many ways from point A to point B; it doesn't need to be a straight line. So many try to find the best solution but need to catch up on what is in Infront of them. Think progress and not perfection.
+

Always learn new things



-Don't always try to use the front door; a backdoor could open. It's nonsense. Don't fight the judo master with judo. Non-action can be action, exposing the weaknesses of others.
+When you learn something new, e.g. a programming language, first gather an overview, learn from multiple sources, play around and learn by doing and not consuming and form your own questions. Don't read too much upfront. A large amount of time is spent in learning technical skills which were never use. You want to have a practical set of skills you are actually using. You need to know 20 percent to get out 80 percent of the results.

+
    +
  • Learn a technology with a goal, e.g. implement a tool. Practice practise practice.
  • +
  • "I know X can do Y, I don't know exactly how, but I can look it up."
  • +
  • Read what experts are writing, for example follow blogs. Stay up to date and spent half an hour per day trading blogs and books.
  • +
  • Pick an open source application, read the code and try to understand it to get a feel of the syntax of the programming language.
  • +
  • Understand, that the standard library makes you a much better programmer.
  • +
  • Self learning is the top skill a programmer can have and is also useful in other aspects in your life.
  • +
  • Keep learning skills every day. Code every day. Don't be overconfident for job security. Read blogs, read books.
  • +
  • If you want to learn, then do it by exploring. Also teach what you learned (for example write a blog post or hold a presentation).
  • +

+Fake it until you make it. But be honest about your abilities or lack of. There is however only time between now and until you make it. Refer to your abilities to learn.

-

Embrace rationality


+Boot camps: The advantage of a boot camp is to pragmatically learn things fast. We almost always overestimate what we can do in a day. Especially during boot camps. Connect to others during the boot camps

-It is a superpower to see things rationally when others are fearful. Focus on the reality of the situation without letting emotions, such as anger, cloud your judgment. This ability will enable you to make better decisions in adversity. Ability to see things what they really are. E.g. wine is old fermented grapes, or other people behaving like animals during a fight. Show the middle finger if someone persists on the stupid rules occasionally.
+

Set goals



-

Control your response


+Your own goals are important but the manager also looks at how the team performs and how someone can help the team perform better. Check whether you are on track with your goals every 2 weeks in order to avoid surprises for the annual review. Make concrete goals for next review. Track and document your progress. Invest in your education. Make your goals known. If you want something, then ask for it. Nobody but you knows what you want.

-You can choose how you respond to obstacles. Focus on what you can control, and don't let yourself feel harmed by external circumstances. Remember, you decide how things affect you; nobody else does. Choose to feel good in response to any situation. Embrace the challenges and obstacles that come your way, as they are opportunities for growth and learning.
+

Ratings



-

Practice emotional and physical resilience


+That's a trap: If you have to rate yourself, that's a trap. That never works in an unbiased way. Rate yourself always the best way but rate your weakest part as high as possible minus one point. Rate yourself as good as you can otherwise. Nobody is putting for fun a gun on his own head.

-Martial artists know the importance of developing physical and emotional strength. Cultivate the art of not panicking; it will help you avoid making mistakes during high-pressure situations.
+
    +
  • Don't do peer rating, it can fire back on you. What if the colleague becomes your new boss?
  • +
  • Cooperate rankings are unfortunately HR guidelines and politics and only mirror a little your actual performance.
  • +

+

Promotions



-Focus on what you can control. Don't choose to feel harmed, and then you won't be harmed. I decide things that affect me; nobody else does. E.g., in prison, your mind stays your own. Don't ignore fear but explain it away, have a different view.
+The most valuable employees are the ones who make themselves obsolete and automate all away. Keep a safety net of 3 to 6 months of finances. Safe at least 10 percent of your earnings. Also, if you make money it does not mean that you have to spent more money. Is a new car better than a used car which both can bring you from A to B? Liability vs assets.

-

Persistence and patience


+
    +
  • Raise or promotion, what's better? Promotion is better as money will follow anyway then.
  • +
  • Take projects no-one wants and make them shine. A promotion will follow.
  • +
  • A promotion is not going to come to you because you deserve it. You have to hunt and ask for it.
  • +
  • Track all kudos (e.g. ask for emails from your colleagues).
  • +
  • Big corporations HRs don't expect a figjit. That's why it's so important to keep track of your accomplishments and kudos'.
  • +
  • If you want a raise be specific how much and know to back your demands. Don't make a thread and no ultimatums.
  • +
  • Best way for a promotion is to switch jobs. You can even switch back with a better salary.
  • +

+

Finish things



-Practice persistence and patience in your pursuits. Focus on the process rather than the prize and take one step at a time. Remember, the journey is about finishing tasks, projects, or workouts to the best of your ability. Never be in a hurry and never be desperate. There is no reason to be rushed; there are all in the long haul. Follow the process and not the price. Take it one step at a time. The process is about finishing (workout, task, project, etc.).
+Hard work is necessary for accomplish results. However, work smarter not harder. Furthermore, working smart is not a substitute for working hard. Work both, hard and smart.

-

Embrace failure


+
    +
  • Learn to finish things without motivation. Things will pay off when you stick to stuff and eventually motivation can also come back.
  • +
  • You will fail if you don't plan realistically. Set also a schedule and follow to it as of life depends on it.
  • +
  • Advances come only of you give more than asked. Consistency, commitment and knowing what you need to do is more key than hard work.
  • +
  • Any action is better than no action. If you get stuck you have gained nothing.
  • +
  • You need to know the unknowns. Identify as many unknown not known things as possible.
  • +

+Hard vs fun: Both engage the brain (video games vs work). Some work is hard and other is easy. Hard work is boring. The harsh truth is you have to put in hard and boring work in order to accomplish and be successful. Work won't be always boring though, as joy will follow with mastery.

-Failure is a natural part of life and can make us stronger. Treat defeat as a stepping stone to success and education. What is defeat? The first step to education. Failure makes you stronger. If we do our best, we can be proud of it, regardless of the result. Do your job, but do it right. Only an asshole thinks he is too good at the things he does. Also, asking for forgiveness is easier than asking for permission.
+Defeat is finally give up. Failure is the road to success, embrace it. Failure does not define you but how you respond to it. Events don't make your unhappy, but how you react to events do.

-

Be adaptable


+

Expand the empire



-There are many ways to achieve your goals; sometimes, unconventional methods are necessary. Feel free to break the rules or go off the beaten path if it will lead to better results. Transform weaknesses into strengths. We have a choice of how to respond to things. It's not about being positive but to be creative. Aim high, but stuff will happen; E.g., surprises will always happen.
+The larger your empire is, the larger your circle of influence is. The larger the circle of influence is, the more opportunities you have.

-

Embrace non-action


+
    +
  • Do the dirty work if you want to expand the empire. That's there the opportunities are.
  • +
  • SCRUM often fails due to the lack to commitment. The backlog just becomes a wish to get completed.
  • +
  • Apply work on your quality standards. Don't cross the line of compromise. Always improve your skills. Never be happy being good enough.
  • +

+Become visible, keep track that you accomplishments. E.g. write a weekly summary. Do presentations, be seen. Learn new things and share your learnings. Be the problem solver and not the blamer.

-We constantly push to the next thing. Sometimes the best course of action is standing still or even going backwards. Obstacles might resolve by themselves. Or going sideways. Sometimes, the best action is to stand still, go sideways, or even go backwards. Obstacles may resolve themselves or present new opportunities if you're patient and observant. People always want your input before you have all the facts. They want you to play after their rules. The question is, do you let them? The English call it the cool head. Being in control of Stress; requires practice. Appear, the absence of fear (Greek).