DevOps&SRE Library
19.8K subscribers
431 photos
2 videos
2 files
5.45K links
Библиотека статей по теме DevOps и SRE.

Реклама: @ostinostin
Контент: @mxssl

РКН: https://www.gosuslugi.ru/snet/67704b536aa9672b963777b3
Download Telegram
diffyml

A fast, structural YAML diff tool with built-in Kubernetes intelligence. One dependency, minimal attack surface, native CI annotations for GitHub, GitLab, and Gitea.


https://github.com/szhekpisov/diffyml
crossview

A modern React-based dashboard for managing and monitoring Crossplane resources in Kubernetes. Visualize, search, and manage your infrastructure-as-code with ease.


https://github.com/crossplane-contrib/crossview
ingress-nginx-migration

The Ingress NGINX Migration is a tool that analyzes Kubernetes NGINX Ingress resources to help with migration planning to Traefik.


https://github.com/traefik/ingress-nginx-migration
kubevirt-benchmark

A comprehensive, vendor-neutral performance testing toolkit for KubeVirt virtual machines running on OpenShift Container Platform (OCP) or any Kubernetes distribution with KubeVirt.


https://github.com/portworx/kubevirt-benchmark
openchoreo

OpenChoreo is a developer platform for Kubernetes offering development and architecture abstractions, a Backstage-powered developer portal, application CI/CD, GitOps, and observability.


https://github.com/openchoreo/openchoreo
Observability at the Edge: OpenTelemetry in Ingress Controllers

https://www.dash0.com/blog/observability-at-the-edge-opentelemetry-in-ingress-controllers
browserly

A smart macOS menu bar app that routes URLs to the right browser based on custom rules.


https://github.com/andyzasl/browserly
5 InfluxDB Alternatives in 2026: An Honest Comparison

https://basekick.net/blog/influxdb-alternatives-2026
Securing CI/CD for an open source project: lessons from Cilium

https://cilium.io/blog/2026/05/06/securing-cicd-open-source-lessons-from-cilium
semble

Semble is a code search library built for agents. It returns the exact code snippets they need instantly, using ~98% fewer tokens than grep+read.


https://github.com/MinishLab/semble
extenddb

A DynamoDB-compatible API adapter, ExtendDB speaks the DynamoDB wire protocol — any AWS SDK, CLI, or tool that works with DynamoDB works with ExtendDB, unchanged.


https://github.com/ExtendDB/extenddb
Monitoring reliably at scale

Designing monitoring that works when everything else doesn’t.


https://medium.com/airbnb-engineering/monitoring-reliably-at-scale-ca6483040930
When AI SRE Fails: Production Reality, Failure Modes, and What They Cost

What you won't find in the marketing collateral is the documented production case where a four-agent AI SRE system runs to €8,500 per month — a 15x multiplier over a simple LLM chat implementation — a number most teams discover only after they've deployed.


https://www.softwareseni.com/when-ai-sre-fails-production-reality-failure-modes-and-what-they-cost
The Pulse: AI load breaks GitHub – why not other vendors?

GitHub's reliability has been beyond unacceptable recently: last month, third party measurements pinned it at one nine (right at 90%).


https://blog.pragmaticengineer.com/the-pulse-ai-load-breaks-github
You've Got (Too Much) Mail: Behind the Scenes of the 3/25/26 Voice Outage

As part of a routine infrastructure change, a configuration update accidentally caused a large portion of Discord's session management servers to shut down simultaneously.


https://discord.com/blog/behind-the-scenes-of-the-3-25-26-voice-outage
Incident Report: May 19, 2026 - GCP Account Suspension

Railway experienced a platform-wide service disruption due to Google Cloud incorrectly placing our account in a suspended status.


https://blog.railway.com/p/incident-report-may-19-2026-gcp-account-outage
Why Your KServe InferenceService Won't Become Ready: Four Production Failures and Fixes

A practitioner's account of the errors the KServe getting-started documentation doesn't tell you about — with exact terminal output, root causes, and working Kustomize patches.


https://sodiq-jimoh.hashnode.dev/why-your-kserve-inferenceservice-won-t-become-ready-four-production-failures-and-fixes
A one-line Kubernetes fix that saved 600 hours a year

Every time we restarted Atlantis, the tool we use to plan and apply Terraform changes, we’d be stuck for 30 minutes waiting for it to come back up.


https://blog.cloudflare.com/one-line-kubernetes-fix-saved-600-hours-a-year
Why Kubernetes Has No Login — And How We Solved It for AuditRadar

When we set out to build the Logins page for AuditRadar — a real-time audit log explorer for OpenShift and Kubernetes — we hit a wall that forced us to deeply understand how authentication actually works on each platform.


https://blog.audit-radar.com/why-kubernetes-has-no-login-and-how-we-solved-it-for-auditradar