DevOps & SRE notes
13.3K subscribers
51 photos
19 files
2.61K links
Helpful articles and tools for DevOps&SRE

WhatsApp: https://whatsapp.com/channel/0029Vb79nmmHVvTUnc4tfp2F

For paid consultation (RU/EN), contact: @tutunak


All ways to support https://telegra.ph/How-support-the-channel-02-19
Download Telegram
This article provides an insightful, framework-driven overview of automated post-mortem generation, defining how AI transforms incident retrospectives from manual reconstruction into automated drafting based on existing artifacts. It introduces a structural model for evaluating tools rather than just summarizing vendor features.

https://www.arvoai.ca/blog/automated-post-mortem-generation
๐Ÿ‘3
Forwarded from AI Vibe Notes
Hands-on comparison against a deliberately broken cluster, with real outputs and failure-mode differences. Practical for deciding where AI Kubernetes tools fit: scanner, agent framework, or natural-language kubectl layer.

https://decodeops.substack.com/p/k8sgpt-vs-kagent-vs-kubectl-ai-what
โค4๐Ÿ‘3๐Ÿ‘2๐Ÿ‘Ž1
Please open Telegram to view this post
VIEW IN TELEGRAM
Containerlab focuses on the containerized Network Operating Systems which are typically used to test network features and designs

https://github.com/srl-labs/containerlab
๐Ÿ‘2๐Ÿ”ฅ2
Lessons from Moving a Live Production Database describes the complex process of migrating a massive dataset while keeping the service available to users.

The case study covers specific strategies for achieving zero downtime during a high-risk infrastructure change.

But is zero downtime always worth the engineering effort?

Imagine that you have two options:
1. Spend several weeks preparing and testing a zero-downtime migration.
2. Schedule 20 minutes of downtime during a low-traffic period.
Considering that 99.9% availability allows approximately 43 minutes of downtime per month, which option would you choose?
What factors would change your decision: revenue loss, SLA penalties, customer expectations, rollback complexity, or the size of the engineering team?
https://www.tines.com/blog/zero-downtime-database-migrations-lessons-from-moving-a-live-production/
๐Ÿ‘3โค1
Please open Telegram to view this post
VIEW IN TELEGRAM
๐Ÿ‘4
Forwarded from AI Vibe Notes
Hands-on vendor-neutral instrumentation for GenAI spans, tool calls, token metrics, and trace exploration. Worth reading if you want AI-agent debugging to fit existing OTel/Grafana/Loki-style observability rather than a separate black box.

https://opentelemetry.io/blog/2026/genai-observability/
๐Ÿ‘3โค1๐Ÿ‘Ž1๐Ÿ‘1
Please open Telegram to view this post
VIEW IN TELEGRAM
โค5
Kubernetes configuration tracking controller.

Wave watches Deployments, StatefulSets and DaemonSets within a Kubernetes cluster and ensures that their Pods always have up to date configuration.

By monitoring mounted ConfigMaps and Secrets, Wave can trigger a Rolling Update of the Deployment when the mounted configuration is changed.

https://github.com/wave-k8s/wave
๐Ÿ‘3
Deep dive into Prometheusโ€™s use-uncached-io work and why page cache behavior can make Kubernetes container memory metrics misleading. Useful for anyone running Prometheus at scale: covers memory predictability, compaction writes, OOM risk, and tradeoffs around direct I/O.

https://prometheus.io/blog/2026/03/05/uncached-io/
๐Ÿ‘3๐Ÿ”ฅ1