This tutorial shows how to replace always-on EC2 runners with autoscaling GitLab runners using the AWS fleeting plugin and an Auto Scaling Group, so instances start per job and shut down when the queue is empty.
More: https://ku.bz/nD6qbNZ3w
More: https://ku.bz/nD6qbNZ3w
Forwarded from LearnKube news
Pumba lets you kill, pause, and stress containers while injecting network delays, packet loss, and corruption.
You can deploy it as a DaemonSet for cluster-wide chaos engineering.
More: https://ku.bz/qcvwrrzn0
You can deploy it as a DaemonSet for cluster-wide chaos engineering.
More: https://ku.bz/qcvwrrzn0
Forwarded from KubeFM
This media is not supported in your browser
VIEW IN TELEGRAM
Delivery tooling matters most when it makes software delivery easier to understand.
Devin Allen points to Argo CD and Octopus Deploy as tools he watches because they make automated delivery more visible. The value is not just shipping faster. It is knowing what changed, why something failed, and how delivery behaves inside the organization.
Watch the full interview: https://ku.bz/8lKHj1C5d
Devin Allen points to Argo CD and Octopus Deploy as tools he watches because they make automated delivery more visible. The value is not just shipping faster. It is knowing what changed, why something failed, and how delivery behaves inside the organization.
Watch the full interview: https://ku.bz/8lKHj1C5d
Goldpinger is a monitoring tool that runs as a DaemonSet and makes inter-pod calls to test connectivity.
More: https://ku.bz/D_P9JG76K
More: https://ku.bz/D_P9JG76K
Forwarded from KubeFM
Media is too big
VIEW IN TELEGRAM
Vitalii Horbachov, Staff Software Engineer at Agoda, explains how Apple's transition to Silicon processors exposed critical flaws in their macOS virtualization approach. He details their original complex architecture that ran Linux on Mac Minis with kubelet, then used QEMU to virtualize macOS on top, creating multiple problematic layers.
Vitalii provides insight into the performance penalties and stability issues of their layered virtualization approach, and how a major hardware shift can expose fundamental architectural weaknesses in production infrastructure.
Watch the full episode: https://ku.bz/q_JS76SvM
Vitalii provides insight into the performance penalties and stability issues of their layered virtualization approach, and how a major hardware shift can expose fundamental architectural weaknesses in production infrastructure.
Watch the full episode: https://ku.bz/q_JS76SvM
Forwarded from KubeFM
Media is too big
VIEW IN TELEGRAM
Tanat Lokejaroenlarb, Staff Site Reliability Engineer @ Adevinta, explains the operational challenges his team faced managing over 2,500 Kubernetes nodes across 30 clusters using EKS Managed Node Groups and Cluster Autoscaler.
He details how the tight coupling between control plane and nodes made version upgrades brittle and noisy, while instance inflexibility created constant maintenance overhead.
Watch the full episode: https://ku.bz/T6hDSWYhb
He details how the tight coupling between control plane and nodes made version upgrades brittle and noisy, while instance inflexibility created constant maintenance overhead.
Watch the full episode: https://ku.bz/T6hDSWYhb
This case study covers migrating a live k3s cluster from a flat network to a VLAN architecture, including an etcd quorum loss caused by moving too many nodes at once and the recovery steps using k3s server
More: https://ku.bz/Yxmxk1dbc
--cluster-reset.More: https://ku.bz/Yxmxk1dbc
Forwarded from LearnKube news
Kubernetes is not difficult because there are too many commands.
It is difficult because networking, scheduling, deployments, storage, autoscaling, and security interact in ways that are hard to see.
Our live Advanced Kubernetes course connects those pieces into one practical mental model.
The next online course runs on 10, 11, 17, and 18 September.
- Four days of live instruction
- 60% hands-on labs
- Small classes
- Lifetime access to the material and private Slack
Joining individually?
https://learnkube.com/online-advanced-september-2026
Need several engineers to build the same baseline? We also deliver private training around your platform, workloads, and goals:
https://learnkube.com/corporate-training
It is difficult because networking, scheduling, deployments, storage, autoscaling, and security interact in ways that are hard to see.
Our live Advanced Kubernetes course connects those pieces into one practical mental model.
The next online course runs on 10, 11, 17, and 18 September.
- Four days of live instruction
- 60% hands-on labs
- Small classes
- Lifetime access to the material and private Slack
Joining individually?
https://learnkube.com/online-advanced-september-2026
Need several engineers to build the same baseline? We also deliver private training around your platform, workloads, and goals:
https://learnkube.com/corporate-training
Forwarded from LearnKube news
🚀 We just published The Technical Guide to Kubernetes Rightsizing in the Age of AI.
The guide follows the complete rightsizing process, from collecting metrics to applying changes safely in production.
- It explains how requests and limits affect scheduling and Linux resource controls.
- It examines how application runtimes change CPU and memory behavior.
- It also shows how KRR and VPA turn historical data into recommendations.
The book also examines where AI can help: collecting evidence, explaining recommendations, drafting policy, and carrying approved changes across systems without breaking prod.
Thank you to Gulcan and @danielepolencic for the research, experiments, writing, and illustrations behind this book.
Download the complete guide for free:
https://learnkube.com/kubernetes-rightsizing
The guide follows the complete rightsizing process, from collecting metrics to applying changes safely in production.
- It explains how requests and limits affect scheduling and Linux resource controls.
- It examines how application runtimes change CPU and memory behavior.
- It also shows how KRR and VPA turn historical data into recommendations.
The book also examines where AI can help: collecting evidence, explaining recommendations, drafting policy, and carrying approved changes across systems without breaking prod.
Thank you to Gulcan and @danielepolencic for the research, experiments, writing, and illustrations behind this book.
Download the complete guide for free:
https://learnkube.com/kubernetes-rightsizing
Forwarded from KubeFM
Media is too big
VIEW IN TELEGRAM
Niels Claeys, Lead data engineer & partner at Dataminded, breaks down the hidden resource overhead in Kubernetes clusters and explains how default scheduling strategies can significantly impact cost efficiency for batch processing workloads.
He explains that Kubernetes node overhead consists of two main components: node-level reservations for the operating system and eviction thresholds, plus daemon set overhead from cluster management tools. Using a concrete example, he shows how a 16GB RAM node only provides about 14GB of actual memory to running jobs.
Watch the full episode: https://ku.bz/hGRfkzDJW
He explains that Kubernetes node overhead consists of two main components: node-level reservations for the operating system and eviction thresholds, plus daemon set overhead from cluster management tools. Using a concrete example, he shows how a 16GB RAM node only provides about 14GB of actual memory to running jobs.
Watch the full episode: https://ku.bz/hGRfkzDJW
This media is not supported in your browser
VIEW IN TELEGRAM
GROOT is a Go CLI that collects Kubernetes logs and cluster context into one archive, with preflight checks, config profiles and archive summaries so you can attach a single file to a ticket.
More: https://ku.bz/rXWYbrY1S
More: https://ku.bz/rXWYbrY1S
Forwarded from KubeFM
Media is too big
VIEW IN TELEGRAM
Introducing Kube Signals: the new KubeFM show that turns keynote trends into direct conversations with the speakers shaping them.
For episode one, Brian Teller sits down with Saiyam Pathak from vCluster after his KubeCon India keynote on AI factories. They examine why the GPU beneath the model is becoming a platform-engineering problem.
They discuss:
- Why whole-GPU allocation wastes capacity
- How DRA, HAMI, MIG, and MPS enable sharing
- What Kubernetes must learn to support AI factories
Watch the full episode: https://ku.bz/4QZDqrnf-
This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.
For episode one, Brian Teller sits down with Saiyam Pathak from vCluster after his KubeCon India keynote on AI factories. They examine why the GPU beneath the model is becoming a platform-engineering problem.
They discuss:
- Why whole-GPU allocation wastes capacity
- How DRA, HAMI, MIG, and MPS enable sharing
- What Kubernetes must learn to support AI factories
Watch the full episode: https://ku.bz/4QZDqrnf-
This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.
This case study shows how one team ran LiteLLM as a single gateway to many model providers on EKS, kept it highly available, and managed the whole thing with ArgoCD.
More: https://ku.bz/YKJG9_NH_
More: https://ku.bz/YKJG9_NH_
Forwarded from LearnKube news
This case study shows how EXANTE replaced manual Saturday releases with a fully automated GitLab CI + Flux + Jira pipeline across 60+ Django modules, 7 GKE environments, and 30+ services to meet fintech regulatory audit requirements.
More: https://ku.bz/8BHV_JGB8
More: https://ku.bz/8BHV_JGB8
Forwarded from LearnKube news
This week on Learn Kubernetes Weekly 198:
🏗️ Data Lakehouse: Infrastructure
🔭 What the Popularity of Emerging Tools Tells Us About Kubernetes' Future
⚡ Kafka on Kubernetes: Performance Lessons for Any Disk-Heavy Data Service
🌐 To Centralise or Not to Centralise: The Questions That Shaped the Kubernetes CODECO Federated Architecture
🚨 Your AI Just Deleted the Wrong Deployment. Now What?
Read it now: https://kube.today/issues/198
⭐️ This newsletter is brought to you by LearnKube — master Kubernetes with hands-on training designed for engineers who want to learn the smart way https://ku.bz/hypSbyc-V
🏗️ Data Lakehouse: Infrastructure
🔭 What the Popularity of Emerging Tools Tells Us About Kubernetes' Future
⚡ Kafka on Kubernetes: Performance Lessons for Any Disk-Heavy Data Service
🌐 To Centralise or Not to Centralise: The Questions That Shaped the Kubernetes CODECO Federated Architecture
🚨 Your AI Just Deleted the Wrong Deployment. Now What?
Read it now: https://kube.today/issues/198
⭐️ This newsletter is brought to you by LearnKube — master Kubernetes with hands-on training designed for engineers who want to learn the smart way https://ku.bz/hypSbyc-V
l9gpu is a GPU telemetry agent that emits OpenTelemetry metrics with workload attribution built in, so you can see which pod, team or Slurm job is using each NVIDIA, AMD or Intel GPU.
More: https://ku.bz/JCtj79Xms
More: https://ku.bz/JCtj79Xms
Forwarded from Kube Architect
This case study shows how Netflix moved millions of batch jobs from its homegrown queueing system to Kueue, without the people submitting those jobs noticing any change.
More: https://ku.bz/3WtH2Fml9
More: https://ku.bz/3WtH2Fml9
Forwarded from KubeFM
This media is not supported in your browser
VIEW IN TELEGRAM
What emerging Kubernetes tools are experts paying attention to right now?
Bart Farrell from KubeFM looks back across 100+ KubeFM conversations to surface the tools guests kept mentioning, including Karpenter, Dapr, Argo CD, Kagent, Agent Gateway, OpenTelemetry, KRO, KCP, KubeVirt, Kueue, Kyverno, Headlamp, KEDA, Crossplane, KServe, ACK, and more.
Bart Farrell from KubeFM looks back across 100+ KubeFM conversations to surface the tools guests kept mentioning, including Karpenter, Dapr, Argo CD, Kagent, Agent Gateway, OpenTelemetry, KRO, KCP, KubeVirt, Kueue, Kyverno, Headlamp, KEDA, Crossplane, KServe, ACK, and more.
This guide shows you how to implement chaos engineering on Amazon EKS using AWS Fault Injection Service to simulate CPU stress and pod terminations while monitoring resilience via ADOT and Grafana.
More: https://ku.bz/cGJ0nyxxS
More: https://ku.bz/cGJ0nyxxS
MatrixHub is a self-hosted model registry you can run in place of Hugging Face, caching weights once and serving them quickly to many GPU nodes, including air-gapped networks.
More: https://ku.bz/6ty6V-QR3
More: https://ku.bz/6ty6V-QR3