KubeFM
331 subscribers
131 photos
1.26K videos
1.8K links
Podcast episodes, fireside chats, roundtables and educational programs about Kubernetes.
Download Telegram
Forwarded from LearnKube news
This week on Learn Kubernetes Weekly 197:

🎛️ Kubernetes Through Control Theory Glasses: HPA
🔭 Engineering End-to-End Observability for Kubernetes Workloads
🌋 Inside Volcano Controllers: Gang Scheduling, State Machines, and Real Kubernetes Logs
🍴 We forked Apache Stateful Functions for Flink 2.x — here's why
📉 Dagster on Kubernetes: When More Nodes Won't Save You

Read it now: https://kube.today/issues/197

⭐️ This newsletter is brought to you by LearnKube — master Kubernetes with hands-on training designed for engineers who want to learn the smart way https://ku.bz/hypSbyc-V
This media is not supported in your browser
VIEW IN TELEGRAM
Jorrick Stempher explains his team's rationale for selecting Facebook's Prophet model over alternatives like ARIMA, SARIMA, and NeuralProphet for their Kubernetes predictive scaling system.

He breaks down the key factors that influenced their decision: the need for a balance between accuracy and training speed, and Prophet's built-in seasonality detection capabilities.

Watch the full episode: https://ku.bz/clbDWqPYp
Media is too big
VIEW IN TELEGRAM
Migrating to the Gateway API on Amazon EKS involves a fundamental shift: annotations no longer live on the load balancer resource.

Sai Vennam explains the new YAML-native CRDs, and why auditing your existing annotations is the first step of any migration.



Watch the full interview: https://ku.bz/VVfQ4G891
Media is too big
VIEW IN TELEGRAM
"Deploying an AI agent or a model for inference is just another app."

Tsahi Duek argues companies don't adopt Kubernetes for AI from scratch — they already have the CI/CD pipelines, autoscaling, and observability in place. The same infrastructure that runs web services can shift to training jobs, then flip back to inference.



Watch the full interview: https://ku.bz/2r41YKBZb
This media is not supported in your browser
VIEW IN TELEGRAM
Delivery tooling matters most when it makes software delivery easier to understand.

Devin Allen points to Argo CD and Octopus Deploy as tools he watches because they make automated delivery more visible. The value is not just shipping faster. It is knowing what changed, why something failed, and how delivery behaves inside the organization.



Watch the full interview: https://ku.bz/8lKHj1C5d
Media is too big
VIEW IN TELEGRAM
Vitalii Horbachov, Staff Software Engineer at Agoda, explains how Apple's transition to Silicon processors exposed critical flaws in their macOS virtualization approach. He details their original complex architecture that ran Linux on Mac Minis with kubelet, then used QEMU to virtualize macOS on top, creating multiple problematic layers.

Vitalii provides insight into the performance penalties and stability issues of their layered virtualization approach, and how a major hardware shift can expose fundamental architectural weaknesses in production infrastructure.

Watch the full episode: https://ku.bz/q_JS76SvM
KubeFM
Vitalii Horbachov, Staff Software Engineer at Agoda, explains how Apple's transition to Silicon processors exposed critical flaws in their macOS virtualization approach. He details their original complex architecture that ran Linux on Mac Minis with kubelet…
This episode is brought to you by Testkube—where teams run millions of performance tests in real Kubernetes infrastructure. From air-gapped environments to massive scale deployments, orchestrate every testing tool in one platform. Check it out at https://ku.bz/lnxYK3s0L
Media is too big
VIEW IN TELEGRAM
At Komodor, they spend 10% of development time building AI features and 80% on validation.

Itiel Shwartz explains the real cost of shipping AI for Kubernetes operations: LLM-as-a-judge, A-B testing, benchmarking — a full suite of evaluation before anything reaches production. The challenge isn't building AI capabilities. It's proving they actually work.

You need internal confidence first. Then you earn your users' confidence.





Watch the full interview: https://ku.bz/b9bDXQ_xq

This interview is a reaction to Mai Nishitani's episode https://ku.bz/3hWvQjXxp
Media is too big
VIEW IN TELEGRAM
Tanat Lokejaroenlarb, Staff Site Reliability Engineer @ Adevinta, explains the operational challenges his team faced managing over 2,500 Kubernetes nodes across 30 clusters using EKS Managed Node Groups and Cluster Autoscaler.

He details how the tight coupling between control plane and nodes made version upgrades brittle and noisy, while instance inflexibility created constant maintenance overhead.

Watch the full episode: https://ku.bz/T6hDSWYhb
Forwarded from LearnKube news
Kubernetes is not difficult because there are too many commands.

It is difficult because networking, scheduling, deployments, storage, autoscaling, and security interact in ways that are hard to see.

Our live Advanced Kubernetes course connects those pieces into one practical mental model.

The next online course runs on 10, 11, 17, and 18 September.

- Four days of live instruction
- 60% hands-on labs
- Small classes
- Lifetime access to the material and private Slack

Joining individually?
https://learnkube.com/online-advanced-september-2026

Need several engineers to build the same baseline? We also deliver private training around your platform, workloads, and goals:
https://learnkube.com/corporate-training
Media is too big
VIEW IN TELEGRAM
App teams say it's networking. Networking says it's the app. Platform teams are stuck in the middle.

Elamaran Shanmugam walks through how organizations troubleshoot today — application metrics, infrastructure metrics, then networking — but by the time you dig into it, the ephemeral pod is gone. With container network observability on Amazon EKS, you get retransmission and flow-level metrics that pinpoint networking issues and can solve in minutes what used to take hours.



Watch the full interview: https://ku.bz/DCxhgWQqS
Media is too big
VIEW IN TELEGRAM
AWS built SOCI and the fast-pull snapshotter to improve container image pull times — but they didn't keep it to themselves. Phil Estes explains why open source and upstream-first development were the only approaches that made sense.

The tools integrate directly with Containerd, work with any cluster (not just EKS), and improvements are flowing back into upstream core projects. You don't even need a custom snapshotter to benefit.



Watch the full interview: https://ku.bz/_ZLldHwVC