KubeFM
332 subscribers
133 photos
1.27K videos
1.82K links
Podcast episodes, fireside chats, roundtables and educational programs about Kubernetes.
Download Telegram
Media is too big
VIEW IN TELEGRAM
Grzegorz Głąb, Kubernetes Engineer at Cloud Kitchens, explains how accumulated pods can negatively impact cluster health.

He describes implementing a simple but effective solution for managing succeeded and evicted pods that were causing ETCD growth and increased API server load. By creating a periodic screening mechanism that removes pods in succeeded phase or those evicted by node pressure after 15 minutes, his team improved cluster stability and reduced user confusion when viewing workloads.

Watch the full episode: https://kube.fmhttps://ku.bz/yg_fkP0LN
Media is too big
VIEW IN TELEGRAM
Shyam Jeedigunta, Principal Engineer at Amazon Web Services (AWS), explains the security challenges and solutions for onboarding Kubernetes nodes from different infrastructure providers.

He discusses how to handle identity management, certificate issuance, and trust establishment when nodes come from edge locations, on-premises infrastructure, or other cloud providers rather than the same infrastructure as the control plane.

Watch the full interview: https://ku.bz/m89tLbgcq
Media is too big
VIEW IN TELEGRAM
Alessandro Pomponio, Research Software Engineer @ IBM Research, explains the operational challenges of managing large bare-metal clusters for scientific teams.

Alessandro details specific problems his team encountered: GPU resource monopolization through interactive pods, large batch jobs overwhelming GPU nodes due to Kubernetes scheduler behavior, and users creating pods with commands like sleep infinity to use them as unofficial VMs.

Watch the full episode: https://ku.bz/5sK7BFZ-8
Media is too big
VIEW IN TELEGRAM
"I don't want to see AI agents autonomously control clusters right now."

Nick Eberts draws a clear line: AI assistants are valuable for read-only troubleshooting — giving hints, explaining what's wrong. But making changes? That should go through pull requests and human review, especially in a GitOps workflow. He also flags an emerging challenge: securing agent-to-agent communication between MCP servers and clients, and extending Istio authorization policies into the agent layer.

The takeaway: AI should assist, not act — until the guardrails catch up.



Watch the full interview: https://ku.bz/G1QSYQTn2

This interview is a reaction to Mai Nishitani's episode https://ku.bz/3hWvQjXxp
Media is too big
VIEW IN TELEGRAM
Andrew Hillier, Co-founder CTO @ Densify, discusses the ongoing debate around setting CPU limits in Kubernetes and shares practical insights from customer deployments.

He explains why many organizations are moving away from CPU limits in environments with abundant CPU resources, arguing that premature throttling often provides no benefit when nodes rarely reach capacity.

Watch the full interview: https://ku.bz/-0wmZX03V
Forwarded from LearnKube news
📕 We published a book on optimising and right-sizing GPUs in Kubernetes.

Most GPU clusters show 100% allocation and single-digit actual usage.

The book helps you:

- Tell whether your GPUs are actually computing or just allocated
- Pick the right metrics instead of trusting nvidia-smi
- Choose between time-slicing, MIG, and dedicated GPUs based on real data
- Stop GPU waste from cascading into CPU and memory waste

Download it for free here: ku.bz/KL4jRvsL4

This book was made possible by Kubex.
This media is not supported in your browser
VIEW IN TELEGRAM
Paul Butler, founder at Jamsocket, discusses the "happy path" for using Kubernetes resources effectively.

He recommends focusing on fundamental resources: Deployments, Services, ConfigMaps, Secrets, and CronJobs.

These resources provide a practical foundation for running workloads in Kubernetes without getting overwhelmed by the platform's complexity.

The advice stems from his experience building Jamsocket and represents what he wished he knew when starting with Kubernetes.

Watch the full episode: https://ku.bz/Dmn93dd7M
Media is too big
VIEW IN TELEGRAM
"The supply chain has become the sharp end of the wedge."

Andrew Martin traces the evolution of software supply chain attacks from boot sector viruses to modern npm-borne worms. His team signs everything, generates SBOMs, and verifies Cosign artifacts at admission time into Kubernetes clusters.

The prediction for 2026: continuous validation of supply chain security metadata at runtime will become a staple in Kubernetes security tooling this year.



Watch the full interview: https://ku.bz/wyMlWGTqf
This media is not supported in your browser
VIEW IN TELEGRAM
Zbyněk Roubalík, Founder & CTO @ Kedify, explains what sets Kedify apart from other autoscaling solutions in the market.

Rather than treating autoscaling as a single-dimensional problem, Kedify's approach recognizes that effective cost optimization and performance improvement on Kubernetes requires coordinated scaling across multiple layers of the infrastructure stack.

Watch the interview: https://ku.bz/qN7BLcYTK

Read the announcement: https://ku.bz/0XVsNHSnK
Media is too big
VIEW IN TELEGRAM
Karpenter can rotate your nodes for three reasons: they're underutilized, they're empty, or the AMI has drifted from what you specified.

You can set a disruption budget for each reason to control how many nodes rotate at once. But here's the catch: if you only set budgets for two reasons and skip the third, Karpenter doesn't disable it. It silently applies a default 10% budget to any reason you didn't mention.

Adhi Sutandi's team found this the hard way — drift events fired during maintenance windows they thought were locked down. The fix? Set a single budget of one node with no reason qualifier, so it applies to everything.



New episode out now: https://ku.bz/XyVfsSQPr
Media is too big
VIEW IN TELEGRAM
@miamorecadenza CEO at Techaro shares their approach to storage architecture in Kubernetes with a practical three-tier solution.

The setup includes Longhorn for critical data requiring replication, NFS for bulk storage of disposable data, and CSI S3 integration with Tigris for infrequently accessed content.

Each tier serves specific use cases, from running SQLite databases on replicated storage to managing media files through S3, demonstrating how to match storage solutions to workload requirements.

Watch the full episode: https://ku.bz/2kzj2MgfH
Media is too big
VIEW IN TELEGRAM
Standardizing a platform across teams sounds simple — until you realize every team needs something slightly different.

Abby Bangser at Syntasso explains that the key is composability: build low-level blocks that give teams freedom, then layer higher-level abstractions for quick starts. She borrows the concept of progressive discovery from Spotify — start users with simple, opinionated defaults, but let them "break glass" to customized components when needed.

The best platforms aren't standardized by force. They're composable by design.





Watch the full interview: https://ku.bz/5Rqq275dl

This interview is a reaction to Ángel Barrera Sánchez's episode https://ku.bz/-5QbzQXJg
Forwarded from LearnKube news
This week on Learn Kubernetes Weekly 173:

🔥 Kubernetes Egress Control with Squid Proxy
💪 How We Turned a Forced OS Migration into a 30% Infrastructure Reduction
Auto-scaling and Load-based Scaling in Kubernetes
🎯 Smart Scheduler: Intelligent Pod Placement for Kubernetes Cost Optimization
🤖 Using Claude Code to Pilot Kubernetes on Autodock

Read it now: https://kube.today/issues/173

⭐️ This newsletter is brought to you by Hadron, the new lightweight secure Linux OS from the Kairos team https://ku.bz/mMZytrj-z
Media is too big
VIEW IN TELEGRAM
Tanat Lokejaroenlarb, Staff Site Reliability Engineer at Adevinta, explains the hidden complexities of blue-green migrations between Kubernetes clusters.

He introduces the concept of "scale mismatch" where new clusters start empty while existing ones are already optimized for production workloads.

Tanat provides concrete examples of components requiring warm-up time:

- Ingress controllers needing to scale to handle production traffic
- Prometheus instances growing from a few gigabytes to 80GB of memory based on cluster size
- AWS elastic load balancers requiring manual pre-warming requests

He also highlights how stateful applications with persistent volumes bound to specific clusters create additional migration challenges, often requiring downtime, backup/restore procedures, and careful traffic management.

Watch the full episode: https://kube.fmhttps://ku.bz/VVHFfXGl_