Media is too big
VIEW IN TELEGRAM
Paul Butler, founder at Jamsocket, explains how to distinguish between necessary and accidental complexity in software systems.
He uses the microwave timer analogy to demonstrate how tools can become over-engineered: while a microwave needs a complex interface for cooking, using it as a timer could be as simple as an egg timer's dial.
This discussion connects to the broader Kubernetes ecosystem where the community is actively working on reducing complexity while preserving essential features.
Watch the full episode: https://ku.bz/Dmn93dd7M
He uses the microwave timer analogy to demonstrate how tools can become over-engineered: while a microwave needs a complex interface for cooking, using it as a timer could be as simple as an egg timer's dial.
This discussion connects to the broader Kubernetes ecosystem where the community is actively working on reducing complexity while preserving essential features.
Watch the full episode: https://ku.bz/Dmn93dd7M
Media is too big
VIEW IN TELEGRAM
Andrew Hillier, Co-Founder & CTO at Densify, reflects on what he would tell himself 10 years ago when starting with Kubernetes. He emphasizes that technology adoption always takes longer than expected and that it took the full decade for organizations to truly understand that misconfigured Kubernetes environments can be costly to run.
Andrew challenges the common misconception that new technologies magically solve existing problems. He explains that Kubernetes doesn't eliminate fundamental spending and performance challenges—it simply presents them in different forms. His key advice centers on patience and maintaining focus on core issues rather than expecting quick transformations.
Watch the full interview: https://ku.bz/YMHdrgqz4
Andrew challenges the common misconception that new technologies magically solve existing problems. He explains that Kubernetes doesn't eliminate fundamental spending and performance challenges—it simply presents them in different forms. His key advice centers on patience and maintaining focus on core issues rather than expecting quick transformations.
Watch the full interview: https://ku.bz/YMHdrgqz4
KubeFM
Andrew Hillier, Co-Founder & CTO at Densify, reflects on what he would tell himself 10 years ago when starting with Kubernetes. He emphasizes that technology adoption always takes longer than expected and that it took the full decade for organizations to truly…
This interview is brought to you with support from Kubex by Densify - automated Kubernetes optimization that cuts costs and improves performance https://ku.bz/8H62chDf9
Media is too big
VIEW IN TELEGRAM
Danielle Cook, Senior Product Marketing Manager @ Akamai, discusses the critical challenge organizations face when moving from AI model training to inference deployment.
She explains how traditional centralized data centers create latency issues when users need to consume AI-generated content, and how Akamai Inference Cloud enables inference at the edge to solve this problem.
Watch the interview: https://ku.bz/twmNqt6wX
Read the announcement: https://ku.bz/MPylRMg8K
She explains how traditional centralized data centers create latency issues when users need to consume AI-generated content, and how Akamai Inference Cloud enables inference at the edge to solve this problem.
Watch the interview: https://ku.bz/twmNqt6wX
Read the announcement: https://ku.bz/MPylRMg8K
Media is too big
VIEW IN TELEGRAM
Zbyněk Roubalík, Founder & CTO @ Kedify, explains the current limitations of Kubernetes cluster autoscaling and his vision for more intelligent infrastructure provisioning.
He discusses how cluster autoscalers like Karpenter work well with pod-level pressure from HPA, scaling based on CPU and memory requests, but argues for a more proactive approach. Zbyněk highlights the technical challenge of connecting different types of metrics to enable proactive cluster scaling - a capability that's currently very difficult to achieve in Kubernetes.
Watch the full interview: https://ku.bz/vc-lBjCr0
This interview is a reaction to Jorrick Stempher's episode https://ku.bz/clbDWqPYp
He discusses how cluster autoscalers like Karpenter work well with pod-level pressure from HPA, scaling based on CPU and memory requests, but argues for a more proactive approach. Zbyněk highlights the technical challenge of connecting different types of metrics to enable proactive cluster scaling - a capability that's currently very difficult to achieve in Kubernetes.
Watch the full interview: https://ku.bz/vc-lBjCr0
This interview is a reaction to Jorrick Stempher's episode https://ku.bz/clbDWqPYp
KubeFM
Zbyněk Roubalík, Founder & CTO @ Kedify, explains the current limitations of Kubernetes cluster autoscaling and his vision for more intelligent infrastructure provisioning. He discusses how cluster autoscalers like Karpenter work well with pod-level pressure…
This interview is brought to you by Kedify, the enterprise autoscaling platform from KEDA's co-creator. Download their free 2025 Kubernetes Autoscaling Playbook to learn how to scale on HTTP intent, OpenTelemetry, and GPU metrics https://ku.bz/z-sMsk3GM
This media is not supported in your browser
VIEW IN TELEGRAM
@miamorecadenza CEO at Techaro explains why traditional password-based authentication is problematic in Kubernetes clusters and how Talos Linux implements a more secure approach using CA certificates.
Watch the full episode: https://ku.bz/2kzj2MgfH
Watch the full episode: https://ku.bz/2kzj2MgfH
Media is too big
VIEW IN TELEGRAM
Alex Arnell, Principal Member Of Technical Staff at Heroku, explains how OpenTelemetry's "T-shape" methodology distinguishes monitoring from observability in practice.
He breaks down how 80% of observability should cover basic golden signals (standard RED metrics that trigger alerts), while the remaining 20% focuses on business-specific, custom telemetry that enables deeper investigation.
Alex walks through a real incident response workflow: starting with high-level monitoring alerts, then using distributed tracing to pinpoint exact failure points across infrastructure components.
Watch the full interview: https://ku.bz/Lsr8gltrH
This interview is a reaction to Miguel Luna's episode https://ku.bz/WwS04jYvv
He breaks down how 80% of observability should cover basic golden signals (standard RED metrics that trigger alerts), while the remaining 20% focuses on business-specific, custom telemetry that enables deeper investigation.
Alex walks through a real incident response workflow: starting with high-level monitoring alerts, then using distributed tracing to pinpoint exact failure points across infrastructure components.
Watch the full interview: https://ku.bz/Lsr8gltrH
This interview is a reaction to Miguel Luna's episode https://ku.bz/WwS04jYvv
Forwarded from LearnKube news
This week on Learn Kubernetes Weekly 164:
📊 Queue-Based Autoscaling Without Flapping: Rethinking App Scaling with Kubernetes, KEDA, and RabbitMQ
🔄 Announcing Changed Block Tracking API support
🐳 Why I Ditched Docker for Podman (And You Should Too)
🔐 That Time I Found a Service Account Token in my Log Files
☁️ Deploying a .NET Weather Forecast App to AKS Using GitHub Actions and Argo CD
Read it now: https://kube.today/issues/164
⭐️ This issue is brought to you by LearnKube — master Kubernetes with hands-on training designed for engineers who want to learn the smart way https://ku.bz/hypSbyc-V
📊 Queue-Based Autoscaling Without Flapping: Rethinking App Scaling with Kubernetes, KEDA, and RabbitMQ
🔄 Announcing Changed Block Tracking API support
🐳 Why I Ditched Docker for Podman (And You Should Too)
🔐 That Time I Found a Service Account Token in my Log Files
☁️ Deploying a .NET Weather Forecast App to AKS Using GitHub Actions and Argo CD
Read it now: https://kube.today/issues/164
⭐️ This issue is brought to you by LearnKube — master Kubernetes with hands-on training designed for engineers who want to learn the smart way https://ku.bz/hypSbyc-V
Media is too big
VIEW IN TELEGRAM
Tanat Lokejaroenlarb, Staff Site Reliability Engineer at Adevinta, explains how his team manages SHIP, Adevinta's internal multi-tenant Kubernetes platform where multiple development teams share underlying clusters with isolated namespaces.
He details their strategic transition from self-managed Kubernetes to Amazon EKS, driven by two key factors: the deprecation of kube-aws (which stopped at Kubernetes 1.15) preventing access to critical features like volume snapshots and ingress class, and the desire to shift team focus from infrastructure management to platform features.
Watch the full episode: https://kube.fmhttps://ku.bz/VVHFfXGl_
He details their strategic transition from self-managed Kubernetes to Amazon EKS, driven by two key factors: the deprecation of kube-aws (which stopped at Kubernetes 1.15) preventing access to critical features like volume snapshots and ingress class, and the desire to shift team focus from infrastructure management to platform features.
Watch the full episode: https://kube.fmhttps://ku.bz/VVHFfXGl_
KubeFM
Tanat Lokejaroenlarb, Staff Site Reliability Engineer at Adevinta, explains how his team manages SHIP, Adevinta's internal multi-tenant Kubernetes platform where multiple development teams share underlying clusters with isolated namespaces. He details their…
This episode is sponsored by LearnKube - get started on your Kubernetes journey through comprehensive online, in-person or remote training https://learnkube.com/training
This media is not supported in your browser
VIEW IN TELEGRAM
Aviv Shukron, VP Product at Komodor, shares three key trends shaping the Kubernetes ecosystem.
He identifies AI SRE as a rapidly growing space with numerous new companies emerging to automate site reliability engineering tasks.
The second trend is checkpointing and restoring of state, a capability that's gaining momentum in the Kubernetes community.
Aviv connects this directly to cost optimization, noting that this has become the primary use case driving adoption of these state management capabilities across organizations.
Watch the full interview: https://ku.bz/HZc2ftY_R
He identifies AI SRE as a rapidly growing space with numerous new companies emerging to automate site reliability engineering tasks.
The second trend is checkpointing and restoring of state, a capability that's gaining momentum in the Kubernetes community.
Aviv connects this directly to cost optimization, noting that this has become the primary use case driving adoption of these state management capabilities across organizations.
Watch the full interview: https://ku.bz/HZc2ftY_R
This media is not supported in your browser
VIEW IN TELEGRAM
Nicholas Eberts, Product Manager at Google, discusses the ongoing debate of multi-tenancy in Kubernetes clusters.
Nicholas emphasizes that there's no one-size-fits-all solution: organizations might need single-tenant clusters for critical workloads while using multi-tenant setups for smaller applications.
Watch the full interview: https://ku.bz/3h7FQWLt1
This interview is a reaction to Artem Lajko's episode https://ku.bz/zp0L7-xM4
Nicholas emphasizes that there's no one-size-fits-all solution: organizations might need single-tenant clusters for critical workloads while using multi-tenant setups for smaller applications.
Watch the full interview: https://ku.bz/3h7FQWLt1
This interview is a reaction to Artem Lajko's episode https://ku.bz/zp0L7-xM4
Media is too big
VIEW IN TELEGRAM
Zain Malik, Software @ Exostellar, describes the massive multi-tenant Kubernetes infrastructure he previously managed at City Storage Systems.
Their environment ran approximately 30,000 pods across different clusters, with sizes ranging from 400 to 950 nodes (the upper limit imposed by AKS constraints). Zain explains how pod density varied dramatically based on workload requirements—some nodes hosted just 1-2 pods while others were densely packed, allowing them to run up to 10,000 pods on 500 nodes during optimal conditions.
Watch the full episode: https://ku.bz/5PLksqVlk
Their environment ran approximately 30,000 pods across different clusters, with sizes ranging from 400 to 950 nodes (the upper limit imposed by AKS constraints). Zain explains how pod density varied dramatically based on workload requirements—some nodes hosted just 1-2 pods while others were densely packed, allowing them to run up to 10,000 pods on 500 nodes during optimal conditions.
Watch the full episode: https://ku.bz/5PLksqVlk
KubeFM
Zain Malik, Software @ Exostellar, describes the massive multi-tenant Kubernetes infrastructure he previously managed at City Storage Systems. Their environment ran approximately 30,000 pods across different clusters, with sizes ranging from 400 to 950 nodes…
This episode is sponsored by LearnKube - get started on your Kubernetes journey through comprehensive online, in-person or remote training https://learnkube.com/training
This media is not supported in your browser
VIEW IN TELEGRAM
Gari Singh, Product Manager, Google Cloud @ Google at Google Cloud, explains how recent AI conformance and agent sandbox announcements are reshaping the Kubernetes ecosystem.
Gari emphasizes that these developments position Kubernetes as the optimal platform not just for AI training and inference, but for running agentic workloads - marking a significant evolution in how autonomous AI operations can be securely deployed and managed at scale.
Watch the interview: https://ku.bz/lwYXpLtPd
Read the announcement: https://ku.bz/s9r4CKqq5
Gari emphasizes that these developments position Kubernetes as the optimal platform not just for AI training and inference, but for running agentic workloads - marking a significant evolution in how autonomous AI operations can be securely deployed and managed at scale.
Watch the interview: https://ku.bz/lwYXpLtPd
Read the announcement: https://ku.bz/s9r4CKqq5
This media is not supported in your browser
VIEW IN TELEGRAM
Itiel Shwartz, Co-Founder & CTO at Komodor, shares his perspective on effective troubleshooting in Kubernetes.
He emphasizes that the key is to learn from past incidents and understand how systems worked before issues arose. Using networking problems as an example, he explains why understanding the underlying architecture and root causes is more valuable than making quick fixes to configurations without proper context.
Watch the full interview: https://ku.bz/35z_flZn3
This interview is a reaction to Alex Movergan's episode https://ku.bz/P5Y-NrSW5
He emphasizes that the key is to learn from past incidents and understand how systems worked before issues arose. Using networking problems as an example, he explains why understanding the underlying architecture and root causes is more valuable than making quick fixes to configurations without proper context.
Watch the full interview: https://ku.bz/35z_flZn3
This interview is a reaction to Alex Movergan's episode https://ku.bz/P5Y-NrSW5
This media is not supported in your browser
VIEW IN TELEGRAM
Michael Levan explains how team silos in technology differ from traditional organizational silos.
He draws a parallel between specialized teams and military trenches, where teams work together in their specific domains.
The trench metaphor helps illustrate how teams can maintain deep expertise while still being part of a larger, coordinated effort.
Watch the full episode: https://ku.bz/qlZPfM-zr
He draws a parallel between specialized teams and military trenches, where teams work together in their specific domains.
The trench metaphor helps illustrate how teams can maintain deep expertise while still being part of a larger, coordinated effort.
Watch the full episode: https://ku.bz/qlZPfM-zr
This media is not supported in your browser
VIEW IN TELEGRAM
Shyam Jeedigunta, Principal Engineer at Amazon Web Services (AWS), discusses two tools:
- KRO (Kubernetes Resource Orchestrator), an open-source AWS project that models resources as connected groups for deployment, enabling replication and application pattern creation.
- Karpenter, an auto-scaling and node lifecycle management tool that's gained substantial adoption and is evolving to address next-generation challenges with AI/ML workloads.
-
Watch the full interview: https://ku.bz/m89tLbgcq
- KRO (Kubernetes Resource Orchestrator), an open-source AWS project that models resources as connected groups for deployment, enabling replication and application pattern creation.
- Karpenter, an auto-scaling and node lifecycle management tool that's gained substantial adoption and is evolving to address next-generation challenges with AI/ML workloads.
-
Watch the full interview: https://ku.bz/m89tLbgcq
This media is not supported in your browser
VIEW IN TELEGRAM
Stefano Doni, CTO @ Akamas, explains what sets Akamas Insights apart from other Kubernetes optimization solutions in the market.
He highlights two key differentiators: the agent-free architecture that leverages existing observability tools instead of requiring additional cluster installations, and the comprehensive full-stack optimization approach that works across multiple levels - from cluster optimization down to pod-level application tuning and even runtime optimization for JVM and Node.js environments.
Watch the interview: https://ku.bz/rdbv-kvWt
Read the announcement: https://ku.bz/mlYTxPC6x
He highlights two key differentiators: the agent-free architecture that leverages existing observability tools instead of requiring additional cluster installations, and the comprehensive full-stack optimization approach that works across multiple levels - from cluster optimization down to pod-level application tuning and even runtime optimization for JVM and Node.js environments.
Watch the interview: https://ku.bz/rdbv-kvWt
Read the announcement: https://ku.bz/mlYTxPC6x
This media is not supported in your browser
VIEW IN TELEGRAM
Stefan Roman discusses building Labs4Grabs, a Kubernetes learning platform that provides root access to users. He explains three critical security challenges:
- Implementing multi-tenancy isolation while supporting privileged workloads and root access
- Learning and configuring new isolation technologies without introducing security vulnerabilities
- Managing resource allocation to balance student experience with infrastructure constraints
The discussion explores the technical trade-offs between providing an authentic learning environment and maintaining platform security.
Watch the full episode: https://ku.bz/Xz-TrmX2F
- Implementing multi-tenancy isolation while supporting privileged workloads and root access
- Learning and configuring new isolation technologies without introducing security vulnerabilities
- Managing resource allocation to balance student experience with infrastructure constraints
The discussion explores the technical trade-offs between providing an authentic learning environment and maintaining platform security.
Watch the full episode: https://ku.bz/Xz-TrmX2F