Find partners
KubeFM

KubeFM

Hosted by KubeFM

TechnologyInterviews guests

Episodes

100

Latest episode

Aug 2026

Language

EN-US

About the show

Discover all the great things happening in the world of Kubernetes, learn (controversial) opinions from the experts and explore the successes (and failures) of running Kubernetes at scale.

Listen to episodes

60 recent
August 19, 2026Episode 436 min

GitOps at Enterprise Scale, with Elad Cohen

At enterprise scale, a deployment pipeline that runs Helm upgrades directly against Kubernetes hides drift, mixes configuration with CI logic, and makes the last pipeline run the source of truth. Elad Cohen explains how WSC Sports moved from Azure DevOps to GitHub Actions and redesigned delivery around Git and Argo CD. The resulting platform separates builds from deployments, keeps service configuration in values files, and continuously reconciles clusters. In this interview: Why CI should change Git instead of the cluster How ApplicationSets create main and shadow deployments from one values file How AppProjects scope permissions and route alerts by team Why reusable Helm contracts make customization compound across services Sponsor This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits. More info Find all the links and info for this episode here: https://ku.bz/wX5H5Mjwv Interested in sponsoring an episode? Learn more.

August 10, 2026Episode 345 min

Automating Pod Disruption Budgets with Kyverno, with Ahmad Asmar

Karpenter can reduce Kubernetes infrastructure costs, but aggressive node consolidation can also expose workloads that lack disruption safeguards. Ahmad Asmar explains how Zencity uses Kyverno to automatically generate Pod Disruption Budgets , while accounting for existing PDBs, percentage-based availability targets, single-replica workloads, and environment-specific policies. In this interview: How Karpenter consolidation changes the availability risks of cluster operations Why Kyverno's generated policies can provide safer defaults than manual enforcement How to handle duplicate PDBs, scaling workloads, and single-replica edge cases How aggregated ClusterRoles keep custom permissions separate from Helm-managed resources Sponsor This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits. More info Find all the links and info for this episode here: https://ku.bz/xrlPJg54D Interested in sponsoring an episode? Learn more.

August 4, 2026Episode 226 min

From KIAM to EKS Pod Identities, with Fabián Sellés Rosa

An unmaintained identity component can remain invisible until a routine Kubernetes upgrade turns it into an incident. Fabián Sellés Rosa , Platform Engineer and Runtime Tech Lead at Adevinta, explains how his team moved from KIAM to EKS Pod Identities without discarding the security boundaries and application interface that their internal platform depended on. In this interview: Why KIAM became urgent to replace after years of stable operation How Crossplane, a custom controller, and KRO with ACK compared against the team's criteria Why managed EKS Capabilities reduced toil but introduced observability and rollout trade-offs How Kyverno preserved namespace-level authorization for IAM roles Sponsor This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits. More info Find all the links and info for this episode here: https://ku.bz/R_06hwnCn Interested in sponsoring an episode? Learn more.

July 28, 2026Episode 147 min

1 Million Tokens Per Second on Kubernetes, with Federico Iezzi

GPU inference throughput depends on more than accelerator generation or count. Memory bandwidth, model parallelism, cache configuration, and the load generator itself all influence measured throughput. Federico Iezzi , Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs. The discussion covers: Why memory bandwidth limits decode performance How Federico chose between tensor and data parallelism What changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization. Sponsor This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits. More info Find all the links and info for this episode here: https://ku.bz/1xD9Md0mb Interested in sponsoring an episode? Learn more.

May 19, 2026Episode 1721 min

The Hidden Cost of Slow Autoscaling, with John Ford

Forced platform migrations are usually treated as something to survive. At Scout24, a mandatory OS migration became an opportunity to rethink Kubernetes autoscaling, node provisioning, and infrastructure efficiency. John Ford explains how Scout24 moved its EKS-based Infinity platform from a polling autoscaler and over-provisioned capacity to Karpenter and Bottlerocket . The result was faster node startup, a safer migration path, and about a 30% infrastructure reduction without major downtime. In this interview: Why two-minute node provisioning forced a 25% capacity buffer How Karpenter made the Bottlerocket migration safer What broke around EC2 metadata, AWS SDKs, and cgroups How the new foundation enables Spot, ARM, and GPU workloads Sponsor This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training. More info Find all the links and info for this episode here: https://ku.bz/DdmVC2_7v Interested in sponsoring an episode? Learn more.

May 12, 2026Episode 1636 min

The Namespaces Scaling Trap, with Brian Stack

Most teams scale Kubernetes by thinking about pods and nodes. At Render, Brian Stack ran into a different dimension: hundreds of thousands of namespaces per cluster, multiplied across DaemonSets that list-watch every namespace. Brian explains how Render traced the issue through Calico and Vector, worked with upstream maintainers, and turned memory profiling into operational wins: lower node costs, lighter API-server load, and faster rollouts. In this interview: Why namespaces can become a hidden scaling bottleneck How DaemonSets multiply memory and control-plane pressure How profiling, staging clusters, and upstream collaboration freed 7 TiB Why pushing from an 80% fix to a complete fix can make teams faster Sponsor This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training. More info Find all the links and info for this episode here: https://ku.bz/0mrvCsXrV Interested in sponsoring an episode? Learn more.

May 5, 2026Episode 1538 min

AI Agents Running Kubernetes, with Mike Solomon

What happens when an AI agent stops generating Kubernetes YAML and starts operating the cluster directly? Mike Solomon , software engineer at AIATELLA, explains how his team moved from a sprawling Helm setup to Markdown-driven infrastructure specs that Claude Code can execute, test, and refine. You will learn Why Helm became hard to maintain for a fast-moving medical infrastructure repo How Claude debugged Argo, TLS conflicts, kubectl patches, and private registry credentials How runbooks plus agent memory files capture failures so deployments become reproducible. It is a practical look at where Kubernetes automation may be heading: less hand-written YAML, more precise intent, and a sharper definition of when the human must stay in the loop. Sponsor This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training. More info Find all the links and info for this episode here: https://ku.bz/y70mLvWNs Interested in sponsoring an episode? Learn more.

April 28, 2026Episode 1435 min

SaaS with Kubernetes Operators and Garbage Collection, with Alexander Held

A single Kubernetes CRD for every service request turns small changes into full-platform reconciliations. Alexander Held , former platform engineer at Mercedes-Benz Tech Innovation, describes a production refactor from a 2,000-line CRD to purpose-built resources and controllers. He shows how teams can model business workflows as Kubernetes APIs and then use owner references , finalizers , and events to keep platform operations predictable. You will learn: Why monolithic CRDs create performance and troubleshooting problems How controllers turn database provisioning and backups into reconciliation loops How finalizers clean up external resources such as S3 backups Why Kubernetes events make platform workflows easier to debug Sponsor This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training. More info Find all the links and info for this episode here: https://ku.bz/TGy4Qn7Qs Interested in sponsoring an episode? Learn more.

April 21, 2026Episode 131 hr 29 min

What Hip-Hop Can Teach Us About Kubernetes, with Kelsey Hightower, Eric Abercrombie, and Julius Payne II

Kelsey Hightower , Eric Abercrombie , and Julius Payne II reflect on life after achievement, entering the Kubernetes world for the first time, and how music, creativity, and lived experience shape the way they think about technology. In this interview: Why fundamentals , patience, and repetition still matter more than shortcuts How Kubernetes , community, and confidence intersect for people entering cloud-native work What hip-hop, production, and storytelling can teach us about ownership , authenticity , and finding your voice Sponsor This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training. More info Find all the links and info for this episode here: https://ku.bz/czrCCXSLt Interested in sponsoring an episode? Learn more.

April 7, 2026Episode 1230 min

Intelligent Kubernetes Load Balancing, with Rohit Agrawal

You're running gRPC services in Kubernetes, load balancing looks fine on the dashboard — but some pods are burning at 80% CPU while others sit idle, and adding more replicas only partially helps. Rohit Agrawal , a Staff Software Engineer on the traffic platform team at Databricks, explains why this happens and how his team replaced Kubernetes's default networking with a proxy-less, client-side load-balancing system built on the xDS protocol. In this episode: Why KubeProxy's Layer 4 routing breaks down under high-throughput gRPC: it picks a backend once per TCP connection, not per request How Databricks built an Endpoint Discovery Service (EDS) that watches Kubernetes directly and streams real-time pod metadata to every client How zone-aware spillover cut cross-availability-zone costs without sacrificing availability Why CPU-based routing failed (monitoring lag creates oscillation) and what signals to use instead The system has been running in production for three years across hundreds of services, handling millions of requests. Sponsor This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training. More info Find all the links and info for this episode here: https://ku.bz/y803JMhBk Interested in sponsoring an episode? Learn more.

Is this your show?

Claim this listing to keep it up to date, reach guests who want to pitch you, and manage bookings with Guestify.

Claim this listing

More Technology podcasts