Kubernetes platform operations
Cluster lifecycle, upgrades, control-plane design, workload placement, node behavior, and day-2 operations.
I work on Kubernetes platforms and the systems around them: Linux, networking, storage, observability, security controls, GitOps workflows, and production operations. My focus is building platforms that engineering teams can safely use, debug, and operate.
I design, operate, and troubleshoot Kubernetes-based platforms. The work usually sits between infrastructure and developer experience: clusters, access, policies, storage, networking, delivery workflows, monitoring, and incident response.
Cluster lifecycle, upgrades, control-plane design, workload placement, node behavior, and day-2 operations.
CNI behavior, Gateway API, service discovery, DNS, MTU, routing, load balancing, and NetworkPolicy troubleshooting.
Persistent storage for Kubernetes workloads, Rook-Ceph operations, failure domains, recovery, and storage-related incidents.
Metrics, logs, traces, dashboards, alerting, capacity signals, and debugging workflows for production platforms.
Security is part of the platform work, not a separate checklist. I work on controls that make Kubernetes safer to use without blocking engineering teams from shipping.
I usually work close to these systems and components.