About this Senior Kubermetes Engineer role at SSC HR Solutions
Senior Kubernetes Engineer
We are looking for a Senior Kubernetes Engineer to design, deploy, and operate production-grade Kubernetes platforms across client-owned infrastructure, including bare metal, private cloud, and sovereign environments.
The role focuses on building reliable and secure Kubernetes reference deployments that client operations teams can install, operate, maintain, and scale, including highly restricted and air-gapped environments.
Key Responsibilities:
- Design and operate production Kubernetes clusters on bare metal, private cloud, and on-premises infrastructure.
- Own the Kubernetes platform architecture from cluster design through deployment and ongoing operations.
- Define reference Kubernetes deployments that can be installed and operated by client infrastructure teams.
- Design and manage CNI networking, CSI integrations, persistent storage, ingress, and load balancing.
- Manage etcd operations, including backup, recovery, health monitoring, and disaster recovery procedures.
- Build and maintain deployment workflows using Helm and GitOps tools such as ArgoCD or Flux.
- Use Terraform to automate infrastructure and Kubernetes platform provisioning where appropriate.
- Design and implement air-gapped and restricted-network Kubernetes deployments.
- Manage container image mirroring, private registries, and offline installation processes.
- Deploy and operate stateful workloads on Kubernetes, including Kafka, Flink, Spark, databases, and their operators.
- Design Kubernetes platforms capable of supporting high availability and business-critical workloads.
- Implement security hardening for regulated and security-sensitive environments.
- Define and enforce RBAC, network policies, secrets management, and cluster security policies.
- Plan cluster capacity, node resources, scaling strategies, and workload placement.
- Design and execute disaster recovery and business continuity approaches for Kubernetes environments.
- Plan and perform Kubernetes upgrades with minimal or zero downtime.
- Support GPU scheduling and node management for AI and data-intensive workloads.
- Work with client infrastructure and operations teams during deployment, handover, and operational readiness.
- Participate in security reviews and technical architecture discussions within client environments.
- Troubleshoot complex networking, storage, scheduling, and cluster-level issues in production.
- Establish operational standards, runbooks, monitoring requirements, and troubleshooting procedures for Kubernetes platforms.
Requirements
Requirements
- Proven experience as a Senior Kubernetes Engineer, Platform Engineer, DevOps Engineer, or similar role.
- Strong hands-on experience building and operating production Kubernetes clusters on bare metal or private cloud.
- Experience with client-owned, on-premises, or sovereign infrastructure environments.
- Strong understanding of Kubernetes architecture and the components that managed Kubernetes services typically abstract away.
- Deep knowledge of Kubernetes networking and CNI behavior.
- Strong experience with CSI, persistent storage, and stateful workloads on Kubernetes.
- Strong understanding of ingress controllers, service networking, and load balancing without relying on a public cloud provider.
- Hands-on experience with etcd operations, including backup, recovery, troubleshooting, and cluster health.
- Experience with Helm and Kubernetes package management.
- Hands-on experience with ArgoCD, Flux, or similar GitOps tooling.
- Strong Terraform experience for infrastructure and platform provisioning.
- Proven experience deploying Kubernetes in air-gapped or highly restricted network environments.
- Experience with private container registries, image mirroring, and offline installation workflows.
- Experience operating stateful technologies on Kubernetes, including Kafka, Flink, Spark, databases, and their operators.
- Strong understanding of Kubernetes security, including RBAC, network policies, secrets management, and policy enforcement.
- Experience hardening Kubernetes platforms for regulated or security-sensitive environments.
- Strong understanding of PKI, certificates, TLS, and certificate lifecycle management in Kubernetes environments.
- Experience with capacity planning, high availability, disaster recovery, and cluster scaling.
- Proven experience performing Kubernetes upgrades while minimizing service disruption.
- Experience with GPU scheduling, node management, and Kubernetes workloads supporting AI or data-intensive applications.
- Strong troubleshooting skills across networking, storage, compute, scheduling, and Kubernetes control-plane components.
- Experience working directly in client data centres and participating in infrastructure or security reviews.
- Ability to create clear deployment documentation, operational runbooks, and handover materials for client operations teams.
- Strong understanding of production reliability, observability, and operational readiness for Kubernetes platforms.