รber diese Senior Platform Engineer - Cloud & Kubernetes Stelle bei Weekday AI
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ฎ๐ฒ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฐ๐ฑ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ฎ๐ฒ-๐ฐ๐ฑ ๐๐ฃ๐)
Experience: 5+ yrs
Location: Abu Dhabi, United Arab Emirates
Job Type: Full-time
We are looking for an experiencedย Senior Platform Engineerย to design, implement, secure, and operate scalable cloud-native platforms usingย Microsoft Azure and Kubernetes. The role focuses on building highly available, resilient, secure, and well-governed production environments while driving automation, observability, infrastructure-as-code, and platform engineering best practices.
The ideal candidate will have strong hands-on expertise acrossย Azure, AKS, Kubernetes, Terraform, CI/CD, cloud governance, security, observability, and Disaster Recovery, along with the ability to provide technical leadership across complex platform environments.
Requirements
Key Responsibilities
- Design, deploy, and manageย Microsoft Azure infrastructure and AKS/Kubernetes platforms.
- Implement and maintainย Disaster Recovery, backup, restore, and business continuityย solutions.
- Build infrastructure automation usingย Terraform, Helm, Azure DevOps, and CI/CD pipelines.
- Manage Kubernetes networking, ingress, storage, namespaces, resource limits, security, and platform configurations.
- Implement observability solutions usingย Prometheus, Grafana, Loki, and related monitoring technologies.
- Manage TLS certificates, secrets, access controls, RBAC, and platform security mechanisms.
- Establish cloud and Kubernetes governance covering resource organisation, naming, tagging, security policies, and operational standards.
- Defineย Infrastructure-as-Code and CI/CD governance, including reusable Terraform modules, state management, code reviews, approval gates, environment promotion, and artifact management.
- Participate in architecture and technical design reviews for platforms, applications, integrations, and infrastructure changes.
- Drive security and compliance readiness through vulnerability remediation, access reviews, security baselines, audit controls, and policy enforcement.
- Define and monitor platformย availability, SLIs, SLOs, capacity, performance, and operational health.
- Lead incident and problem management, including root-cause analysis, corrective actions, and prevention of recurring issues.
- Perform capacity planning, performance optimisation, andย cloud cost optimisation / FinOpsย activities.
- Own DR testing, RTO/RPO validation, backup and recovery standards, and periodic recovery exercises.
- Plan and execute platform migrations, infrastructure upgrades, and cloud transformation initiatives.
- Evaluate emerging platform technologies, conduct POCs, and establish approved patterns for production adoption.
- Maintain technical documentation includingย HLDs, LLDs, architecture diagrams, SOPs, runbooks, troubleshooting guides, and DR procedures.
- Provide technical leadership, mentoring, and knowledge sharing across platform engineering teams.
What Makes You a Great Fit
- 5+ years of experienceย in platform engineering, cloud infrastructure, DevOps, SRE, or related roles.
- Strong hands-on expertise inย Microsoft Azure and Kubernetes, particularly AKS.
- Strong experience withย Docker, Helm, Terraform, and Azure DevOps / CI/CD.
- Good understanding of Azure networking, includingย VNets, NSGs, Private Endpoints, and Azure Firewall.
- Experience implementing observability usingย Prometheus, Grafana, Loki, or similar platforms.
- Strong knowledge ofย Linux, Bash scripting, Git, and GitOps practices.
- Proven experience withย Disaster Recovery, backup/restore, RTO/RPO planning, and recovery testing.
- Strong understanding of Kubernetes security, governance, RBAC, policy enforcement, and Azure Policy.
- Experience establishing reusableย Infrastructure-as-Code patterns and platform engineering standards.
- Strong understanding of CI/CD governance, release controls, environment promotion, and source-control standards.
- Experience working with databases such asย PostgreSQL, MongoDB, MySQL, or Azure SQL.
- Strong knowledge of incident management, problem management, RCA, capacity planning, performance optimisation, and cloud cost management.
- Experience working inย security-, compliance-, audit-, or governance-controlled environments.
- Ability to create and maintain HLDs, LLDs, architecture diagrams, operational runbooks, SOPs, and technical documentation.
- Strong experience participating in architecture and technical design reviews.
- Excellent troubleshooting, analytical, and problem-solving capabilities.
- Strong technical leadership, mentoring, communication, and cross-functional collaboration skills.
- Experience inย banking or other regulated industriesย would be an advantage.
- Strong understanding of high availability, production operations, platform modernisation, and cloud transformation initiatives.