Jobs Companies AIFT Senior Site Reliability Engineer, Vulcan (AI Security product)

À propos de ce poste Senior Site Reliability Engineer, Vulcan (AI Security product) chez AIFT

AIFT · Sur site · UAE

Job Overview

We are looking for a hands-on infrastructure engineer to own the deployment, migration and troubleshooting of on-premise Kubernetes environments for enterprise and government clientsincluding airgapped, high-security data center environments where remote access is not possible.  

This is a client-facing, on-site role: you will be the technical authority in the room, responsible for executing complex infrastructure changes correctly the first time, diagnosing failures independently under pressure and communicating clearly with client stakeholders throughout. 

This role carries real ownership; you will be expected to understand the systems deeply enough to make sound judgment calls when things don't go to plan, without waiting on remote support.

In this role, you will play a vital part in supporting our Cybersecurity business, Vulcan. Vulcan is a cybersecurity solution for GenAI, providing red and blue team services to ensure compliance and security. 

Learn more about us 👉

 

Responsibilities

  • Plan and execute on-prem Kubernetes cluster deployments, upgrades and infrastructure migrations (including IP re-addressing, certificate rotation and cluster reconfiguration) in production and airgapped environments 
  • Diagnose and resolve failures independently on-site 
  • Own the full infrastructure stack end-to-end: Kubernetes control plane and data plane, PostgreSQL (primary/replica replication), distributed storage (e.g. SeaweedFS/Ceph/similar), private container registries and centralized logging (ELK or equivalent)
  • Validate deployment tooling (scripts, installers, automation) thoroughly in lab/staging environments before any client-facing execution 
  • Represent the technical work directly to client stakeholders on-site: explain status, failures and remediation plans clearly 
  • Travel to client data centers (including airgapped/restricted-access sites) as required, sometimes on short notice, for deployment and go-live support
  • Write clear, structured runbooks, decision trees and incident reports that others (including less experienced engineers) can follow under pressure 
  • Escalate risks proactively to internal leadership, not just after something has gone wrong 

Requirements

Technical: 

  • 5-6 years of hands-on experience with Kubernetes in production, including at least one on-premise (not purely cloud-managed) deployment 
  • Solid understanding of etcd internals. Quorum, peer membership, failure recovery, not just kubectl-level familiarity 
  • Experience with kubeadm-based cluster bootstrapping and certificate management (SANs, CA rotation, renewal) 
  • Working knowledge of PostgreSQL replication, Linux networking fundamentals (DNS, NTP, firewalls) and container registries (Docker Distribution or similar) 
  • Comfortable working entirely from the Linux command line, writing and debugging bash scripts and reading unfamiliar automation tooling under time pressure 
  • Experience with at least one distributed storage system (SeaweedFS, Ceph, MinIO or similar) is a strong plus 
  • GPU-enabled Kubernetes nodes (NVIDIA device plugin, container toolkit) experience is a plus, not required 

Working style:

  • Demonstrated ability to work independently in high-pressure, high-stakes environments without live support 
  • Strong incident communication. Can explain technical failures to non-technical stakeholders factually and calmly, without over-promising or minimizing 
  • A track record of validating changes in test environments before touching production and the judgment to insist on this even under deadline pressure 
  • Comfortable with travel, including to secure/restricted facilities where personal devices, internet access or remote assistance may not be available

Nice to have: 

  • Prior consulting, systems integration or professional services experience, ideally on enterprise or government accounts 
  • Experience specifically in the GCC/Middle East region, or with government-sector clients 
  • Security background (the ability to reason about access controls, credential handling and airgapped operational discipline is valuable given the environments involved) 

Interview Process

  • HR phone interview: 1 hour
  • Online interview: 1.5 - 2 hours, meet with hiring manager
  • Online interview: 1 hour, meet with hiring team

Why Join Us? 

  • Innovative Environment: Be part of a company at the forefront of technology to provide security in GenAI, with opportunities to work on groundbreaking projects. 
  • Growth Opportunities: Take your career to new heights with our career development programs and growth-focused culture. 
  • Dynamic Team: Join a multi-cultural and dynamic team of dedicated professionals who inspire and support each other. 
  • CompensationCompetitive salary and benefits package, commensurate with experience and performance. 

 

Prêt à postuler chez AIFT ?
Postuler chez AIFT

À propos de AIFT

Established in 2016, the group’s vision is to “Secure the Future”, a future that will increasingly be shaped by artificial intelligence.  AIFT provides services across key markets in Asia and the Middle East. We are continuously expanding our global footprint and actively recruiting international talent to join our growing team.

Voir tous les emplois chez AIFT →

Emplois similaires

Anduril Industries
Technical Site Reliability Engineer
Anduril Industries
⚡ Postuler tôt Abu Dhabi, United Arab Emirate... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
AIFT
Senior Site Reliability Engineer, Vulcan (AI Security product)
AIFT
⚡ Postuler tôt Taipei / Japan / UAE Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 mois
Dyson
Lead Reliability Engineer
Dyson
⚡ Postuler tôt China - Shanghai Office Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 4 h
B&G Foods
Sr. Reliability Engineer
B&G Foods
⚡ Postuler tôt Ankeny, IA Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
Morningstar
Senior Site Reliability Engineer
Morningstar
⚡ Postuler tôt Toronto Hybride $90,489–$132,711
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
DuPont
Reliability Engineer - Vibration
DuPont
⚡ Postuler tôt Richmond, Virginia Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 6 h
Ajaib
Senior/Staff DevOps Engineer / SRE
Ajaib
⚡ Postuler tôt Jakarta, Jakarta, Indonesia Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 9 h
Tailor
SRE (Site Reliability Engineer)
Tailor
⚡ Postuler tôt Tokyo Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 10 h
Tailor
SRE (Site Reliability Engineer) [業務委託]
Tailor
⚡ Postuler tôt Tokyo Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 10 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez AIFT

Voir tous les emplois chez AIFT →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit