À propos de ce poste Senior Platform Engineer (Agentic) - #35290 chez Manila Recruitment
As a Senior Platform Engineer (Agentic), you will be responsible for building and evolving the Azure platform that engineering teams use to develop, deploy, and run their services. You’ll design and provision secure, scalable, resilient infrastructure as code using Terraform and Azure-native technologies, while ensuring environments are reproducible, well-governed, and ready for production.
The role is focused on making the platform reliable, secure, performant, and cost-efficient while enabling delivery teams to self-serve their infrastructure and deployments. You’ll also help establish safe practices for using AI agents to develop infrastructure and scripts, including the environments, permissions, guardrails, and review processes needed to ensure automated changes are controlled and auditable.
Duties and Responsibilities:
Cloud and Infrastructure Craft
- You bring the depth the platform runs on: Azure networking including virtual networks, private endpoints, DNS and load balancing; identity, managed identities and role based access; compute across app services, containers and Kubernetes; storage and data services; Windows and Linux hosting; certificates and TLS. You are expected to know how these fail as well as how they are configured, and to design for resilience, scale and cost efficiency rather than for the default.
Infrastructure as Code
- Own the estate as code. Terraform is the primary tool: reusable modules, remote state, clean plan and apply discipline, drift detected and corrected rather than tolerated, and environments that differ by variables rather than by history. Bicep or ARM where the Azure native path is better. Everything is version controlled and reviewed before it is applied.
Configuration Management and Automation
- Own what lives inside the machines and the repeat work around them. Configuration management with Ansible, Chef or an equivalent where it is the right tool, and scripting in PowerShell, Bash or Python written to the standard you would ship. Manual steps are treated as defects to be removed rather than as procedures to be documented
Environments and Developer Experience
- Give the delivery teams the infrastructure they work on: environments they can get without waiting, templates and modules that make the right thing the quickest thing, and documentation that means they do not have to ask you twice. Measure yourself on what teams can do without you.
Pipelines and Safe Change
- Own the paths to production for infrastructure and applications: build, test, release, and the deployment approach that makes a bad change survivable. Rollback is a capability you provide, not a plan someone improvises during an incident
Security, Access and Standards
- Hold identity, access control, secrets and network boundaries across the estate, with least privilege and policy as code applied as a matter of course. Work to the shared Azure standards rather than a local variation, and contribute back to them so the estate stays consistent as it grows
Agents on the Platform
- Two sides to this. Use coding agents to produce infrastructure code and scripts, and review what comes back with the same scepticism you would apply to any change that can take an environment down. And provide the infrastructure agentic development itself needs: execution environments and sandboxes, tool servers, credential handling, and clear limits on what an automated change may touch without a person agreeing to it.
Requirements
- Deep cloud infrastructure engineering on Azure. Networking including virtual networks, subnets, private endpoints, DNS and load balancing; identity and role based access; compute across app services, containers, Kubernetes and virtual machines; storage and data services; Windows and Linux administration; certificates and TLS. This is the requirement we do not flex on.
- Infrastructure as code, with Terraform as the primary tool. Module design, remote state, plan and apply discipline, drift detection, and a pull request workflow for infrastructure change. Bicep or ARM for Azure native work. The tool matters less than the habit of building everything as reviewed, version controlled code.
- Configuration management and scripting. Ansible, Chef, Puppet or an equivalent for what runs inside the machines, and PowerShell, Bash or Python written to a standard you would ship. Containers and Kubernetes packaging with Helm or similar.
- Pipelines, security and operability. CI/CD with Azure DevOps, GitHub Actions or equivalent, safe deployment approaches and rollback that has been used in anger, policy as code, least privilege and secrets management, and enough observability practice to hand over something the reliability engineers can actually run
- Knowledge and awareness of agentic engineering: what these tools are, where they add value, and where an automated change must not go unsupervised. Hands- on experience, including the operational side of running agents safely with sandboxes, permissions and credential handling, is beneficial rather than required.