Sobre este puesto de Senior Platform Engineer en ZILO
At ZILO™, we're redefining what’s possible in technology. ZILO™ is the UK-based FinTech specialising in global asset and wealth management software, designed to scale and transform businesses of all types using our own developed AI Technology. Our mission is to digitalise the future of the global asset management industry.
We are a team of experts with decades of combined experience at leading firms globally, who thrive in fast-paced environments and want to shape the future of technology. Every individual plays a key role in driving progress and making a real impact. We continuously strive to innovate and improve.
Why work with us? At ZILO™, you'll be part of a dynamic and inclusive environment where creativity thrives. We offer the opportunity to work on cutting-edge technology, collaborate with talented individuals, and contribute to projects that have a real-world impact. We value continuous learning, personal growth, and providing our team with the resources they need to succeed.
Requirements
Key Responsibilities
Technical Leadership
- Provide senior technical leadership within the Platform Engineering team, remaining hands-on with the technology and setting a high standard for engineering excellence.
- Provide technical guidance and mentorship to Platform Engineers, fostering a culture of collaboration, continuous learning and engineering excellence.
- Work closely with Site Reliability Engineering and Software Engineering teams to design and evolve scalable, resilient and easy-to-use platform capabilities.
- Champion Platform Engineering principles, infrastructure automation, Infrastructure as Code and self-service engineering practices across the engineering organisation.
- Influence technical direction and platform architecture through design reviews, technical discussions and engineering standards.
- Encourage the adoption of AI-assisted engineering practices, enabling engineers to deliver more effectively while maintaining high standards of quality, security and reliability.
Platform Engineering
- Design, build and evolve ZILO's cloud platform and shared infrastructure capabilities.
- Develop platform services and capabilities that enable Software Engineering teams to build, test and deliver applications efficiently.
- Own and maintain non-production platform environments, ensuring they are reliable, consistent and representative of the capabilities required by engineering teams.
- Work closely with SRE to ensure platform designs and changes are suitable for reliable operation in production.
- Develop reusable platform components, patterns and abstractions that reduce complexity for engineering teams.
- Identify and remove platform bottlenecks that impact developer productivity and software delivery.
- Evaluate new technologies and approaches that can improve the scalability, security, maintainability and efficiency of the platform.
Cloud Infrastructure & Kubernetes
- Design and engineer cloud infrastructure primarily within AWS.
- Build and evolve Kubernetes-based platform capabilities using Amazon EKS.
- Develop reusable Terraform modules and Infrastructure as Code patterns that enable consistent and repeatable infrastructure provisioning.
- Design scalable Kubernetes architectures and platform services that can support future business and engineering growth.
- Improve Kubernetes scheduling, resource management and scalability, including the use of technologies such as Karpenter.
- Implement appropriate container security, networking, identity and access management controls.
- Collaborate with SRE on platform architecture and operational requirements for capabilities that will ultimately support production workloads.
Developer Platform & Self-Service
- Build platform capabilities that enable Software Engineering teams to provision and manage development infrastructure safely and independently.
- Develop self-service workflows that reduce dependencies on Platform and SRE teams for routine engineering activities.
- Create reusable templates, tooling and platform abstractions that simplify common engineering tasks.
- Improve the developer experience by reducing friction across local development, testing, integration and deployment workflows.
- Work closely with Software Engineering teams to understand developer pain points and translate them into platform capabilities.
- Establish clear platform interfaces, documentation and paved-road approaches that encourage consistent engineering practices.
Infrastructure Automation
- Eliminate manual infrastructure processes through automation and engineering-led improvements.
- Design, develop and maintain tooling that improves engineering productivity and platform efficiency.
- Build reusable automation for infrastructure provisioning, configuration and lifecycle management.
- Develop and maintain Infrastructure as Code using Terraform and associated tooling.
- Improve infrastructure testing, validation and deployment processes.
- Identify opportunities to simplify engineering workflows through automation and self-service capabilities.
- Ensure infrastructure changes are repeatable, testable, auditable and version controlled.
CI/CD & Software Delivery
- Design, build and improve CI/CD capabilities using GitHub Actions and associated tooling.
- Develop reusable pipelines and workflows that enable engineering teams to build, test and deliver software consistently.
- Improve the security, performance and maintainability of software delivery pipelines.
- Build automated validation and testing into infrastructure and platform delivery processes.
- Work with Software Engineering and SRE teams to establish consistent deployment patterns and engineering standards.
- Reduce software delivery friction through reusable tooling, automation and platform capabilities.
Observability & Platform Insights
- Build observability capabilities into the platform using metrics, logs and distributed tracing.
- Develop reusable observability patterns and tooling using technologies such as Grafana, Prometheus, OpenTelemetry and ClickHouse.
- Enable Software Engineering and SRE teams to instrument applications and services consistently.
- Improve visibility into non-production platform health, capacity, performance and resource utilisation.
- Develop dashboards and telemetry that provide insight into platform usage and developer experience.
- Work with SRE to ensure platform capabilities provide the telemetry required for effective production operation.
AI-Assisted Engineering
- Identify opportunities to leverage AI to improve Platform Engineering efficiency, infrastructure automation and developer productivity.
- Develop and adopt AI-assisted tooling to support infrastructure development, troubleshooting and platform engineering workflows.
- Explore the use of AI agents and agentic workflows to automate repetitive engineering activities.
- Use AI-assisted engineering tools such as Claude Code, GitHub Copilot, ChatGPT or equivalent as part of day-to-day engineering.
- Establish practical patterns for using AI safely and effectively when working with infrastructure and engineering systems.
- Champion responsible adoption of AI across the engineering organisation.
Non-Production Environments
- Own the engineering and evolution of ZILO's non-production platform environments.
- Ensure development, integration and testing environments are scalable, consistent and fit for purpose.
- Improve environment provisioning and lifecycle management through Infrastructure as Code and automation.
- Enable engineering teams to create and use environments efficiently while maintaining appropriate governance and cost controls.
- Improve consistency between environments through reusable infrastructure patterns and configuration management.
- Identify opportunities for ephemeral or dynamically provisioned environments where they improve engineering productivity.
- Optimise non-production infrastructure utilisation and cloud costs without compromising developer productivity.
Security & Governance
- Embed security and compliance requirements into platform architecture and automation.
- Implement secure-by-default infrastructure patterns and reusable platform components.
- Work with Security and engineering teams to improve cloud security, container security and infrastructure governance.
- Ensure platform infrastructure and configuration are auditable and managed through appropriate engineering controls.
- Identify and remediate infrastructure vulnerabilities and technical debt within platform-owned non-production environments.
- Support engineering requirements associated with operating within a regulated Financial Services environment.
Technical Collaboration
- Work closely with SRE to ensure platform capabilities can be operated effectively when adopted within production environments.
- Collaborate with Software Engineering teams to understand application requirements and provide appropriate platform capabilities.
- Participate in architecture and technical design discussions across engineering.
- Review infrastructure designs, Terraform changes and platform engineering proposals.
- Provide technical mentorship and support to other engineers.
- Document platform architecture, engineering standards, reusable patterns and technical decisions.
- Share knowledge through technical sessions, documentation, pairing and engineering communities of practice.
Continuous Improvement
- Promote a culture of continuous improvement, engineering excellence and platform ownership.
- Identify opportunities to improve developer experience, infrastructure delivery and engineering efficiency.
- Reduce platform complexity and technical debt through pragmatic engineering improvements.
- Improve platform scalability, maintainability and security through automation and engineering best practices.
- Use engineering metrics and developer feedback to identify areas where the platform can better support Software Engineering and SRE teams.
Required Skills & Experience
Essential
- Significant experience in Platform Engineering, Infrastructure Engineering, DevOps or Site Reliability Engineering.
- Strong hands-on experience designing and building cloud infrastructure and engineering platforms.
- Experience operating as a Senior Engineer and providing technical leadership and mentorship to other engineers.
- Strong AWS experience.
- Strong Kubernetes experience (Amazon EKS preferred).
- Strong Infrastructure as Code experience (Terraform preferred).
- Experience designing reusable Terraform modules and infrastructure patterns.
- Experience building and maintaining CI/CD pipelines using GitHub Actions or similar.
- Experience designing developer platforms, self-service infrastructure or reusable engineering capabilities.
- Experience with observability technologies such as Grafana, Prometheus, OpenTelemetry and ClickHouse.
- Strong understanding of Linux, containers and cloud infrastructure.
- Strong understanding of networking, DNS, TLS, load balancing and cloud networking.
- Experience with infrastructure security, IAM and secure cloud architecture.
- Experience automating infrastructure and engineering workflows using Python, Go, Bash or similar.
- Experience working with highly available, customer-facing SaaS platforms and understanding the platform requirements necessary to support them.
- Experience using modern AI-assisted engineering tools and agents such as Claude Code, GitHub Copilot, ChatGPT or similar to improve engineering productivity and software delivery.
- Excellent communication and stakeholder management skills.
- Strong collaboration skills and experience working across Platform, SRE and Software Engineering teams.
Desirable
- Financial Services or FinTech experience.
- Experience engineering platforms within regulated environments.
- Experience with Karpenter.
- Experience with container security.
- Strong OpenTelemetry experience.
- Experience building Internal Developer Platforms or developer self-service capabilities.
- Experience with ephemeral environments or automated environment provisioning.
- Experience designing reusable GitHub Actions workflows.
- Experience with policy-as-code and infrastructure governance.
- AWS Professional or Specialty certifications.
- Kubernetes certifications (CKA preferred).
- Experience with Chaos Engineering and resilience testing as part of platform validation.
- Experience building or working with AI agents, AI-assisted automation or agentic engineering workflows.
Technical Stack
You'll ideally have experience with many of the following technologies:
- AWS
- Kubernetes (Amazon EKS)
- Terraform
- GitHub
- GitHub Actions
- Grafana
- Prometheus
- OpenTelemetry
- ClickHouse
- Linux
- Docker
- Python, Go or Bash
- Networking & DNS
- IAM and cloud security
- Infrastructure as Code
- Developer self-service platforms
- AI-assisted engineering tools (Claude Code, GitHub Copilot, ChatGPT or equivalent)
Benefits
- 23 Annual days holiday (Start and Fix at 23 days)
- 15 Public Holidays
- Provident Fund
- Health insurance (including immediate family)