About this Practice Manager - Platform Reliability, Operations Hub & Automation role at Datacom
Our Why
Datacom works with organisations and communities across Australia and New Zealand to make a difference in people’s lives and help organisations use the power of tech to innovate and grow.
About the role (your why)
We're seeking an exceptional technology leader to transition from manual operations to an automated and AI led Mode 2 platform engineering model for our Reliability Operations Hub (ROH), a critical capability at the centre of our transformation from traditional IT operations to an AI-enabled, automation-first Platform Engineering model.
This is more than an operational leadership role. You'll lead the evolution of a 24x7 Reliability Operations Hub that not only drives operational excellence across complex hybrid and cloud environments but also helps shape the future commercialisation of our operational services through an innovative "SRE-as-a-Service" model.
You'll combine strategic vision, technical expertise, people leadership, and commercial acumen to deliver measurable outcomes including significant toil reduction, enhanced observability, accelerated incident resolution, and increased automation across the enterprise.
In this role you will
Lead 24x7 Technical Operations
- Lead and scale regional 24x7 technical operations, ensuring effective follow-the-sun support, triage, handovers, and operational excellence.
- Own and continuously optimise the ROH Front Door operating model, ensuring efficient intake, routing, segmentation, and tracking of operational work.
- Apply Lean principles to improve operational flow, reducing queue wait times, cycle times, and delivery bottlenecks.
- Define clear engagement models between the ROH and specialist engineering teams, protecting engineering capacity from low-value operational noise.
- Drive integrated customer outcomes across modern platform environments.
Drive SRE & Automation Excellence
- Lead Automation Swarms focused on converting repetitive operational work into reusable automated solutions using Ansible platform.
- Champion Site Reliability Engineering (SRE) practices including Service Level Objectives (SLOs) and Error Budget frameworks.
- Expand self-service capabilities and automate common operational tasks through standardised Golden Paths.
- Drive adoption of automation-first and AI-assisted operational practices across the organisation.
Own Major Incident & Platform Resilience
- Act as escalation leader during critical incidents, ensuring rapid service restoration and effective executive communication.
- Lead root cause analysis and problem management processes that convert operational learnings into long-term improvements.
- Ensure high-severity incidents result in identified and prioritised preventative automation opportunities.
Build Operational Standards & Knowledge Management
- Govern the development and continuous improvement of operational runbooks and AI-ready documentation.
- Reduce operational variance through standardisation, automated patching, role-based access controls, and gold-standard platform configurations.
- Foster a culture focused on knowledge sharing and continuous improvement.
Shape Commercial Service Growth
- Lead the transition of operational capabilities from a traditional cost centre to a scalable, productised service offering.
- Partner with product and commercial teams to package observability, automated triage, and SRE capabilities into customer-facing services.
- Drive operational efficiency, service margin growth, and the creation of repeatable, high-value offerings.
Inspire Teams & Transformation
- Lead, coach, and develop a distributed Platform Reliability Engineering (PRE) team across multiple regions and drive process standardisation and unification.
- Champion cross-skilling of PRE’s to build capabilities across infrastructure, cloud, database, middleware, and platform technologies.
- Champion a culture of psychological safety, accountability, innovation, and continuous improvement.
- Lead organisational change initiatives that shift teams from reactive operations to engineering-led automation practices.
- Support right-shoring strategies across onshore, nearshore, and offshore delivery teams.
What You'll Bring
Experience
- 15+ years' experience in large-scale technology operations environments.
- At least 3 years leading Site Reliability Engineering, Platform Engineering, or similar operational engineering teams.
- Proven success leading 24x7 operational functions across multiple regions and time zones.
- Experience delivering transformational operating model change and driving automation-first ways of working.
- Commercial leadership experience with service-based delivery models, service pricing, margin optimisation, and operational economics.
- Strong experience managing major incidents, crisis response, and enterprise resilience programmes.
Technical Expertise
- Deep understanding of SRE principles, platform reliability, and operational engineering.
- Experience with enterprise observability platforms and ITSM tooling.
- Knowledge of automation and orchestration technologies, including Ansible Automation Platform and AI-assisted operational workflows.
- Strong understanding of APIs, systems integration, scripting, coding, and database automation.
- Experience working across hybrid technology environments, including:
- Windows and Linux platforms
- Databases
- Middleware technologies
- Containers and virtualisation
- AWS, Azure, and Google Cloud
- Familiarity with CI/CD, version control, cloud-native operations, automation frameworks, and modern infrastructure platforms.
Culture and Benefits
Datacom is one of Australia and New Zealand’s largest suppliers of Information Technology professional services. We have managed to maintain a dynamic, agile, small business feel that is often diluted in larger organisations of our size. It's our people that give Datacom its unique culture and energy that you can feel from the moment you meet with us.
We care about our people and provide a range of perks such as social events, chill-out spaces, remote working, flexi-hours and professional development courses to name a few. You’ll have the opportunity to learn, develop your career, connect and bring your true self to work. You will be recognised and valued for your contributions and be able to do your work in a collegial, flat-structured environment.