Platform Support Engineering is a new team within Kraken’s Platform department. It will be the primary contact point for most day-to-day engineering interactions between the company’s 2,000 staff and the 120 engineers in the Platform department. The team will handle a high volume of tickets each month, supporting the business across a myriad of clients in the EU, North America, and Australia-Pacific regions on a 24/5 basis.
This role sits in the middle of that function, owning investigations end-to-end and beginning to shape the team's tooling and processes. You'll work with enough independence to own complex investigations, and close enough to the platform to keep developing your technical depth.
What you'll do
Act as the primary contact point for platform-related requests and non-critical incidents from engineers, Technical Account Managers, and Client Delivery Leads
Triage incoming tickets and alerts, gathering context and reproducing issues where possible
Monitor platform alerting channels and respond to non-critical alerts in line with defined processes
Resolve routine platform requests (access, configuration changes, environment queries, etc.)
Escalate complex technical issues to the relevant Platform Engineering teams for investigation and resolution, capturing their input and formalising it into runbooks for future use
Escalate to Senior Platform Support Engineers or the Platform Support Lead where pushback or seniority is required
Communicate clearly across a range of technical and non-technical audiences, translating between the two where needed
Identify recurring problems and flag opportunities to improve tooling, automation or documentation
Maintain and improve runbooks, FAQs and internal documentation
Work closely with Platform Engineers to build your understanding of how the platform is built and operated
Take ownership of ticket investigations end-to-end, including identifying when to escalate severity without prompting
Proactively flag risks and scope changes to stakeholders before they become problems
Identify and close gaps in team documentation without being asked
Begin contributing to tooling and automation improvements: scoped, well-defined pieces that reduce toil or improve the support workflow
What you'll need
Hands-on experience with at least one of: AWS (console + CLI), Kubernetes (kubectl, real triage experience), Terraform (reading plans, running applies)
Comfortable triaging platform issues independently, knowing what to look for without being told and able to investigate before escalating
Experience communicating technical problems and resolutions to non-technical stakeholders
Can engage confidently in technical conversations with Platform Engineers when investigating complex issues, contributing context and understanding the response
Proactively flags risks and scope changes to stakeholders before they become problems
Able to identify gaps in documentation and close them, not just follow what exists
It would be great if you had
Familiarity with observability tooling (Datadog or similar), using it to understand system behaviour, not just confirm alerts
Experience in a support, SRE, or platform ops role in a high-scale environment
Experience with GitOps workflows, understanding how infrastructure changes flow from code to production
Familiarity with incident management tooling (PagerDuty, incident.io, Rootly, or similar), not just receiving alerts but owning the response flow
Exposure to a client-facing or stakeholder-heavy environment (TAMs, CDLs, or similar non-engineering audiences)
Scripting ability in Python, Bash, or similar, enough to automate a repetitive task or parse logs efficiently
What success looks like in this role
You're comfortable with ambiguity. Platform Support sits at the intersection of engineering, clients, and product. Not every ticket has a clean answer. The people who do well here move through uncertainty rather than waiting for it to resolve.
You communicate before you're asked to. Stakeholders (engineers, TAMs, CDLs) care most about knowing what's happening. Proactive, accurate updates matter more than perfect resolutions.
You treat runbooks as a floor, not a ceiling. Following a runbook is the starting point. Finding the gap in it, fixing it, and making the next person's job easier is the standard we're aiming for.
You escalate with context, not just symptoms. When something needs to go up the chain, it goes up with a clear problem statement, what's been tried, and what's needed. You don't just hand off the stress.