Sobre esta vaga de Senior Platform Engineer - Platform Metal | UK | Remote na Grafana Labs
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions. Today, more than 35 million users and 7,000+ customers – including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce – trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly, and optimize their telemetry to reduce noise and cost. We are a 100% remote company with 1,600+ team members across 40+ countries, and we’re backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital. Learn more at grafana.com and follow us on LinkedIn and X.
We’re scaling fast and staying true to what makes us different: an open-source legacy, a global collaborative culture, and a passion for meaningful work. Our team thrives in an innovation-driven environment where transparency, autonomy, and trust fuel everything we do.
You may not meet every requirement, and that’s okay. If this role excites you, we’d love you to raise your hand for what could be a truly career-defining opportunity.
This is a remote position. We are looking for candidates in the UK, Spain and Ireland only.
The Opportunity:
About our Platform:
Grafana Cloud moves millions of metrics, log lines, and traces per second from our customers' environments into a highly available, low-latency stack that processes and stores this data, and serves them to dashboards and alerting tools. We aim to grow this to hundreds of millions per second, and it's critical that as we grow, we improve our performance, increase our reliability, and, of course, do it efficiently and effectively.
The Internal Engineering Platform (IEP) delivered by the Foundations group provides application engineers with the tools, systems and Kubernetes clusters they need to build, deploy and run their workloads.
We are hiring an engineer as part of the creation of a new Foundation’s “Metal” squad. This squad will be responsible for building and provisioning resources including storage, networking and Kubernetes clusters in new environments, starting with our own physical hardware. They will provide a Kubernetes control plane and underlying resources, outside public Cloud, to the rest of Engineering.
What Makes You a Great Fit:
You enjoy working with engineers, as well as with the management structures that are there to support you and enable you and your team to do your very best.
You are comfortable working in a remote-first company; communication is key. For us, working together means being collaborative, friendly, kind, and respectful. We operate by consensus, you can contribute to a discussion but then commit to the team decision.
As such, being such a highly distributed company, means we would love someone who is keen on working with distributed systems, too.
You are eager to learn and grow. There is a lot of room for growth and development, and the team has quite a lot of knowledge to share for those who are wanting to learn.
You approach development holistically. The team owns the full life cycle of our code; from writing design docs, to looking at developer feedback, and integration testing. We appreciate engineers who enjoy looking at the big picture, and also notice the details of the brush strokes. The Platform team mainly works with Go, Python, and Shell.
You have experience with operating your code. Since a lot of operators and developers use our software, having some grounding in both of these spaces really helps us with building better platforms for our users.
Datacenter experience.
Participation in an on-call rotation.
What You’ll Be Doing:
- Management of our physical “Metal” environment from bare metal to kubernetes.
- Management of cluster networking components: load balancing, NAT, DNS, CNIs, cross-cluster communication, Routing protocols, Architecture.
- Management of scheduling and autoscaling.
- Maintaining Crossplane compositions and Terraform modules for CSP resources common to our users. As well the management of versioning and compatibility for Crossplane and Terraform core as well as providers.
- Work with our users (Grafana Cloud application teams) to help understand their needs and ensure we’re investing in the right capabilities.
Bonus Points For:
- Experience with Cluster-API, Tinkerbell, Talos and/or Ceph
- You’ve worked in or on open source, or other community-based projects previously. At Grafana Labs, “OSS is in our DNA”.
- Experience with a few CSPs. We run Grafana Cloud on AWS, GCP, and Azure using each’s managed Kubernetes service - EKS, GKE, AKS.
- Experience operating and managing workloads on Kubernetes. We use Tanka for configuration management with Jsonnet.
- Familiarity with Kubernetes scheduling and projects like Karpenter.
- Terraform and/or Crossplane experience. We have mixed usage - each has its strengths.
- Enjoys programming in Go! We love building our own tools, utilities, exporters, etc. that suit our needs and otherwise don’t exist (and open sourcing them).
- Maybe there’s something missing in this job description that you think you might bring to our table? Feel free to let us know!
In UK, the compensation range for this role is £91,755 - £110,106. Actual compensation may vary based on level, experience, and skillset as assessed throughout the interview process. All of our roles include Restricted Stock Units (RSUs), giving every team member ownership in Grafana Labs' success. We believe in shared outcomes—RSUs help us stay aligned and invested as we scale globally.
*Compensation ranges are country specific. If you are applying for this role from a different location than listed above, your recruiter will discuss your specific market’s defined pay range & benefits at the beginning of the process.
Why You’ll Thrive at Grafana Labs:
- 100% Remote, Global Culture - As a remote-only company, we bring together talent from around the world, united by a culture of collaboration and shared purpose.
- Scaling Organization – Tackle meaningful work in a high-growth, ever-evolving environment.
- Transparent Communication – Expect open decision-making and regular company-wide updates.
- Innovation-Driven – Autonomy and support to ship great work and try new things.
- Open Source Roots – Built on community-driven values that shape how we work.
- Empowered Teams – High trust, low ego culture that values outcomes over optics.
- Career Growth Pathways – Defined opportunities to grow and develop your career.
- Approachable Leadership – Transparent execs who are involved, visible, and human.
- Passionate People – Join a team of smart, supportive folks who care deeply about what they do.
- In-Person onboarding - We want you to thrive from day 1 with your fellow new ‘Grafanistas’ to learn all about what we do and how we do it.
- Balance is Key - We operate a global annual leave policy of 30 days per annum. 3 days of your annual leave entitlement are reserved for Grafana Shutdown Days to allow the team to really disconnect. *We will comply with local legislation where applicable.
Equal Opportunity Employer: Grafana Labs is an equal opportunities employer. We welcome applications from everyone regardless of race, colour, nationality, origin, caste, sex, gender reassignment identity or expression, sexual orientation, age, religion or belief, disability, veteran status, genetic information, pregnancy, maternity, marital, family or carer status, or any other characteristic which is protected by local law. We believe that equality and diversity build a strong organisation, and we work hard to ensure that is the foundation of our organisation as we grow.
Grafana Labs may utilize AI tools in its recruitment process to assist in matching information provided in CVs to job postings. The recruitment team will continue to review inbound CVs manually to identify alignment with current openings.
#LI-Remote