Sobre esta vaga de Database Reliability Engineer - DBRE na Cognite - AI for Industry
What Cognite is: Relentless to achieve
Cognite operates at the forefront of industrial digitalization, building AI, and data solutions that solve the world’s hardest, highest-impact problems. With unmatched industrial heritage and a comprehensive suite of AI capabilities, including low-code AI agents, Cognite accelerates the digital transformation to drive operational improvements.
We thrive in challenges. We challenge assumptions. We execute with speed and ownership. If you view obstacles as signals to step forward - not backwards - you’ll feel right at home here.
Our Moonshot is bold: Unlock $100B in customer value by 2035, and redefine how global industry works. Join us in this venture where AI and data meet ingenuity, and together, we will forge the path to a smarter, more connected industrial future.
About the Role
We are looking for a Senior Database Reliability Engineer (DBRE) to join our Cloud deployment team and take ownership of the reliability, scalability, automation, and operational excellence of our core database infrastructure.
You will work across PostgreSQL, Elasticsearch, and Kafka, supporting a multi-cloud, Kubernetes-based platform at significant scale. This role is ideal for someone who enjoys solving complex infrastructure and database challenges, eliminating operational toil, and building highly automated systems.
You will partner closely with Software Engineering, SRE, Platform Engineering, and Product teams to ensure our data platforms are highly available, scalable, secure, and resilient.
What You’ll Own
Managed PostgreSQL Fleet Orchestration
- Standardize and automate the lifecycle management of a fleet of 1000+ PostgreSQL instances across Azure, AWS, and GCP.
- Build and maintain infrastructure using Infrastructure as Code (IaC) and automation frameworks.
- Automate provisioning, configuration, patching, upgrades, backups, and operational workflows.
- Improve consistency and reliability across our PostgreSQL fleet while reducing manual operational effort.
- Work with cloud-managed PostgreSQL services including Azure Database for PostgreSQL – Flexible Server, Amazon RDS, and GCP Cloud SQL.
- Evolve our provisioning and configuration approach using Kubernetes-based and Terraform-driven workflows.
Hybrid Elasticsearch Operations
- Design, operate, and scale high-performance Elasticsearch clusters across Elastic Cloud (SaaS) and ECK (self-managed Kubernetes) environments.
- Own the reliability and performance of Elasticsearch infrastructure supporting search, semantic retrieval, vector workloads, and high-volume operational data.
- Establish best practices around cluster sizing, capacity planning, shard management, upgrades, backups, monitoring, and disaster recovery.
- Support Elasticsearch deployments where PostgreSQL is the source of truth as well as environments where Elasticsearch is the authoritative datastore.
- Build automation to simplify provisioning and lifecycle management across Elastic Cloud and ECK.
- Help improve security, connectivity, zone awareness, and operational resilience across multi-cloud deployments.
Kafka Streaming Infrastructure
- Own the reliability, scalability, and performance of Kafka clusters supporting high-throughput, low-latency event streaming across the platform.
- Operate Kafka in both self-managed (Kubernetes-based, e.g. Strimzi/Kafka Operator) and managed (e.g. Confluent Cloud, MSK) configurations across multi-cloud environments.
- Drive capacity planning, partition and topic design, replication strategy, and broker-level performance tuning.
- Establish and automate best practices around upgrades, rolling restarts, backups/disaster recovery, and schema evolution (e.g. Schema Registry).
- Build monitoring and alerting for consumer lag, broker health, ISR status, and throughput to enable early incident detection.
- Partner with Software Engineering to ensure producer/consumer patterns, retention policies, and topic ownership scale cleanly as usage grows.
What You Bring
- 6+ years of experience in Database Reliability Engineering, Database Engineering, SRE, Platform Engineering, or a closely related role.
- Strong hands-on experience operating PostgreSQL at scale, preferably in cloud-managed environments.
- Strong experience with Elasticsearch, including cluster administration, performance tuning, scaling, shard management, and troubleshooting.
- Experience operating databases and stateful workloads on Kubernetes.
- Experience with Kafka or another distributed database/storage system is highly valuable.
- Strong Infrastructure as Code experience with Terraform or similar tools.
- Experience with at least one major cloud platform — Azure, AWS, or GCP — with exposure to multi-cloud environments preferred.
- Strong understanding of database reliability principles including high availability, disaster recovery, backups, replication, capacity planning, and observability.
- Proficiency in scripting or programming using Python, Go, or a similar language.
- Experience automating repetitive operational tasks and building self-service infrastructure.
- Strong troubleshooting and incident-management skills, with the ability to diagnose complex distributed-system failures.
Good to Have
- Experience with Kafka and fdb-kubernetes-operator.
- Experience with Elastic Cloud and ECK.
Proficiency in Go for building automation, tooling, or operators; Python or another scripting language is a plus.
- Experience managing PostgreSQL through Kubernetes operators or custom resources.
- Experience with Kubernetes operators and building automation for stateful workloads.
- Experience operating databases across multi-cloud and private-cloud environments.
- Experience with observability platforms and database performance monitoring.
- Familiarity with distributed systems, data replication, consistency, and failure-domain design.
What Success Looks Like
- Database provisioning and lifecycle operations become highly automated and repeatable.
- Operational toil across PostgreSQL, Elasticsearch, and Kafka is significantly reduced.
- Database platforms can scale reliably with business and customer growth.
- Incidents are detected early and resolved through strong automation and observability.
- Database infrastructure is consistent and resilient across Azure, AWS, GCP, and private-cloud environments.
- Engineering teams can consume reliable database infrastructure without requiring significant manual DBRE intervention.
Why This Role?
This is an opportunity to work on large-scale, distributed data infrastructure at the heart of a modern Knowledge Graph platform. You will have significant ownership across three critical database technologies and will play a key role in making our data infrastructure more automated, reliable, and scalable.
- Impact 2025
- Cognite's Industrial AI: Moonshot
- We’re globally recognized domain experts with an international presence that spans Phoenix, Houston, Oslo Tokyo, Bengaluru, and Abu Dhabi.