Jobs Companies GE Vernova SRE Observability SLO Engineer

À propos de ce poste SRE Observability SLO Engineer chez GE Vernova

GE Vernova · Sur site · Mexico City

Job Description Summary

GE Vernova's GridOS Platform Engineering team is building the next generation of SaaS reliability for critical energy infrastructure.The Observability & SLO Engineer is the eyes and ears of the GridOS SRE team. In this role you will build and own the full telemetry stack — from instrumentation standards to SLO dashboards to synthetic monitors — that give GE Vernova and its utility customers real-time confidence in the reliability of mission-critical energy management systems. This is a cyclical, high-impact position: you will drive an intensive initial ramp to establish v1.0 observability coverage across all customer environments, then shift into an ongoing improvement cadence aligned to new product releases and customer onboarding.

Job Description

Roles and Responsibilities

Telemetry Standards & Architecture

  • Implement organization-wide telemetry standards covering metrics, logs, and distributed traces across all GridOS SaaS services.

  • Implement metrics collection for Kubernetes-hosted services (EKS/Rancher) including pod-level, namespace-level, and cluster-level metrics.

  • Working with the SRE Lead and SRE Platform Engineers help define and implement data retention policies, cardinality budgets, and telemetry cost controls to keep observability economically sustainable.

  • Publish and maintain an Observability Runbook library covering onboarding, alert tuning, and dashboard standards for Platform SRE and Production DevOps teams.

SLO Definition, Tooling & Governance

  • Partner with product engineering, Platform SRE, and customer stakeholders to define meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) per product and customer tier.

  • Build and maintain SLO tooling — error budget burn-rate alerts, burn-rate dashboards, and automated SLO compliance reports.

  • Govern the SLO review cycle: facilitate monthly SLO reviews, identify reliability risks early, and drive prioritization of reliability work with the SRE Lead.

  • Translate SLOs into SLAs for customer-facing commitments in coordination with the SRE Team Lead.

Dashboards & Alerting

  • Design and build operational dashboards covering availability, latency, error rates, and saturation (the 'Golden Signals') for every GridOS SaaS product.

  • Implement alert policies with noise-reduction practices: symptom-based alerting, multi-window burn-rate rules, and alert deduplication.

  • Create executive-level dashboards for SRE leadership and customer-facing uptime/availability reports aligned to contractual SLAs.

  • Establish and maintain alert routing, escalation policies, and on-call schedules in coordination with the incident response workflow.

Synthetic Monitoring

  • Design and implement a synthetic monitoring plan covering critical user journeys for each GridOS SaaS product and customer environment.

  • Build synthetic checks for API health, UI flows, and integration endpoints using AWS CloudWatch Synthetics or equivalent tooling.

  • Define alerting thresholds for synthetic monitors and integrate them into the broader incident detection pipeline.

Continuous Improvement Cadence

  • After v1.0 delivery, transition into a roadmap-aligned improvement cycle: expand coverage for new features, tune alert signal-to-noise, and retire stale monitors.

  • Conduct periodic observability health reviews to identify gaps in coverage, reduce MTTD (Mean Time to Detect), and improve MTTR (Mean Time to Resolve).

  • Collaborate with the Production DevOps engineer on FinOps validation — correlate infrastructure cost metrics with performance and reliability data.

Required Experience

  • 2–3 years in SRE, observability engineering, or infrastructure reliability roles.}

  • Fluent in English.

  • Experience with at least one major observability platform — Datadog, Grafana + Prometheus, AWS CloudWatch, Dynatrace, or New Relic.

  • Decent understanding of distributed systems telemetry: metrics (Prometheus/CloudWatch), structured logging (CloudWatch

  • Logs Insights, ELK), and distributed tracing (OpenTelemetry, AWS X-Ray).

  • Experience with Kubernetes observability — kube-state-metrics, node exporters, Helm deployed monitoring stacks, and namespace-level resource metrics.

  • Proficiency in at least one query/visualization language: PromQL, Splunk SPL, Datadog Query Language, or CloudWatch Logs Insights query syntax.

  • Experience enabling monitoring alerts to provide visibility to system health.

  • Scripting skills in Python and/or Bash for automation of monitoring configuration and report generation.

Key Skills and Technologies

  • Cloud Technologies - AWS Cloud Infrastructure - EKS, RDS, MSK, S3, EC2, EBS, SQS, etc. Kubernetes - EKS, Rancher

  • Deployment and Configuration Tools - Ansible, Chef or Puppet

  • Observability tools and technology - Datadog, Splunk, NewRelic, etc.

  • Alerting and notification - AWS and Azure alerting notification

  • Scripting - Go, Python, Groovy, Bash

  • Linux Administration Skills

Nice to Have

  • Familiarity with OpenTelemetry (OTel) for vendor-agnostic instrumentation.

  • Experience with synthetic monitoring tools — AWS CloudWatch Synthetics, Datadog Synthetics, or Catchpoint.

  • Experience in regulated industries — energy, utilities, healthcare — where compliance grade audit trails are required.

  • AWS certifications: CloudWatch / Observability specialty, Solutions Architect Associate or Professional.

Education Qualification

Bachelor's Degree in Computer Science


Personal Attributes:

• Critical thinker; able to quickly adapt to changing environments
• A hacker or tinkerer at heart
• Risk taker, not afraid to think outside the box or challenge the status quo
• Emotional Intelligence, ability to influence up and out and the ability to work independently
• Must be a team player with a strong desire to win
• Passionate about continuously learning
• Highly organized and efficient; able to balance competing priorities and execute accordingly
• Strong oral and written communication skills.

Additional Information

Relocation Assistance Provided: Yes

Prêt à postuler chez GE Vernova ?
Postuler chez GE Vernova

À propos de GE Vernova

Addressing the climate crisis is an urgent global priority and we take our responsibility seriously. That is our singular mission at GE Vernova: continuing to electrify the world while simultaneously working to help decarbonize it. If we want our energy future to be different…we must be different. Our mission is embedded in our name. We retain our treasured legacy, “GE,” in our name as an enduring and hard-earned badge of quality and ingenuity. “Ver” / “verde” signal Earth’s verdant and lush ecosystems. “Nova,” from the Latin “novus,” nods to a new, innovative era of lower carbon energy that GE Vernova will help deliver. Together, we have The Energy to Change the World. www.gevernova.com

Voir tous les emplois chez GE Vernova →

Emplois similaires

Bluelightconsulting
DevOps / SRE / DevSecOps Engineer (AWS) - Remote, Latin America
Bluelightconsulting
⚡ Postuler tôt Puebla City, Mexico · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
Bluelightconsulting
DevOps / SRE / DevSecOps Engineer (AWS) - Remote, Latin America
Bluelightconsulting
⚡ Postuler tôt Mexico City, Mexico · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
Bluelightconsulting
DevOps / SRE / DevSecOps Engineer (AWS) - Remote, Latin America
Bluelightconsulting
⚡ Postuler tôt Chihuahua City, Mexico · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
GV
SRE Platform Engineer
GE Vernova
⚡ Postuler tôt Mexico City Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 5 j
Mastercard
Lead Site Reliability Engineer
Mastercard
⚡ Postuler tôt Mexico City, Mexico Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 5 j
Mastercard
Senior Site Reliability Engineer
Mastercard
⚡ Postuler tôt Mexico City, Mexico Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
Crunchyroll, LLC
Senior Engineer, Insights & Reliability Engineering
Crunchyroll, LLC
⚡ Postuler tôt Mexico City, Mexico City, Mexi... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 sem.
NCR Atleos
Site Reliability Engineer I
NCR Atleos
⚡ Postuler tôt MEXICO CITY, MEX Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 sem.
Mastercard
Site Reliability Engineer
Mastercard
⚡ Postuler tôt Mexico City, Mexico Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 sem.

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez GE Vernova

Voir tous les emplois chez GE Vernova →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit