Jobs Companies NextGen Healthcare Sr. Cloud Operations Reliability Engineer (SRE)

Sobre este puesto de Sr. Cloud Operations Reliability Engineer (SRE) en NextGen Healthcare

NextGen Healthcare · Remoto · Remote GA

Job Description:

The Senior Cloud Operations Reliability Engineer is responsible for driving operational excellence and strengthening the reliability posture of cloud-based services and supported platforms. This role owns critical reliability initiatives, establishes observability and service health practices, and leads incident response coordination to improve service availability, resiliency, and recovery. The Senior Cloud Operations Reliability Engineer partners with engineering, security, and operations teams to advance reliability practices, support and mature service-level objectives, improve production readiness, and develop reliability-focused automation that reduces operational toil and accelerates incident recovery.

  • Own service reliability and operational health—establish and maintain SLOs/SLIs, design monitoring and alerting strategies, and drive improvements that enhance service availability and performance across cloud platforms.
  • Lead incident response coordination and post-incident processes, including troubleshooting complex production issues, conducting root cause analysis, and driving remediation activities with accountability for timeline and resolution quality.
  • Design and implement reliability-focused automation, operational tooling, and runbooks to reduce manual toil, improve response consistency, and strengthen production readiness and resilience; apply Infrastructure as Code practices where appropriate to support recovery, reliability, and operational consistency.
  • Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures to ensure rapid problem detection and response.
  • Conduct performance and capacity analysis, evaluate utilization trends, identify bottlenecks; provide recommendations for reliability-focused scaling, performance improvement, capacity planning, and operational readiness of cloud-based services.
  • Partner with development and engineering teams to evaluate deployment readiness, support deployment reliability improvements, and implement operational best practices that strengthen service reliability, rollback readiness, and production supportability.
  • Contribute to disaster recovery and business continuity planning, conduct operational readiness exercises, and ensure recovery procedures and documentation reflect current production state and evolving business requirements.
  • Mentor team members and establish reliability standards and practices within Cloud Operations and supported service areas; create and maintain operational documentation, standard operating procedures, and knowledge base materials.
  • Support operational adherence to cloud governance, compliance, and security initiatives; including access control, tagging, logging, and audit readiness, and reliability-related documentation.
  • Perform other duties that support the overall objective of the position.

Education Required:

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
  • Or, any combination of education and experience which would provide the required qualifications for the position.

Experience Required:

  • 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline with demonstrated ownership of production systems.
  • Extensive hands-on experience supporting production cloud environments using Google Cloud Platform (GCP), AWS, or equivalent cloud service providers.
  • Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
  • Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.) and version control best practices.
  • Strong background in incident management and post-incident review processes; experience driving corrective actions and establishing reliability improvements.
  • Experience with Kubernetes operations, containerization, and orchestration platforms.
  • Experience with application performance monitoring (APM) and distributed tracing.
  • Experience mentoring junior engineers or leading operational improvements initiatives.

License/Certification Required:

  • Google Cloud certifications: Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer.
  • AWS certification: AWS SysOps Administrator or equivalent.
  • Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering (SRE), or ITIL.

Knowledge, Skills & Abilities:

  • Knowledge of: Working knowledge of CI/CD practices, cloud governance, compliance frameworks, disaster recovery, and business continuity planning. Deep technical knowledge of Google Cloud Platform (GCP), AWS, or similar cloud providers; understanding of cloud-native services, networking, security, and compute models.Familiarity with observability tools such as Grafana, Prometheus, Cloud Monitoring, or similar platforms. Security operations, compliance auditing, or audit readiness processes, preferred.
  • Skill in: Hands-on expertise with monitoring platforms (Datadog, New Relic, Prometheus, Cloud Monitoring, etc.); ability to design effective dashboards, alerts, and health checks. Proficiency in scripting languages (Python, Bash, Go, etc.) to develop automation solutions that reduce manual effort.
  • Ability to: Advanced ability to diagnose complex, multi-layered infrastructure issues and coordinate timely recovery. Ability to translate complex technical findings into actionable recommendations; experience influencing cross-functional teams on reliability practices.

The company has reviewed this job description to ensure that essential functions and basic duties have been included. It is intended to provide guidelines for job expectations and the employee's ability to perform the position described. It is not intended to be construed as an exhaustive list of all functions, responsibilities, skills and abilities. Additional functions and requirements may be assigned by supervisors as deemed appropriate. This document does not represent a contract of employment, and the company reserves the right to change this job description and/or assign tasks for the employee to perform, as the company may deem appropriate.

NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

¿Listo para postularte en NextGen Healthcare?
Postúlate en NextGen Healthcare

Sobre NextGen Healthcare

NextGen is the healthcare technology partner for diverse ambulatory practices with evolving business needs. We are uniquely equipped to make it easier when things get complex—combining performance at scale with operational flexibility. This empowers practices to become confident, capable and ready for whatever comes next.

Ver todos los empleos en NextGen Healthcare →

Empleos similares

PrizePicks
Senior Site Reliability Engineer (SRE)
PrizePicks
⚡ Postúlate pronto Atlanta, GA preferred, Remote · restringido por ubicación $120,000–$175,000
● Nuevo 👁 Visto ✓ Postulado hace 3d
Carrier
Staff Site Reliability Engineer
Carrier
⚡ Postúlate pronto CAFLO: Carrier-Home Florida Re... · restringido por ubicación $96,000–$192,000
● Nuevo 👁 Visto ✓ Postulado hace 1sem
Swift
Lead Site Reliability Engineer
Swift
⚡ Postúlate pronto Kuala Lumpur, Malaysia Presencial
● Nuevo 👁 Visto ✓ Postulado hace 7h
Tailor
SRE (Site Reliability Engineer)
Tailor
⚡ Postúlate pronto Tokyo · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 11h
Okta
Staff SRE for K8s Platform Team (AWS, Kubernetes, Platform Creation, Helm, Karpenter, Istio)
Okta
⚡ Postúlate pronto Bengaluru, India Presencial
● Nuevo 👁 Visto ✓ Postulado hace 12h
Okta
Senior Site Reliability Engineer
Okta
⚡ Postúlate pronto Bengaluru, India Presencial
● Nuevo 👁 Visto ✓ Postulado hace 12h
Okta
Senior Manager, Site Reliability Engineering - Infrastructure Platform
Okta
⚡ Postúlate pronto Bellevue, Washington; Chicago,... Presencial $232,000–$319,000
● Nuevo 👁 Visto ✓ Postulado hace 12h
Leidos
System Administrator
Leidos
⚡ Postúlate pronto Remote, US · restringido por ubicación $92,300–$166,850
● Nuevo 👁 Visto ✓ Postulado hace 5d
Clio
Solutions Engineer, Cloud Communications
Clio
⚡ Postúlate pronto Remote - CO, USA Híbrido $122,800–$166,000
● Nuevo 👁 Visto ✓ Postulado hace 1sem

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en NextGen Healthcare

Ver todos los empleos en NextGen Healthcare →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis