Jobs Companies Redwood Software Site Reliability Engineer

Sobre este puesto de Site Reliability Engineer en Redwood Software

Redwood Software · Presencial · Canada

OUR MISSION

At Redwood, we empower our customers with lights-out automation for their mission-critical business processes.

 

ABOUT US

Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next.

 

CORE VALUES

  • One Team. One Redwood
  • Make Your Own Weather
  • Obsess over Customer Success
  • Work the Problem
  • Be Curious
  • Own the Outcome
  • Respect Each Other

 

YOUR IMPACT

As a Site Reliability Engineer (SRE) / DevOps Engineer you will be responsible for ensuring the stability, performance, scalability, and reliability of Redwood’s mission-critical SaaS platform. You will apply engineering principles to operational challenges, automate repetitive work, strengthen observability, and collaborate across teams to deliver a resilient and high-performing customer experience.

  • Provide day-to-day management of system alerts, monitor system health, and escalate issues as necessary to maintain high availability.
  • Participate in a 24x7, team-shared on-call rotation for critical SaaS platform incidents and provide support during emergencies.
  • Lead incident response efforts to ensure fast and effective mitigation and resolution of production issues.
  • Perform thorough Root Cause Analysis (RCA) and lead blameless post-mortems to identify systemic weaknesses and establish corrective actions that prevent recurrence.
  • Collaborate with engineering teams to establish and enforce error budgets derived from Service Level Objectives (SLOs), balancing development velocity with system stability.
  • Automate routine operational tasks to reduce manual effort and toil while increasing team efficiency.
  • Design, deploy, and maintain cloud infrastructure using Infrastructure as Code (IaC), leveraging Terraform and Helm for deployment to EKS/Kubernetes clusters.
  • Design, secure, and troubleshoot AWS cloud network architecture, including VPCs, subnetting, routing, security groups/NACLs, load balancers (ALB/NLB), and VPN/Transit Gateway connectivity across a multi-region, multi-account environment supporting EKS and Docker Swarm on EC2.
  • Improve infrastructure health by developing and implementing checks, scripts, and automated remediation to proactively address known issues and enable platform self-healing.
  • Maintain, develop, and evolve Continuous Integration/Continuous Delivery (CI/CD) deployment code and pipelines.
  • Maintain existing infrastructure running on Docker and Docker Swarm while contributing to migration strategies toward EKS/Kubernetes.
  • Implement and integrate new technologies and services into Redwood’s cloud infrastructure to enhance platform capabilities and resilience.
  • Design and implement comprehensive observability strategies across metrics, logs, and traces.
  • Create and refine robust monitoring and alerting configurations within the EKS/Kubernetes ecosystem.
  • Utilize and maintain Datadog to gather performance data, create synthetic tests, configure monitoring, and visualize system health through dashboards.
  • Leverage existing monitoring solutions, including Grafana and Prometheus, while supporting the migration or integration of data into a unified observability platform.
  • Document issues, remediation steps, system architecture, and runbooks to support knowledge transfer and rapid incident response.
  • Collaborate closely with Support, Customer Success, Migration, and Professional Services teams to deliver a high level of SaaS service and minimize customer impact during changes.
  • Maintain a strong customer focus when planning deployments and updates, considering the impact on end users before implementing changes.

 

YOUR EXPERIENCE

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, ideally supporting a production SaaS platform.
  • Hands-on AWS Cloud Engineering experience with strong working knowledge of the AWS ecosystem, including AWS IAM roles and policies.
  • Strong knowledge of AWS cloud networking, including VPC design and peering, subnetting, security groups/NACLs, route tables, ALB/NLB load balancers, Transit Gateway/VPN connectivity, and Route 53 DNS.
  • Ability to design and troubleshoot connectivity across multi-region, multi-account AWS environments.
  • Proficiency with EKS/Kubernetes (K8s) and container orchestration technologies.
  • Demonstrated experience with Infrastructure as Code (IaC), specifically Terraform and Helm.
  • Working experience with Docker and maintaining systems using Docker Swarm.
  • Expertise in implementing and managing logging, monitoring, and observability solutions.
  • Direct experience with Datadog, including APM, infrastructure monitoring, synthetic testing, and custom dashboards.
  • Experience with Grafana and/or Prometheus is a plus.
  • Proficiency working in Linux environments with strong Bash and/or Python scripting skills for automation and troubleshooting.
  • Operational experience with relational databases in production environments, such as AWS RDS/PostgreSQL, including performance tuning, backup and restore, and troubleshooting under load.
  • Experience providing product or application support for high-availability SaaS products.
  • Experience designing, implementing, and operating within a DevSecOps environment.
  • Excellent written and verbal communication skills, with the ability to clearly explain complex technical issues and Root Cause Analyses to both technical and customer-facing audiences.

 

Pay Transparency: The typical base salary range for this position is $125,000 - $145,000 CAD annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable. Final compensation will be determined based on the candidate's skills, experience, and relevant qualifications.

Vacancy Status: Existing vacancy

AI Transparency Disclosure: As part of our commitment to hiring transparency and in compliance with Ontario law, please be advised that we utilize artificial intelligence (AI) tools during our interview process. Specifically, we may use an AI-enabled notetaking assistant to help capture and summarize interview discussions. This tool is used solely to assist our hiring team in summarizing candidate information; all final hiring and selection decisions are made exclusively by individuals on the hiring team.

 

If you like growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

 

THE LEGAL BIT
Redwood is an equal opportunity employer. Redwood prohibits unlawful discrimination based on race, colour, religion, sex, gender identity, marital or veteran status, age, national origin, ancestry, citizenship, physical or mental disability, medical condition, genetic information or characteristics (or those of a family member), sexual orientation, pregnancy or any other consideration made unlawful by regional or local laws. We also prohibit discrimination based on a perception that anyone has any of those characteristics or is associated with a person who has or is perceived as having any of those characteristics. All such discrimination is unlawful and will have a zero tolerance policy applied to it.
 

Redwood will comply with all local data protection laws, including GDPR when it comes to the handling and processing of personal data. Should you wish for us to remove your personal data from our recruitment database, please email us directly at [email protected]

¿Listo para postularte en Redwood Software?
Postúlate en Redwood Software

Cómo se compara este salario de SRE

Este puesto paga $96,983/yren línea con el rango típico para los puestos de SRE.

$78,558 la mediana de $123,407 $223,620

Rango típico $89,896–$170,113/yr, a partir de 57 ofertas comparables de SRE en JobsRadar (salario anualizado en USD). Ver datos salariales de SRE →

Sobre Redwood Software

About Redwood Software

Redwood Software delivers IT, finance and business process automation to help modern enterprises excel in the digital age. Redwood orchestrates and automates business processes across complex hybrid IT environments so enterprise organizations can focus on business agility, cost efficiency, and customer experiences. Our automation solutions help thousands of organizations across 150 countries execute with speed and precision.
 
We pride ourselves on having an inclusive, supportive, coaching culture, giving you the opportunity to build lasting relationships with all our team members and senior management.   

 

Important: We have been made aware that individuals are posing as Redwood recruiters in an attempt to deceive candidates into sharing personal information.  Redwood employees will only contact you from an “@redwood.com” email domain.  If you have questions or suspect an email is fraudulent, please contact us at [email protected].

 

 

Ver todos los empleos en Redwood Software →

Empleos similares

Baseten
Site Reliability Engineer
Baseten
⚡ Postúlate pronto San Francisco Híbrido $165,000–$330,000
● Nuevo 👁 Visto ✓ Postulado hace 7h
MongoDB
Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
MongoDB
⚡ Postúlate pronto Montreal; Toronto Presencial CA$144,000–CA$200,000
● Nuevo 👁 Visto ✓ Postulado hace 8h
MongoDB
Site Reliability Engineer (Senior or Staff), Deployments
MongoDB
⚡ Postúlate pronto Toronto Presencial CA$144,000–CA$200,000
● Nuevo 👁 Visto ✓ Postulado hace 8h
Credit Genie
Senior DevOps/SRE Engineer
Credit Genie
⚡ Postúlate pronto Plymouth Meeting, PA Presencial $170,000–$225,000
● Nuevo 👁 Visto ✓ Postulado hace 23h
CA
Senior Site Reliability and DevOps Engineer
Capco
⚡ Postúlate pronto Canada - Toronto Presencial CA$118,000–CA$152,000
● Nuevo 👁 Visto ✓ Postulado hace 1d
CA
Site Reliability and DevOps Engineer
Capco
⚡ Postúlate pronto Canada - Toronto Presencial CA$92,000–CA$118,000
● Nuevo 👁 Visto ✓ Postulado hace 1d
Braze
Senior Site Reliability Engineer I
Braze
⚡ Postúlate pronto Vancouver Presencial CA$172,000–CA$308,400
● Nuevo 👁 Visto ✓ Postulado hace 1d
Braze
Senior Site Reliability Engineer I
Braze
⚡ Postúlate pronto Toronto Presencial CA$172,000–CA$308,400
● Nuevo 👁 Visto ✓ Postulado hace 1d
Finning
Reliability Engineer
Finning
⚡ Postúlate pronto Edmonton, AB, CA Presencial
● Nuevo 👁 Visto ✓ Postulado hace 2d

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Redwood Software

Ver todos los empleos en Redwood Software →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis