Jobs Companies OnBoard Site Reliability Engineer

Über diese Site Reliability Engineer Stelle bei OnBoard

OnBoard · Vor Ort · United States

Cloud Operations Engineer III

Function:  Engineering 

Reports to:  Manager, Cloud Operations 

Location: Remote - United States 

Position Summary 

The Cloud Operations Engineer III is a senior member of the cloud operations team, responsible for the reliability, observability, performance, and operational security of our multi-product SaaS platform. This role owns our Datadog observability practice — instrumentation standards, dashboards, SLOs, monitors, and alert routing — and leads the migration off our legacy monitoring stack. 

It is an engineering role, not a ticket-queue role: the expectation is that recurring operational work gets replaced with code. The Cloud Operations Engineer III participates in on-call, incident response and is measured on fewer customer-impacting incidents, faster detection and recovery, and less manual work year over year. 
 
The ideal candidate is a proactive problem-solver who thrives in dynamic, evolving environments and works effectively across departments to address complex challenges.  They have experience partnering with cross-functional teams to understand and document requirements, then translating those needs into meaningful dashboards that improve service visibility (Observability) and support informed decision-making.  They are passionate about automation, process improvement, and eliminating unnecessary manual effort.  They confidently propose better approaches when opportunities for improvement arise. 

Key Responsibilities 

Observability and Datadog Ownership 

  • Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration. 
  • Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner with engineering teams to close instrumentation gaps in both legacy and modern codebases. 
  • Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product. 
  • Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook. 

Automation and Toil Elimination 

  • Develop and maintain automation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting. 
  • Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free. 
  • Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning. 

Security, Documentation, and Mentorship 

  • Implement and maintain platform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning. 
  • Author and maintain runbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response. 

Skills and Experience Needed 

  • Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience. 
  • 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering for production SaaS systems. 
  • Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs. 
  • Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safely operate. 
  • Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures. 
  • Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud. 
  • Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm or Kustomize; Azure DevOps preferred). 
  • Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews. 
  • Experience operating multi-region, multi-tenant systems. 
  • Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation. 
  • Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals. 
  • Strong communication, and teamwork skills, including the ability to work effectively with legacy systems and their constraints. 
  • Nice to have: experience migrating from a legacy monitoring stack to a consolidated observability platform; relevant Azure, Kubernetes, or Datadog certifications. 

Competencies 

Accountability 

Adaptability 

AI Curiosity/Innovation 

Applied Learning 

Business Acumen 

Collaboration 

Customer Focus 

Dealing w/Ambiguity 

Decision Making 

Driving for Results 

Initiating Action 

Planning and Organizing 

Technical/Professional Knowledge 

 

 

 

About the Company:

Boards set the standard for what organizations can achieve. At OnBoard, our board management software helps boards function at a higher level so every organization can make a bigger difference in the world.

Launched in 2011, today, OnBoard serves as the board intelligence platform for more than 5,000 organizations and their 12,000 boards and committees in 60 countries worldwide. With customers in higher education, nonprofit, healthcare systems, government, and enterprise business, OnBoard is the leading board management provider.

OnBoard has grown from a class project at Purdue University in West Lafayette, Indiana in 2003 into the world’s leading board management software platform today. Backed by JMI Equity and the acquisitions of eScribe and Govenda, OnBoard is positioned to become the industry leader in Board Management and Meeting Solutions for private and public sector entities.

Benefits and Perks: 

  • Fully remote work with company provided equipment (laptop, software, etc.) 
  • Employment with a growing, casual, fun, philanthropic minded company
  • US Based Employees
    • Comprehensive, high-quality medical/prescription drug plan options, as well as dental and vision plan offerings.   
    • An employer contribution to your Health Savings Account (HSA) if you participate in a High Deductible Healthcare Plan.  
    • Medical Flexible Spending Accounts available.   
    • Dependent Care Flexible Spending Accounts available.  
    • Basic life insurance in the amount of $50,000 or 1 X’s your salary (whichever is higher) 
    • Short and long-term disability and Accidental Death and Dismemberment benefits at no cost to you.  
    • 401K Retirement Savings Plan with automatic enrollment at the first of the month following 60 days of employment at 5% to help you secure your financial freedom. We offer a generous company match that starts on the first of the month following 60 days of employment. The company match is dollar for dollar on the first 3% of your pay that you contribute and $0.50 on the dollar on the next 2%, for a total match of 4%. 
    • Paid Time Off (PTO)/Holiday 
  • CAN Based Employees
    • Employer paid Life and Accidental Death Insurance
    • Contribution to Health Care Spending Account
    • Dependent Life Insurance
    • Optional Life Insurance
    • LTD Insurance
    • Drug and Paramedical Coverage
    • Dental Insurance
    • Vision Insurance
    • EAP
  • AUS Based employees
    • Superannuation rate of 12% 
    • Monthly stipend of $400 AUD to purchase private medical insurance 
  • UK Based Employees (via EPG)
    • Pension - Aegon
      • Passageways/OnBoard contributes 8% of the employee's basic salary
      • Employees can contribute up to 100% of salary subject to max limits
      • Enrolled from Day 1 of employment
    • Private Medical Insurance
    • Life Assurance
    • Income Protection
    • Critical Illness
    • Employee Assistance Programme
    • Serious Illness Benefit
    • Help@Hand
    • Cashplan

Diversity Statement - Culture of Togetherness:  

At OnBoard, our mission is to encourage and celebrate a culture of togetherness. We acknowledge that uniqueness is powerful, and we welcome, foster, and appreciate all. Diversity, Equity, and Inclusiveness fuel the Pathfinder atmosphere and all our efforts. Our power is in our people and we Pledge 1% to give back to our communities and across the globe.

OnBoard is an equal opportunity employer and committed to a diverse and inclusive working environment. We  do not discriminate based on race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status.

Interview Transparency & Technology Disclosure
We use video/audio recordings and artificial intelligence (AI) tools during our interview process to transcribe responses, evaluate skills, and streamline evaluations. Your data is processed securely and handled in line with our Privacy Policy and local data protection laws

Bereit, sich bei OnBoard zu bewerben?
Bei OnBoard bewerben

Wie sich dieses Gehalt für SRE vergleicht

Diese Stelle zahlt $50,000/yrunter der üblichen Spanne für SRE Stellen.

$124,080 dem Median $186,420 $253,766

Übliche Spanne $155,275–$213,628/yr, aus 286 vergleichbaren SRE Anzeigen auf JobsRadar (Vergütung auf USD hochgerechnet). Gehaltseinblicke für SRE ansehen →

Über OnBoard

 

 

Alle Jobs bei OnBoard ansehen →

Ähnliche Jobs

Apex Technology, Inc.
Site Reliability Engineer
Apex Technology, Inc.
⚡ Früh bewerben Los Angeles Vor Ort $155,000–$195,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Waabi
Vehicle Reliability Engineer
Waabi
⚡ Früh bewerben Dallas, TX Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Roblox
Senior Site Reliability Engineer, Compute
Roblox
⚡ Früh bewerben San Mateo, CA, United States Vor Ort $243,290–$295,250
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Varda Space Industries
Spacecraft Build Reliability Engineer II
Varda Space Industries
⚡ Früh bewerben El Segundo, California, United... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Varda Space Industries
Senior Spacecraft Build Reliability Engineer
Varda Space Industries
⚡ Früh bewerben El Segundo, California, United... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Peak Energy
Senior Cell Validation & Reliability Engineer
Peak Energy
⚡ Früh bewerben Burlingame, California Vor Ort $150,000–$190,000
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Anduril Industries
Director, Site Reliability Engineering
Anduril Industries
⚡ Früh bewerben Costa Mesa, California, United... Vor Ort $253,000–$336,000
● Neu 👁 Gesehen ✓ Beworben vor 7 Std.
Anduril Industries
Senior Site Reliability Engineer - Undersea Dominance
Anduril Industries
⚡ Früh bewerben Costa Mesa, California, United... Vor Ort $166,000–$250,000
● Neu 👁 Gesehen ✓ Beworben vor 7 Std.
Anduril Industries
Staff Site Reliability Engineer
Anduril Industries
⚡ Früh bewerben Costa Mesa, California, United... Vor Ort $191,000–$253,000
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei OnBoard

Alle Jobs bei OnBoard ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos