Jobs › Companies › Clearwater Analytics › Sr. Site Reliability Engineer

About this Sr. Site Reliability Engineer role at Clearwater Analytics

Clearwater Analytics · Onsite · Office - Boise

Job Summary

We are seeking a highly skilled Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems and applications. This role drives automation, monitoring, and incident management practices while operating Kubernetes platforms (Amazon EKS) and leveraging observability tools such as Prometheus, Grafana, Dynatrace, and OpenSearch to maintain high availability and operational excellence.


Key Responsibilities

• Design, build, and maintain highly available, scalable, and reliable production systems.

• Define and manage SLIs, SLOs, and SLAs to drive system reliability.

• Automate infrastructure provisioning and operations using Terraform (IaC).

• Operate and manage cloud-native platforms, including Amazon EKS.

• Implement and maintain monitoring, logging, and alerting using Prometheus, Grafana, Dynatrace, and OpenSearch.

• Lead incident management — on-call rotation, production troubleshooting, and root cause analysis (RCA).

• Drive AI-assisted investigations as a core part of incident response, and build and maintain the prompts, integrations, and guardrails that make AI-driven triage and RCA reliable.

• Improve system reliability through automation, self-healing mechanisms, and performance tuning.

• Collaborate with development teams to improve application reliability, scalability, and deployment processes.

• Build and maintain CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions) for fast, reliable software delivery.

• Perform capacity planning and cost optimization for infrastructure and services.

• Ensure security, compliance, and best practices across infrastructure and applications.


Required Qualifications

• Bachelor’s degree in computer science or a related field, or equivalent practical experience.

• 7+ years in Site Reliability Engineering or Platform Engineering.

• Proven ownership of incident management, on-call support, and root cause analysis (RCA) for production systems.

• Strong expertise in Terraform and Infrastructure as Code.

• Hands-on experience with AWS and EKS.

• Strong understanding of monitoring, logging, and observability (Prometheus, Grafana, Dynatrace, OpenSearch).

• Proficiency in Python, Java, Go, or Bash.

• Experience with Agile development and CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions).

• Strong problem-solving, documentation, and communication skills.

• Proven ability to troubleshoot effectively in high-pressure production environments.

• Experience with autoscaling, performance tuning, and cost optimization.

• Familiarity with AI-assisted automation tools and a track record of using them to reduce toil and improve reliability.


Preferred Skills

• Docker and Linux administration.

• Build systems and dependency management (Maven, Gradle, npm).

• Additional AWS services: Cognito, WAF, Elasticsearch, SNS, SQS, S3, Systems Manager.

• Database infrastructure knowledge (RDS, MySQL, SQL Server).

• Cloud or Kubernetes certifications.


Ready to apply to Clearwater Analytics?
Apply to Clearwater Analytics

About Clearwater Analytics

Clearwater Analytics (NYSE: CWAN) is transforming investment management with the industry’s most comprehensive cloud-native platform for institutional investors across global public and private markets. While legacy systems create risk, inefficiency, and data fragmentation, Clearwater’s single-instance, multi-tenant architecture delivers real-time data and AI-driven insights throughout the investment lifecycle. The platform eliminates information silos by integrating portfolio management, trading, investment accounting, reconciliation, regulatory reporting, performance, compliance, and risk analytics in one unified system. Serving leading insurers, asset managers, hedge funds, banks, corpora

See all jobs at Clearwater Analytics →

Similar jobs

AX
Site Reliability Engineer II
Axon
⚡ Apply early Washington, United States Hybrid $135,000–$154,000
● New 👁 Seen ✓ Applied 23h ago
FT
Reliability Engineer
Flow Traders
⚡ Apply early Hong Kong Onsite
● New 👁 Seen ✓ Applied 1d ago
CyberCube
Site Reliability Engineer
CyberCube
⚡ Apply early Tallinn Office Hybrid €42,000–€54,000
● New 👁 Seen ✓ Applied 2d ago
Zscaler
Staff Site Reliability Engineer (Production Engineer)- Federal
Zscaler
⚡ Apply early Bellevue, Washington, USA; Bos... Hybrid $119,000–$170,000
● New 👁 Seen ✓ Applied 3d ago
Planet
Senior Site Reliability Engineer
Planet
⚡ Apply early Berlin, Germany; Haarlem, Neth... Hybrid €77,000–€96,300
● New 👁 Seen ✓ Applied 4d ago
SC
Site Reliability Engineer
SS&C Technologies
⚡ Apply early Boston MA - One Post Office Sq... Hybrid
● New 👁 Seen ✓ Applied 4d ago
Zscaler
Staff Site Reliability Engineer
Zscaler
⚡ Apply early San Jose, California, USA Hybrid $122,500–$175,000
● New 👁 Seen ✓ Applied 6d ago
Zscaler
Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)
Zscaler
⚡ Apply early Bangalore, IND Onsite
● New 👁 Seen ✓ Applied 6d ago
Nabla
SRE / Backend Engineer
Nabla
⚡ Apply early New York office Hybrid $160,000–$220,000
● New 👁 Seen ✓ Applied 6d ago

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at Clearwater Analytics

See all jobs at Clearwater Analytics →

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free