Jobs Companies Guidewire Lead, Site Reliability Engineer (Platform)

Über diese Lead, Site Reliability Engineer (Platform) Stelle bei Guidewire

Guidewire · Vor Ort · Malaysia - Kuala Lumpur

Summary

The Opportunity

Site Reliability Engineering (SRE) brings together software and systems engineering to design and operate large-scale, highly distributed, and fault-tolerant systems. As a Lead Site Reliability Engineer, on the Platform team, you will play a key role in designing, building, and operating the foundational infrastructure that powers Guidewire's SaaS platform. You will focus on automation, reliability, scalability, and operability — ensuring our multi-tenant cloud platform consistently meets both functional and non-functional requirements. You will work closely with product development teams to influence system design, improve resilience, and ensure production readiness, while building the tools and frameworks that enable efficient, global, follow-the-sun operations.

Job Description

What You'll Do

Drive Reliability, Automation & Scale

  • Design, build, and operate highly reliable, scalable infrastructure for a multi-tenant SaaS platform.

  • Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.

  • Develop internal tools, services, and frameworks to improve efficiency and reduce manual effort.

  • Participate in a follow-the-sun rotation to support critical production systems.

Improve Platform & Infrastructure

  • Contribute to core platform systems by building features, resolving issues, and enhancing reliability.

  • Partner with development teams to ensure systems meet availability, performance, and scalability requirements.

  • Proactively identify risks, bottlenecks, and failure modes, and implement solutions before they impact customers.

Observability, Incident Management & Resilience

  • Build and maintain observability systems (metrics, logging, tracing, dashboards).

  • Define and track Service Level Objectives (SLOs) and reliability metrics.

  • Lead or contribute to incident response, root cause analysis, and blameless postmortems.

  • Drive improvements toward self-healing systems and reduced operational toil.

Security & Identity

  • Design and support secure access patterns, including SSO, SAML, and OAuth-based authentication systems.

  • Ensure platform services meet security and compliance standards.

Enablement & Collaboration

  • Collaborate across engineering teams, providing guidance, feedback, and hands-on contributions.

  • Create and maintain documentation, runbooks, and training materials.

  • Mentor engineers and promote best practices in reliability engineering and automation.

Who You Are

Infrastructure Development

  • Strong programming skills in Python or Go (Java/Spring Boot is a plus).

  • Deep experience with AWS and building/operating production systems at scale.

  • Hands-on expertise with Kubernetes (EKS), Docker, Helm, CNI, and Ingress networking.

  • Strong understanding of Kubernetes primitives and patterns (deployments, services, operators, etc.).

  • Experience with Infrastructure as Code (Terraform, Terragrunt, or similar).

  • Solid understanding of Linux systems and networking fundamentals.

Observability & Operations

  • Experience with observability platforms such as Datadog, Prometheus, OpenTelemetry, or CloudWatch.

  • Familiarity with incident management practices and production support in a microservices environment.

  • Experience with messaging/streaming systems (e.g., Kafka, SQS) and relational databases (e.g., Aurora, RDS) is a plus.

Security & Identity

  • Working knowledge of SSO, SAML, OAuth, and identity providers (Okta is a plus).

  • Experience with AWS IAM (roles, policies, IRSA), VPC security groups, and Kubernetes security primitives (RBAC, network policies, pod security standards, secrets management).

DevOps & Delivery

  • Experience with CI/CD and GitOps tools such as GitHub Actions, TeamCity, Jenkins, FluxCD, or Bitbucket.

  • Comfortable working in agile environments (Scrum, Kanban).

Mindset & Collaboration

  • Strong troubleshooting and problem-solving skills with a proactive, systems-thinking mindset.

  • Passion for automation: “If you have to do it more than once, automate it.”

  • Excellent communication skills and ability to work across distributed teams.

  • A collaborative team player who can influence, mentor, and lead through technical expertise.

  • Demonstrated ability to leverage AI and data-driven insights to improve productivity and outcomes.

Preferred Qualifications

  • Bachelor's degree in Computer Science or related field, or equivalent experience

  • Experience supporting large-scale SaaS platforms

  • AWS or Kubernetes certifications

  • Exposure to modern platform frameworks such as KubeVela (OAM) or Crossplane

  • Contributions to open-source projects

Why Guidewire?

  • Work on a mission-critical global platform used by leading P&C insurers worldwide

  • Solve complex, real-world infrastructure problems at genuine scale

  • Be part of a collaborative, high-impact engineering culture grounded in integrity, rationality, and collegiality

  • Opportunity to shape the future of a rapidly evolving cloud platform

  • A culture of curiosity and innovation where engineers are empowered to leverage AI and emerging technologies.

#LI-AA1

About Guidewire

Guidewire is the platform P&C insurers trust to engage, innovate, and grow efficiently. We combine digital, core, analytics, and AI to deliver our platform as a cloud service. More than 540+ insurers in 40 countries, from new ventures to the largest and most complex in the world, run on Guidewire.

As a partner to our customers, we continually evolve to enable their success. We are proud of our unparalleled implementation track record with 1600+ successful projects, supported by the largest R&D team and partner ecosystem in the industry. Our Marketplace provides hundreds of applications that accelerate integration, localization, and innovation.

For more information, please visit www.guidewire.com and follow us on Twitter: @Guidewire_PandC.

Guidewire Software, Inc. is proud to be an equal opportunity and affirmative action employer. We are committed to an inclusive workplace, and believe that a diversity of perspectives, abilities, and cultures is a key to our success. Qualified applicants will receive consideration without regard to race, color, ancestry, religion, sex, national origin, citizenship, marital status, age, sexual orientation, gender identity, gender expression, veteran status, or disability. All offers are contingent upon passing a criminal history and other background checks where it's applicable to the position.

Bereit, sich bei Guidewire zu bewerben?
Bei Guidewire bewerben

Über Guidewire

At Guidewire, we are utterly committed to customer success. We combine digital, core, analytics, and AI to deliver our platform as a cloud service to the P&C Insurance industry. And with the largest R&D team, services team, and partner ecosystem in the industry, we continually evolve and innovate to meet our customers’ needs. We put our values of Integrity, Rationality, and Collegiality first, harboring a culture of honesty and openness that our people never want to lose. And we each bring a little quirkiness—and a little genius—to the table. As the landscape of our industry continues to shift, we respond with flexibility and skill. We’re braving uncharted territory, pushing past the convent

Alle Jobs bei Guidewire ansehen →

Ähnliche Jobs

Anduril Industries
Senior Infrastructure Reliability Engineer
Anduril Industries
⚡ Früh bewerben Costa Mesa, California, United... Vor Ort $166,000–$220,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Anduril Industries
Site Reliability Engineer
Anduril Industries
⚡ Früh bewerben Waltham, Massachusetts, United... Vor Ort $166,000–$220,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Anduril Industries
Site Reliability Engineer, Intelligence Systems
Anduril Industries
⚡ Früh bewerben Reston, Virginia, United State... Vor Ort $146,000–$194,000
● Neu 👁 Gesehen ✓ Beworben vor 3 Std.
Roku
Senior Software Engineer, MLOps/SRE
Roku
⚡ Früh bewerben Bengaluru, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Roku
Senior Software Engineer,  SRE
Roku
⚡ Früh bewerben Bengaluru, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Roblox
Senior Site Reliability Engineer, Compute
Roblox
⚡ Früh bewerben San Mateo, CA, United States Vor Ort $243,290–$295,250
● Neu 👁 Gesehen ✓ Beworben vor 5 Std.
Genesys
Senior Operations Reliability Engineer - IAM
Genesys
⚡ Früh bewerben Ontario, Canada Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 19 Std.
LSEG
Senior Manager, Site Reliability Engineering
LSEG
⚡ Früh bewerben IND-BLR-Divyasree Technopolis Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 19 Std.
NL
Senior Cloud Platform & Site Reliability Engineering Lead
National Life Insurance Company
⚡ Früh bewerben Addison, TX; Montpelier, VT Vor Ort $136,875–$200,750
● Neu 👁 Gesehen ✓ Beworben vor 20 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Guidewire

Alle Jobs bei Guidewire ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos