Über diese Senior Infrastructure SRE Stelle bei Pointclickcare
About the role
PointClickCare builds cloud platforms that power safer, more connected care for millions of patients. You'll design and operate resilient infrastructure services across Azure, AWS, GCP, and on-prem environments, driving reliability, automation, and operational excellence for identity, compute, storage, messaging, and shared services with an SRE-first mindset.
What you'll do
- Design and implement highly available infrastructure solutions for compute, storage, identity, messaging, and shared services
- Build and maintain Infrastructure as Code using Terraform or Pulumi; establish best practices and standards
- Automate operational workflows to eliminate toil: auto-remediation, self-healing systems, capacity planning
- Define and track SLIs and SLOs for critical services; manage error budgets and reliability targets
- Participate in the on-call rotation and lead incident response for complex infrastructure issues; conduct blameless post-mortems and drive systemic fixes
- Develop observability strategies: metrics, logs, distributed tracing, alerting frameworks
- Apply AI-assisted tooling to reduce toil and speed up investigation: log analysis, alert triage, runbook and post-mortem drafting, automation scaffolding
- Own reliability for one or more infrastructure domains end to end, driving multi-team initiatives with product engineering from problem definition through adoption
- Mentor intermediate SREs; review infrastructure changes; establish operational best practices
What you'll bring
Must-haves
- 5+ years of hands-on experience operating and designing cloud infrastructure
- Expert-level understanding of the core services of either Azure or AWS
- Working proficiency in at least one additional platform (Azure, AWS, or GCP)
- Experience designing and supporting production infrastructure that spans multiple cloud platforms
- 3+ years of production experience with Infrastructure as Code (Terraform, Pulumi, CloudFormation)
- Ability to design scalable, reusable IaC modules and enforce GitOps workflows
- Experience managing IaC across multiple cloud providers, including module design, state layout, and provider-specific resource differences
- Strong proficiency in at least one programming language (Python, Go, Bash) for production automation
- Demonstrated ability to write tested, maintainable automation and tooling
- Practical application of SRE principles in production environments
- Experience defining and managing SLIs/SLOs, error budgets, toil metrics
- Track record of improving system reliability (e.g., MTTR reduction, availability improvements)
- Demonstrated depth across key infrastructure services:
- Expert-level experience running Kubernetes and containerized workloads in production on managed Kubernetes (AKS, EKS, or equivalent), plus VM-based compute
- Practical experience operating a service mesh in production (Istio preferred; Linkerd or equivalent)
- Strong proficiency with enterprise identity and SSO: SAML, OAuth/OIDC, LDAP, and cloud IAM
- Working knowledge of an enterprise federation platform such as PingFederate, Entra ID, Okta, or ADFS
- Strong proficiency in storage solutions (object, block, and file storage; Kubernetes persistent volumes)
- Working knowledge of designing and operating infrastructure in a regulated environment (HIPAA, SOC 2, PCI, FedRAMP, or equivalent)
- Familiarity with audit evidence, access controls, encryption in transit and at rest, and data residency constraints
- Proven track record of measurably reducing operational toil through automation (e.g. ticket volume, manual runbook executions, hours reclaimed)
- Strong communication and documentation skills; demonstrated ability to influence engineering teams
Nice-to-haves
- 2+ years in healthcare technology or highly regulated SaaS environments (HIPAA, SOC 2, HITRUST)
- Cloud certifications: Azure Solutions Architect Expert, AWS Solutions Architect Professional, GCP Professional Cloud Architect, or equivalent
- Working experience operating Kubernetes at scale across multiple clusters (CKA or CKAD certification a plus)
- Working knowledge of messaging and event-streaming platforms (Kafka, Azure Service Bus, Event Hubs, SQS, Pub/Sub)
- Practical experience building CI/CD pipelines and deployment automation (GitLab, GitHub Actions, ArgoCD)
- Familiarity with AI-assisted engineering and operations tooling: LLM-based coding assistants, AIOps, agentic incident investigation
- Judgment about where AI belongs in an operational workflow, including safe handling of sensitive data in prompts and reviewing generated changes before they reach production contribution to open-source SRE tools or infrastructure projects (GitHub profile, PRs merged)
Education
- Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field
- OR equivalent practical experience with a proven track record in infrastructure and SRE practices evidence of continuous learning and staying current with SRE and cloud-native trends
How we work
- Transparent collaboration: We work in the open using OKRs, cross-functional retrospectives, and public roadmaps so everyone knows priorities and progress
- Blameless culture: We confront problems courageously through structured post-mortems and root cause analyses, focusing on systems improvement not individual blame
- Data-driven decisions: We use metrics, APM, and observability data to make evidence-based choices about reliability and performance investments
- Continuous learning: We learn from incidents through retrospectives and RCAs, sharing knowledge across teams to prevent repeat issues
- Outcome accountability: We're accountable for customer and business results, not just completing tasks, measuring success by impact
- Thoughtful experimentation: We hold strong opinions loosely, testing concepts and running small experiments before scaling solutions
- Iterative delivery: We think big but act small, using Scrum, delivery plans, and frequent milestones to ship incrementally and learn fast
- Inclusive environment: We create space to listen and learn, actively growing our Ally Community to support equity, belonging, and career growth for all