Jobs Companies Chalice Custom Algorithms Senior Site Reliability Engineer

Über diese Senior Site Reliability Engineer Stelle bei Chalice Custom Algorithms

Chalice Custom Algorithms · Hybrid · Hybrid/Remote

Location: Remote (US); Hybrid/Remote (NYC)

About Chalice

Chalice Custom Algorithms (chalice.ai) is the leading AI solution for brands applying their own data and analytics to real-time decisioning in ad buying. Our platform automates data ingestion, predictive analytics, and the deployment of custom bidding logic across all major DSPs, Meta, and YouTube. Chalice has been recognized as "Best Demand Side Tech" by AdExchanger and powered AdWeek's "Best Use of Programmatic" in the 2023 Media Plan of the Year awards.

Organizational Context

This role reports to the VP of engineering and works cross-functionally with Engineering, Data Science, Machine Learning, and Product. The Senior SRE sits at the intersection of infrastructure, ML systems, and platform governance, with broad influence across teams.

What "Senior" Means at Chalice

This role focuses more on architectural leverage than operational work. You will:

• Recommend architectural direction

• Reduce systemic complexity

• Introduce durable patterns

• Identify architectural risk early

• Retire services when necessary

• Define reliability standards across teams

• Shape how our AI platform is delivered to customers

This role will influence infrastructure, ML operations, governance, and platform strategy.

We are evolving toward:

• Event-driven system design

• Container deployments to customer and partner infrastructure

• Reduced architectural rigidity

• Strong internal platform standards

Mission

As Senior Site Reliability Engineer, you will define and operate the architectural backbone of Chalice's AI platform, reporting directly to the VP of Engineering and working cross-functionally with Engineering, Data Science, Machine Learning, and Product.

In addition to building systems, you will mentor and elevate other engineers in infrastructure best practices, operational rigor, and architectural thinking. You will help establish a culture of reliability, ownership, and continuous improvement across the organization.

You will design and operate a scalable, event-driven, multi-tenant ML infrastructure platform that supports:

• Distributed ML training (Databricks, Ray, Flyte on EKS)

• Containerized product delivery to external customers

• Internal event-driven services across AWS

• Centralized state-store-driven orchestration

• Governance across adtech integrations and third-party APIs

What You'll Own

1. Event-Driven Platform Architecture

You will build and support Chalice's event-driven / API first platform and participate the build-out of a scalable SRE function to support it. You will design and implement AWS event-driven systems using:

  • EventBridge

  • MSK / Kafka

  • Kinesis

  • Lambda / Fargate

  • SQS / SNS / Step Functions

Architect centralized state stores (DynamoDB, Redis, Postgres) that:

  • React to signals

  • Trigger downstream services

  • Maintain system integrity

Establish architectural standards for:

  • Idempotency

  • Replay safety

  • Event schema governance

  • Operational clarity and traceability

2. Kubernetes & Control Plane Ownership

• Operate multi-cluster Kubernetes environments in production

• Understand and tune: API server scaling, etcd performance, RBAC architecture, admission controllers

• Implement: GitOps patterns, progressive delivery, cluster-level security policies, multi-tenant isolation

Bonus: have built internal developer platforms, managed customer-facing container workloads, operated ML workloads in Kubernetes.

3. Infrastructure as Code, Governance & CI/CD Evolution

• Participate in Terraform module standards and create reusable infrastructure primitives

• Enforce GitHub guardrails (branch protections, CI gates)

• Evolve and standardize CI/CD pipelines to support automated infrastructure testing, policy validation, progressive deployment, and rollback mechanisms

• Standardize federated identity management, Azure SSO, API authentication patterns, key management, and resource isolation

4. External Model Serving Architecture

We are enabling customers to run our models without sending us their data. You will have a decision in:

• Model packaging standards (OCI images)

• Secret injection patterns

• Network isolation models

• Telemetry back to Chalice

• Upgrade and compatibility strategy

• Runtime configuration contracts

This is effectively building a deployable AI product platform.

5. Observability & Reliability Standards

• Define SLIs, SLOs, and error budgets

• Separate ML reliability from infrastructure reliability

• Implement distributed tracing

• Design golden signals for event pipeline health, data freshness, model serving reliability, and control plane stability

• Define on-call structure, escalation paths, incident response standards, and postmortem processes

We use Datadog for observability, but you will define what "good" looks like.

6. Databricks as a Platform

• SSO implementation and maintenance

• Authentication and provisioning

• Terraform-based deployments

• Cluster policies

• Unity Catalog governance

You’ll be a part of a team that owns Databricks as infrastructure, not as a notebook environment.

Ideal Background

• 8-10 years of experience designing and operating production infrastructure at scale

• Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience

• Deep AWS architecture expertise across networking, SSO, compute, storage, and event-driven services

• Experience managing Kubernetes control planes in production environments

• Experience building or migrating event-driven systems at scale

• Experience in ML-heavy or data-heavy environments

• Has replaced or eliminated legacy infrastructure components and simplified system design

• Has owned real production failures and led postmortems that resulted in systemic improvements

• Strong architectural judgment and the ability to challenge assumptions constructively

• Certifications (AWS, Databricks, Kubernetes, etc.) are a plus but not required.

What We're Looking For

The ideal candidate:

• Has refactored/replaced legacy architectures in the past

• Has migrated systems safely

• Has designed something from zero

• Challenges leadership constructively

Interview Process

We highly value architectural roles and aim to evaluate real-world systems thinking rather than trivia or syntax knowledge.

1. Conversational + Architecture Discussion: A live discussion focused on past systems, decision-making frameworks, and architectural tradeoffs.

2. Architecture + process deep dive: A practical exercise evaluating structure, reliability, testing, and operational clarity.

3. CTO Strategic fit: Technical and strategic discussion on platform evolution, organizational design, and infrastructure at scale.

4. CEO Conversation: Final conversation on company vision, long-term platform direction, and cultural alignment.

Why Join Us

  • Innovative environment: Work with cutting-edge technology in a fast-paced, exciting industry.

  • Real scope: You'll own the relationship, the strategy, and the growth plan on accounts that matter to the business.

  • Genuinely novel work: Custom models built for a specific client's outcome, in a category that is still being defined.

  • Collaborative culture: A diverse, inclusive team where creativity and collaboration thrive in a relaxed, welcoming space — complete with the occasional four-legged coworker.

  • Career growth: Ample opportunities for professional development and advancement.

  • Competitive compensation: An attractive salary and benefits package, including unlimited PTO.

Benefits

  • Medical, dental, and vision insurance

  • 401(k) options

  • Unlimited PTO

  • 11 company holidays

  • Office-wide closure between Christmas Eve and New Year's

  • Office start-up stipend and in-office meal allowance

Chalice is an equal opportunity employer. We celebrate diversity and are committed to an inclusive workplace.

To learn more about Chalice, visit our website: www.chalice.ai. Chalice participates in E-Verify.

Bereit, sich bei Chalice Custom Algorithms zu bewerben?
Bei Chalice Custom Algorithms bewerben

Wie sich dieses Gehalt für SRE vergleicht

Diese Stelle zahlt $185,000/yrim Einklang mit der üblichen Spanne für SRE Stellen.

$96,500 dem Median $158,000 $228,950

Übliche Spanne $122,200–$195,000/yr, aus 808 vergleichbaren SRE Anzeigen auf JobsRadar (Vergütung auf USD hochgerechnet). Gehaltseinblicke für SRE ansehen →

Ähnliche Jobs

OpenTeams
Site Reliability Engineer / DevSecOps Engineer
OpenTeams
⚡ Früh bewerben Washington, DC Metro; Denver,... Hybrid $145,000–$250,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Saviynt
Principal Site Reliability Engineer
Saviynt
⚡ Früh bewerben Vancouver Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Offchainlabs
Site Reliability Engineer
Offchainlabs
⚡ Früh bewerben United States · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Rain
Site Reliability Engineer
Rain
⚡ Früh bewerben Remote · standortgebunden $170,000–$225,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Nametag
Senior Software Engineer, Reliability
Nametag
⚡ Früh bewerben Remote · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 3 Std.
Recorded Future
Site Reliability Engineer
Recorded Future
⚡ Früh bewerben Gothenburg, Sweden Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 3 Std.
Pinterest
Site Reliability Engineer II, tvScientific
Pinterest
⚡ Früh bewerben San Francisco, CA, US; Remote,... · standortgebunden $114,297–$235,319
● Neu 👁 Gesehen ✓ Beworben vor 3 Std.
Pinterest
Sr. Site Reliability Engineer, tvScientific
Pinterest
⚡ Früh bewerben San Francisco, CA, US; Remote,... · standortgebunden $139,764–$287,749
● Neu 👁 Gesehen ✓ Beworben vor 3 Std.
Prove
Site Reliability Engineer, Sr. Site Reliability Engineer
Prove
⚡ Früh bewerben United States (Remote) · standortgebunden $153,000–$171,000
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Chalice Custom Algorithms

Alle Jobs bei Chalice Custom Algorithms ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos