Jobs › Companies › Crusoe › Senior Staff Software Engineer, Cloud Availability Platform

Über diese Senior Staff Software Engineer, Cloud Availability Platform Stelle bei Crusoe

Crusoe · Vor Ort · San Francisco, CA - US

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

Build the operating system for the AI datacenter

Crusoe operates one of the world's largest managed GPU fleets, and it is growing fast. A fleet at this scale cannot be run the way GPU clouds have traditionally been run: runbooks, war rooms, and heroics. It has to be run by a unified platform that senses, reasons about, and acts on the entire fleet, so infrastructure that used to take a team to operate takes a service instead. That is what our cloud platform team builds. You will work on the control plane for one of the largest AI fleets in the world, at a point where fleet autonomy is still an open problem: nobody has fully solved this at this scale, in this market.

A true platform, not internal tooling: We want to be explicit about that heading, because it is the thing most infrastructure roles get wrong. Everything we build ships as a product, fleet engineers, SREs, and product teams across Crusoe build their own services and workflows on top of what we ship. A platform team does not scale by doing everyone's work; it scales by making everyone's work self-serve. Concretely:

  • API-first. Every capability is exposed through well-designed, versioned APIs behind a single gateway. If it isn't an API, it doesn't exist. No side doors, including for us.

  • SDKs and paved paths. First-class client libraries, workflow templates, and golden paths so a fleet or SRE engineer can ship a new remediation or lifecycle workflow in days without asking the platform team.

  • Micro frontends and a self-serve portal. Teams plug their own UI surfaces into one developer portal instead of building one-off dashboards. One console for the fleet, extensible by every team.

  • Platform as product. Internal teams are customers. We own contracts, versioning, deprecation policy, quotas, documentation, and support. Adoption is our success metric: the platform wins when other teams choose it because it is the fastest path, not because it is mandated.

What this platform is

Four layers, built as one system:

  • Agents on every site and host that collect telemetry and execute commands.

  • A distributed infra graph: Models system connections down to the rack, fabric, power, and cooling layers. By integrating these connections with telemetry signals, the platform can precisely trace events to identify their blast radius and root cause.

  • A reconciliation core: workflow engine, policy engine, and state reconciler that continuously close the gap between intended state and reality, exposed through the API gateway.

  • Domain services: Services spanning provisioning, firmware upgrade, validation, deployment, repair and RMA, capacity, power and thermal, and Day-2 operations. Built once, run fleet-wide, consumable by any team through APIs and SDKs.

We operate on a continuous autonomy loop—sense, correlate, reason, act, learn—incorporating guardrails that evolve from recommendation to full automation. We treat every recurring manual intervention as a signal to engineer the next automation.

You'll thrive here if you

  • Want to build a platform, not integrate one. This is core distributed-systems engineering: event buses, graph models, reconciliation loops, policy evaluation.

  • Treat internal engineers as customers and sweat API ergonomics, docs, and onboarding the way product teams sweat UX.

  • Like owning a hard abstraction and defending it as ten teams build on top of you.

  • Believe the interesting problems are where physical infrastructure meets software: a firmware counter, a thermal event, and a scheduling decision are one problem, not three.

  • Measure yourself by what stops paging humans, and by how fast another team ships on your platform.

What you'll do

  • Design and build core platform services: RBAC, tenancy, the workflow engine, policy engine, and state reconciler that drive fleet actions safely at scale.

  • Design the public face of the platform: the API gateway, resource model, and versioned API contracts that fleet, SRE, and product teams build against.

  • Build SDKs, workflow templates, and golden paths that make the platform self-serve, plus the developer portal and micro frontend framework that let teams bring their own UI surfaces.

  • Build the inventory and topology graph as the fleet's source of intended truth, and the pipelines that keep it honest against reality (metadata drift is one of our top verified incident root causes; you will kill it).

  • Build site, GPU, and network agents and the event bus that moves fleet telemetry and commands reliably.

  • Deliver the platform roadmap: pilot site on the foundation layer, first site deployed entirely through the platform, zero-downtime firmware upgrades, first fully automated RMA, then 100K+ GPUs on platform with MTTD under 60 seconds and MTTR under 30 minutes.

  • Work with embedded engineers from fleet and production engineering who bring the operational scar tissue, and turn it into services other teams extend.

Requirements

  • 10+ years building distributed systems, control planes, or infrastructure platforms.

  • Strong software engineering skills in Go, Python, or Rust.

  • Experience building platforms other engineers consume: public or internal APIs, SDKs, or developer tooling with real adoption.

  • Depth in at least one of: workflow/orchestration engines (Temporal or similar), event-driven architectures, graph data models, policy/rules engines, or reconciliation-based control loops (Kubernetes operator patterns).

  • Experience running what you build: you have carried a pager for a platform other teams depend on.

  • Systems thinking across the hardware/software boundary.

Bonus experience

  • Internal developer platforms: API gateways, service catalogs, Backstage-style portals, micro frontend architectures.

  • GPU or bare-metal fleet infrastructure: DCGM, Redfish/IPMI, firmware lifecycle.

  • High-cardinality observability platforms (per-GPU telemetry at fleet scale).

  • InfiniBand or RoCE fabrics.

  • AI agents applied to infrastructure triage and autonomous remediation.

About CAPE

Vision. Crusoe's infrastructure runs as a self-aware, self-healing system: anticipating and auto-remediating failures, shaping its own power demand, and tuning silicon-to-orchestration as one instrument. The world's most reliable, efficient, and sustainable AI compute platform.

Mission. We design, build, and operate the world's most reliable and energy-efficient AI infrastructure platform by treating the physical and digital layers as one software-defined system. Every day, for every workload, we automate away the latency, waste, and fragility between stranded energy and delivered intelligence.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $250,000 - $300,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Bereit, sich bei Crusoe zu bewerben?
Bei Crusoe bewerben

Wie sich dieses Gehalt für Staff Engineer vergleicht

Diese Stelle zahlt $275,000/yr — im Einklang mit der üblichen Spanne für Staff Engineer Stellen.

$197,460 dem Median $250,700 $445,000

Übliche Spanne $226,250–$362,500/yr, aus 273 vergleichbaren Staff Engineer Anzeigen auf JobsRadar (Vergütung auf USD hochgerechnet). Gehaltseinblicke für Staff Engineer ansehen →

Ähnliche Jobs

Redwood Materials
Staff Electrical Engineer, Balance of Plant
Redwood Materials
⚡ Früh bewerben San Francisco, California, Uni... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Waabi
Senior / Staff Electrical Engineer
Waabi
⚡ Früh bewerben Pittsburgh, PA Vor Ort $147,000–$250,000
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Reddit
Staff Site Reliability Engineer, Ads
Reddit
⚡ Früh bewerben San Francisco, CA Vor Ort $217,000–$303,900
● Neu 👁 Gesehen ✓ Beworben vor 11 Std.
Gusto, Inc.
Staff Software Engineer, Cloud Infrastructure
Gusto, Inc.
⚡ Früh bewerben Denver, CO - Hybrid; New York,... Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 11 Std.
SA
Senior/Staff Machine Learning Research Engineer, General Agents, Enterprise GenAI
Scale AI
⚡ Früh bewerben San Francisco, CA; New York, N... Vor Ort $290,400–$363,000
● Neu 👁 Gesehen ✓ Beworben vor 13 Std.
SA
Staff Software Engineer, Enterprise GenAI
Scale AI
⚡ Früh bewerben San Francisco, CA; New York, N... Vor Ort $252,000–$315,000
● Neu 👁 Gesehen ✓ Beworben vor 13 Std.
SA
Senior Staff Frontier Agents Engineer
Scale AI
⚡ Früh bewerben San Francisco, CA; New York, N... Vor Ort $288,000–$360,000
● Neu 👁 Gesehen ✓ Beworben vor 13 Std.
CHAOS Industries
Staff Software Engineer
CHAOS Industries
⚡ Früh bewerben San Francisco, California, Uni... Vor Ort $190,000–$260,000
● Neu 👁 Gesehen ✓ Beworben vor 13 Std.
EC
Senior Staff Software Engineer, Linux Kernel
Efficient Computer
⚡ Früh bewerben San Francisco, Bay Area OR Pit... Vor Ort $220,000–$270,000
● Neu 👁 Gesehen ✓ Beworben vor 14 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Crusoe

Alle Jobs bei Crusoe ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos