Jobs Companies Crusoe Principal Engineer, CAPE

Über diese Principal Engineer, CAPE Stelle bei Crusoe

Crusoe · Vor Ort · San Francisco, CA - US

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

Conductor is the platform we're building: a self-driving fleet.


Most infrastructure teams watch dashboards and replace dead nodes after they fail. We're building something different — a control plane that predicts, decides, and acts on its own, keeping tens of thousands of accelerators doing useful work at the highest goodput in the industry while optimizing power and cost in real time. The mandate for this role is to make the fleet self-driving.


This is a Principal Engineer role reporting into Cloud Availability, working directly on one of the highest-priority technical charters in the org.


Your Charter — The Problems You'll Own

  1. Unified observability plane. One pane that correlates GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals — so any engineer can diagnose and recover jobs fast.

  2. The fleet as one logical computer. Tens of thousands of accelerators across sites behaving as a single programmable system: one health model, one scheduler, one source of truth.

  3. Closed-loop autonomy, not alerting. Diagnose, decide, and remediate with no human in the loop — drain, checkpoint, replace, and resume a live job automatically. The hard part: earning enough trust to act on a running job.

  4. Goodput as an objective function. Don't just measure useful compute — continuously maximize it, trading scheduling, placement, and maintenance decisions against it as the north-star metric.

  5. Predict failures hours ahead, not seconds after. Forecast GPU, NVLink, optics, and thermal degradation before it stalls a job, and pre-emptively migrate work. An ML problem on noisy hardware telemetry at fleet scale.

  6. Straggler & silent-failure detection. One slow GPU stalls an entire distributed job — isolate the exact rank, GPU, and node from collective-operation signals and surface root cause fast.

  7. Energy-aware compute. Because we own the power stack, schedule, throttle, and place workloads against real-time energy availability, cost, and thermal headroom. The problem you can work on here and nowhere else.

  8. A digital twin of the fleet. Simulate failures, scheduling policies, and remediation logic before they touch production — the foundation that makes real autonomy safe.

  9. Agentic operations. An agent that doesn't just answer "why is this job slow" but proposes and executes the fix, with guardrails and a full audit trail.

  10. Zero-trust, fully auditable multi-tenancy. Every action — human or autonomous — identity-scoped, policy-checked, and streamed to customer audit systems in real time.

  11. Self-qualifying hardware. New and repaired nodes prove themselves through automated burn-in before taking customer load. Fleet growth becomes a software-gated pipeline.

Why This Role Is Different

  • Scale almost no one gets to work on — distributed systems and control theory across tens of thousands of accelerators.

  • A problem unique to Crusoe — we own the stack from energy to compute, so the platform can treat power as a control input, not a constraint.

  • Greenfield authority — you define the operating standard for how we run AI at scale. You don't inherit one.

What You'll Bring

  • 10+ years building infrastructure-layer systems at scale — fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation. This is a systems-builder role, not a consumer of managed cloud services.

  • Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and systems that make autonomous decisions against live production infrastructure.

  • Hands-on fluency with GPU/HPC infrastructure — GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior at the hardware level.

  • Track record of designing and shipping large-scale observability or telemetry platforms that correlate signals across compute, network, and storage layers.

  • Comfort operating in ambiguity and defining the architecture and standards for a system that doesn't exist yet — this is a 0→1 charter, not a maintenance role.

  • Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) and the judgment to know when to build vs. adopt existing tooling.

  • Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus.

  • Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus but not required.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $285,000 - $335,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Bereit, sich bei Crusoe zu bewerben?
Bei Crusoe bewerben

Wie sich dieses Gehalt für Principal Engineer vergleicht

Diese Stelle zahlt $310,000/yrüber der üblichen Spanne für Principal Engineer Stellen.

$209,500 dem Median $275,000 $325,980

Übliche Spanne $235,500–$289,000/yr, aus 27 vergleichbaren Principal Engineer Anzeigen auf JobsRadar (Vergütung auf USD hochgerechnet). Gehaltseinblicke für Principal Engineer ansehen →

Ähnliche Jobs

Fluidstack
Principal Operations Engineer, Reliability
Fluidstack
⚡ Früh bewerben Austin, TX Vor Ort $220,000–$260,000
● Neu 👁 Gesehen ✓ Beworben vor 12 Std.
Faire
Principal Applied AI / ML Engineer
Faire
⚡ Früh bewerben San Francisco, CA Vor Ort $339,000–$466,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Intercom
Senior / Principal Demo Engineer
Intercom
⚡ Früh bewerben San Francisco, California Vor Ort $492,960–$492,960
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Fastly
Principal Engineer - Customer Identity Access Management (CIAM)
Fastly
⚡ Früh bewerben Denver, CO; New York City, NY;... Hybrid $246,550–$295,000
● Neu 👁 Gesehen ✓ Beworben vor 5 Tg.
CoreWeave
Principal Software Engineer, Developer Experience
CoreWeave
⚡ Früh bewerben Livingston, NJ / New York, NY... Vor Ort $227,000–$303,000
● Neu 👁 Gesehen ✓ Beworben vor 5 Tg.
CoreWeave
Principal Operational Readiness Engineer
CoreWeave
⚡ Früh bewerben Livingston, NJ / New York, NY... Vor Ort $198,000–$264,000
● Neu 👁 Gesehen ✓ Beworben vor 5 Tg.
CoreWeave
Principal Security Engineer
CoreWeave
⚡ Früh bewerben Livingston, NJ / New York, NY... Vor Ort $275,000–$330,000
● Neu 👁 Gesehen ✓ Beworben vor 5 Tg.
SoFi
Principal Engineer, Digital Identity
SoFi
⚡ Früh bewerben WA - Seattle; CA - San Francis... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 5 Tg.
BetterUp
Principal Solutions Engineer
BetterUp
⚡ Früh bewerben Austin, TX Hybrid $198,800–$248,500
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Crusoe

Alle Jobs bei Crusoe ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos