Jobs › Companies › DigitalOcean › Principal Engineer, Rack-Scale GPU Architecture

Sobre este puesto de Principal Engineer, Rack-Scale GPU Architecture en DigitalOcean

DigitalOcean · Híbrido · Seattle

Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here.  We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world. 

We are looking for a Principal Engineer to own how rack-scale GPU systems are architected, orchestrated, and operated for inference at DigitalOcean.

DigitalOcean is the Inference Cloud. The unit of GPU capacity is changing underneath the entire industry: for a decade the server was the boundary, and now it is the rack. GB300 NVL72 and the Vera Rubin generation behind it present 72 or more GPUs inside a single coherent NVLink domain, which breaks most of the assumptions our scheduling, networking, failure-handling, and capacity models were built on. Getting this right determines whether we can serve trillion-parameter models with the economics our customers expect.

This role owns that transition. You will define the reference architecture for rack-scale inference at DigitalOcean—how NVLink domains are carved and allocated, how the scale-out RDMA fabric is designed and tuned around them, how Kubernetes is taught to schedule against topology it was never designed to understand, and what happens when one GPU in a 72-GPU coherent domain fails at 3am. You'll work across our data center, networking, and inference platform teams, and directly with NVIDIA on pre-silicon enablement and bringup.

What You'll Be Doing:

  • Owning the end-to-end reference architecture for rack-scale inference systems—GB300 NVL72 today, Vera Rubin–class NVL144 and successors next—from fabric design through to what a customer can actually schedule
  • Defining how NVLink domains are partitioned, allocated, and isolated across tenants and workloads, including multi-node NVLink and IMEX domain provisioning, fabric manager configuration, and the blast radius each choice implies
  • Architecting the scale-out RDMA fabric that connects racks—RoCEv2 or InfiniBand, rail-optimized topologies, congestion control, GPUDirect RDMA paths—and tuning NCCL to actually use it well
  • Making Kubernetes topology-aware for these systems: DRA for GPU and NVLink-domain allocation, multi-node inference primitives such as LeaderWorkerSet, gang scheduling, Topology Manager and NUMA alignment, and the GPU and Network Operators that underpin all of it
  • Driving multi-node disaggregated serving on this hardware in partnership with the inference platform team—how prefill and decode pools map onto NVLink domains and rail boundaries, and where the KV cache moves between them
  • Owning failure semantics and serviceability at rack scale: health checking and DCGM-based diagnostics, drain and repair workflows, degraded-domain scheduling, and firmware and driver lifecycle across a fleet of racks
  • Establishing the validation and burn-in methodology that qualifies a rack for production—collective bandwidth and latency characterization, NCCL and end-to-end inference benchmarks, and the acceptance gates a rack must pass
  • Partnering with our data center engineering team on power, liquid cooling, and rack density constraints, and translating those into what is actually deployable and at what cost
  • Working directly with NVIDIA and ODM partners on pre-silicon enablement, early hardware bringup, and roadmap feedback
  • Setting technical direction, mentoring senior and staff engineers across teams, and representing DigitalOcean upstream, at conferences, and in customer architecture reviews

What You'll Add to DigitalOcean:

  • 12+ years in large-scale compute, HPC, or GPU infrastructure engineering, including hands-on responsibility for systems in production
  • Deep understanding of GPU interconnect and system topology—NVLink and NVSwitch, NVLink domains, PCIe, NUMA, and how each shapes what a workload can be placed where
  • Strong RDMA networking background: RoCEv2 or InfiniBand at scale, GPUDirect RDMA, congestion control and lossless fabric tuning, and rail-optimized cluster topologies
  • Practical NCCL expertise—not just running it, but debugging it: algorithm and protocol selection, topology detection, and diagnosing collectives that are slow for non-obvious reasons
  • Substantial Kubernetes depth for accelerated workloads: device plugins and DRA, scheduling extensions, operators, and a clear-eyed view of where Kubernetes' model breaks down against rack-scale hardware
  • Strong systems programming and automation skills in Go and Python, and comfort at the driver, firmware, and BMC layer when the problem lives there
  • Experience operating fleets where hardware failure is routine, including the diagnostics, automation, and blast-radius thinking that makes that survivable
  • Excellent written and verbal communication, and a track record of leading architecture across organizational boundaries—hardware, networking, platform, and product

Bonus:

  • Direct experience bringing up GB200 or GB300 NVL72 systems, or comparable rack-scale platforms, in a production environment
  • Experience with multi-node inference for very large models—expert or pipeline parallelism spanning NVLink domains, and the placement constraints that follow
  • Familiarity with liquid-cooled and high-density rack deployments, and the operational realities of servicing them
  • Contributions to relevant open source: Kubernetes SIG-Node or scheduling, NVIDIA GPU or Network Operator, DRA drivers, Kueue, LeaderWorkerSet, or NCCL
  • Experience running heterogeneous fleets spanning multiple accelerator vendors and generations under a single orchestration model

Compensation Range: 

  •  $249,600 - $312,000

*This is a hybrid role

JR: 2026-8161

#LI-Hybrid

Why You’ll Like Working for DigitalOcean

  • We innovate with purpose. You’ll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of responsibility for customers, products, employees, and decisions.
  • We prioritize career development. At DO, you’ll do the best work of your career. You will work with some of the smartest and most interesting people in the industry. We are a high-performance organization that will always challenge you to think big. Our organizational development team will provide you with resources to ensure you keep growing. We provide employees with reimbursement for relevant conferences, training, and education. All employees have access to LinkedIn Learning's 10,000+ courses to support their continued growth and development.
  • We care about your well-being. Regardless of your location, we will provide you with a competitive array of benefits to support you from our Employee Assistance Program to Local Employee Meetups to flexible time off policy, to name a few. While the philosophy around our benefits is the same worldwide, specific benefits may vary based on local regulations and preferences.
  • We reward our employees. The salary range for this position is based on market data, relevant years of experience, and skills. You may qualify for a bonus in addition to base salary; bonus amounts are determined based on company and individual performance. We also provide equity compensation to eligible employees, including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
  • DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.

Application Limit: You may apply to a maximum of 3 positions within any 180-day period. This policy promotes better role-candidate matching and encourages thoughtful applications where your qualifications align most strongly.

¿Listo para postularte en DigitalOcean?
Postúlate en DigitalOcean

Cómo se compara este salario de Principal Engineer

Este puesto paga $280,800/yr — en línea con el rango típico para los puestos de Principal Engineer.

$192,741 la mediana de $242,250 $345,000

Rango típico $206,875–$288,800/yr, a partir de 98 ofertas comparables de Principal Engineer en JobsRadar (salario anualizado en USD). Ver datos salariales de Principal Engineer →

Empleos similares

DigitalOcean
Principal Engineer, Inference Memory and Storage Systems
DigitalOcean
⚡ Postúlate pronto Seattle Híbrido $249,600–$312,000
● Nuevo 👁 Visto ✓ Postulado hace 2h
DigitalOcean
Principal Engineer, Model Optimizations
DigitalOcean
⚡ Postúlate pronto Seattle Híbrido $249,600–$312,000
● Nuevo 👁 Visto ✓ Postulado hace 2h
Mastercard
Principal Software Development Engineer
Mastercard
⚡ Postúlate pronto Arlington, Virginia Presencial $195,000–$323,000
● Nuevo 👁 Visto ✓ Postulado hace 11h
Blue Origin
Sr Principal Thermal Engineer, Project Sunrise Orbital Data Centers
Blue Origin
⚡ Postúlate pronto Denver, CO Presencial $205,695–$287,972
● Nuevo 👁 Visto ✓ Postulado hace 11h
Coupang
Principal Engineer, ML
Coupang
⚡ Postúlate pronto Seattle, USA Presencial $207,900–$207,900
● Nuevo 👁 Visto ✓ Postulado hace 18h
CI
Principal, Security Engineer
Coupang Internal
⚡ Postúlate pronto Seattle, USA Presencial
● Nuevo 👁 Visto ✓ Postulado hace 18h
Coupang
Principal, Security Engineer
Coupang
⚡ Postúlate pronto Mountain View, USA; Seattle, U... Híbrido $209,000–$209,000
● Nuevo 👁 Visto ✓ Postulado hace 18h
Yoodli AI Roleplays
Principal Software Engineer- Backend (Admin Workflows)
Yoodli AI Roleplays
⚡ Postúlate pronto Seattle, WA Híbrido $180,000–$210,000
● Nuevo 👁 Visto ✓ Postulado hace 23h
Nordstrom
Principal Engineer - Agentic AI Platform (Hybrid - Seattle,WA)
Nordstrom
⚡ Postúlate pronto Seattle, WA Presencial $191,000–$297,000
● Nuevo 👁 Visto ✓ Postulado hace 1d

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en DigitalOcean

Ver todos los empleos en DigitalOcean →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis