Jobs Companies NVIDIA Principal Software Engineer – Large-Scale LLM Memory and Storage Systems

Sobre este puesto de Principal Software Engineer – Large-Scale LLM Memory and Storage Systems en NVIDIA

NVIDIA · Presencial · US, CA, Santa Clara

NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of any single GPU, this platform enables efficient, resilient deployment of cutting-edge LLM workloads.


We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of large-scale LLM and storage systems.


What you'll be doing:

  • Design and evolve a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage to support large-scale LLM inference.

  • Architect and implement deep integrations with leading LLM serving engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters.

  • Co-design interfaces and protocols that enable disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage (GPU, CPU, local disk, and remote memory) for high-throughput, low-latency inference.

  • Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools.

  • Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives).

What we need to see:

  • Masters or PhD or equivalent experience

  • 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of delivering production services.

  • Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency.

  • Distributed caching or key-value systems, especially designs optimized for low latency and high concurrency.

  • Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters.

  • Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput.

  • Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams.

Ways to stand out from the crowd:

  • Prior contributions to open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse.

  • Experience designing unified memory or storage layers that expose a single logical KV or object model across GPU, host, SSD, and cloud tiers, especially in enterprise or hyperscale environments.

  • Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML.

With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our special engineering teams are growing fast. If you're a creative and autonomous engineer with a genuine passion for technology, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until January 13, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

¿Listo para postularte en NVIDIA?
Postúlate en NVIDIA

Sobre NVIDIA

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA .

Ver todos los empleos en NVIDIA →

Empleos similares

NVIDIA
Distinguished Engineer, Production Engineering, Cluster Management
NVIDIA
⚡ Postúlate pronto US, CA, Santa Clara Presencial
● Nuevo 👁 Visto ✓ Postulado hace 3d
NVIDIA
Distinguished Engineer, Production Engineering, Data Center Automation
NVIDIA
⚡ Postúlate pronto US, CA, Santa Clara Presencial
● Nuevo 👁 Visto ✓ Postulado hace 4d
NVIDIA
Principal Firmware Engineer - Data Center Server Management
NVIDIA
⚡ Postúlate pronto US, CA, Santa Clara Presencial
● Nuevo 👁 Visto ✓ Postulado hace 5d
Lloyds Banking Group
Principal Engineer – AI Productivity Lab
Lloyds Banking Group
⚡ Postúlate pronto London Presencial £92,701–£169,403
● Nuevo 👁 Visto ✓ Postulado hace 2h
HP
Principal Hypervisor Engineer
HP
⚡ Postúlate pronto Cambridge, Cambridgeshire, Uni... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 2h
HP
Principal Embedded Firmware and Software Engineer
HP
⚡ Postúlate pronto Spring, Texas, United States o... Presencial $147,050–$230,850
● Nuevo 👁 Visto ✓ Postulado hace 2h
Wellington Management
Principal, Workday Integrations Engineer
Wellington Management
⚡ Postúlate pronto Other Remote · restringido por ubicación $90,000–$180,000
● Nuevo 👁 Visto ✓ Postulado hace 3h
AstraZeneca
Principal Product Engineer - Evinova
AstraZeneca
⚡ Postúlate pronto US - Gaithersburg - MD Presencial $172,466–$258,700
● Nuevo 👁 Visto ✓ Postulado hace 4h
AstraZeneca
Associate Principal AI Engineer
AstraZeneca
⚡ Postúlate pronto Mexico - Guadalajara Presencial
● Nuevo 👁 Visto ✓ Postulado hace 4h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en NVIDIA

Ver todos los empleos en NVIDIA →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis