Jobs Companies NVIDIA DL Performance Software Engineer - LLM Inference

Über diese DL Performance Software Engineer - LLM Inference Stelle bei NVIDIA

NVIDIA · Vor Ort · Canada, Toronto

We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference software, optimize GPU kernels, drive industry benchmarks, and work with state-of-the-art research techniques to improve serving efficiency. You’ll collaborate across inference performance, kernels, training, large-scale serving, and research teams to push the frontier of accelerated computing for AI.
 

What you’ll be doing:

  • Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features and serving runtime algorithms.

  • Profile and optimize the inference framework (vLLM) with methods like speculative decoding, 5D Parallelism, and prefill-decode disaggregation.

  • Architect novel frameworks and runtime optimizations for inference infrastructure, benchmarking, and kernels.

  • Conduct and publish original research that advances the Pareto frontier in ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into production-grade, open-source software.

  • Develop, optimize, and benchmark GPU kernels (both hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization.

What we need to see:

  • Bachelor’s, Master’s, or PhD degree in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE).

  • 5+ years of industry experience in software engineering or equivalent research experience. 

  • Strong programming skills in Python and one of C/C++, Go, or Rust. Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, deep learning theories.

  • Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).

  • Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).

  • Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.

Ways to Stand out from the Crowd

  • Experience developing major features and optimizations for LLM inference engines (e.g., vLLM, SGLang).

  • Hands-on work with LLM inference and training runtimes (deploying LLMs to production, large-scale LLM pre-training and RL), ML compilers and DSLs (e.g., Triton, CuTe, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).

  • Experience with speculative decoding training and runtime features: tree-structured drafting, parallel drafting, diffusion LLMs, DFlash, EAGLE.

  • Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.

  • At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential.

Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you’re excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we’d love to hear from you.


#LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 135,000 CAD - 185,000 CAD for Level 3, and 170,000 CAD - 220,000 CAD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 10, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

Bereit, sich bei NVIDIA zu bewerben?
Bei NVIDIA bewerben

Über NVIDIA

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA .

Alle Jobs bei NVIDIA ansehen →

Ähnliche Jobs

Kong
Software Engineer
Kong
⚡ Früh bewerben Toronto, Canada Hybrid CA$118,000–CA$140,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Geotab
Lead Software Developer
Geotab
⚡ Früh bewerben Oakville, Ontario - Canada; To... Vor Ort CA$116,200–CA$155,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Geotab
Senior Software Developer (Geotab Vitality)
Geotab
⚡ Früh bewerben Oakville, Ontario - Canada; To... Vor Ort CA$101,600–CA$132,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Geotab
Senior Software Developer - Mobile
Geotab
⚡ Früh bewerben Oakville, Ontario - Canada; To... Vor Ort CA$104,400–CA$135,700
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Internship List
Software Developer Intern (Winter/January 2027, 8 Months)
Internship List
⚡ Früh bewerben Oakville, Ontario - Canada; To... Vor Ort CA$68,640–CA$81,120
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Internship List
Software Developer Intern, MyGeotab (Winter/January 2027, 4-8 Months)
Internship List
⚡ Früh bewerben Oakville, Ontario - Canada; To... Vor Ort CA$49,920–CA$81,120
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Internship List
Software Developer Intern, Geotab Vitality (Winter/January 2027, 4 Months)
Internship List
⚡ Früh bewerben Oakville, Ontario - Canada; To... Vor Ort CA$68,640–CA$81,120
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Stripe
Staff Software Engineer, Financial Crimes
Stripe
⚡ Früh bewerben Toronto Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Stripe
Software Engineer, Machine Learning Infrastructure
Stripe
⚡ Früh bewerben Toronto, Canada Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei NVIDIA

Alle Jobs bei NVIDIA ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos