Jobs Companies Qode Senior Software Engineer (Performance)

À propos de ce poste Senior Software Engineer (Performance) chez Qode

Qode · Sur site · Vietnam, Vietnam, Vietnam

Job Description:

We are looking for a Senior Inference Engineer with a strong foundation in software engineering, distributed systems, and performance optimization to build and optimize inference engines for large-scale LLM serving systems. You will work across both research and production environments, ensuring our LLM serving systems are fast, scalable, and efficient. The role spans the entire inference stack — from kernel and runtime to scheduling, memory management, and distributed execution


Key Responsibilities:

  • Profile, benchmark, and analyze bottlenecks for LLM inference workloads across multiple layers: kernel, memory, networking, and scheduler
  • Optimize inference engines (vLLM, SGLang, TensorRT-LLM) for throughput, latency, memory efficiency, GPU utilization, and cost
  • Implement and fine-tune inference optimization techniques including batching, KV-cache management, quantization, speculative decoding, parallelism strategies, and disaggregated serving
  • Build instrumentation and profiling tools to identify bottlenecks
  • Ensure the reliability of the inference pipeline through A/B launches, rollback, model versioning, and fault tolerance
  • Collaborate with the Platform Engineering team to improve serving architecture based on performance findings
  • Document and share knowledge, contributing to internal best practices and AI open-source projects whenever possible

Requirements

1 - Mandatory:

  • At least 5 years of experience as a Software Engineer, Performance Engineer, or equivalent.
  • Strong foundation in Software Engineering, Software Architecture, and Distributed Systems.
  • Proficiency in at least one of the following languages: Python, Go, or C++.
  • Experience developing or optimizing distributed systems, high-throughput backends, or large-scale serving systems.
  • Experience with benchmarking, profiling, and performance tuning in production environments.
  • Ability to analyze CPU, Memory, Network, or Storage bottlenecks.
  • Strong systems thinking, Root Cause Analysis capabilities, and the ability to solve complex performance problems.
  • Strong ownership mindset and the ability to work independently.

2 - Nice to Have:

  • Experience with Linux internals, kernel tuning, or custom Linux kernel.
  • Understanding of GPU Architecture or CUDA Programming.
  • Experience with AI/ML Serving Systems or LLM Inference.- Have worked with one of the inference engines such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
  • Understanding of batching, KV Cache, quantization, speculative decoding, tensor/pipeline parallelism, or disaggregated serving.
  • Experience with the NVIDIA inference stack (TensorRT, Triton, CUTLASS, NCCL, cuBLAS, cuDNN).
  • Experience with observability stacks such as Prometheus, Grafana, or OpenTelemetry.
  • Open-source contributions or research related to AI Infrastructure, ML Systems, or Performance Optimization.


Prêt à postuler chez Qode ?
Postuler chez Qode

À propos de Qode

Qode is dedicated to helping technical talent around the world find meaningful careers that match their skills and interests. Our platform provides a range of resources and tools that empower job seekers to take control of their careers and connect with top employers across a variety of industries. We believe that every individual deserves to find work that they're passionate about, and we are committed to making that vision a reality.

Qode's team of experienced professionals is passionate about creating a better world of work by providing innovative solutions that improve the job search process for both job seekers and employers. We believe in transparency, trust, and collaboration, and we strive to build strong relationships with our customers and partners. Through our platform, we aim to create a more engaged and fulfilled global workforce that drives innovation and growth.

Voir tous les emplois chez Qode →

Emplois similaires

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Qode

Voir tous les emplois chez Qode →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit