Sobre este puesto de Senior Software Engineer (Performance) en Qode
Job Description:
We are looking for a Senior Inference Engineer with a strong foundation in software engineering, distributed systems, and performance optimization to build and optimize inference engines for large-scale LLM serving systems. You will work across both research and production environments, ensuring our LLM serving systems are fast, scalable, and efficient. The role spans the entire inference stack — from kernel and runtime to scheduling, memory management, and distributed execution
Key Responsibilities:
- Profile, benchmark, and analyze bottlenecks for LLM inference workloads across multiple layers: kernel, memory, networking, and scheduler
- Optimize inference engines (vLLM, SGLang, TensorRT-LLM) for throughput, latency, memory efficiency, GPU utilization, and cost
- Implement and fine-tune inference optimization techniques including batching, KV-cache management, quantization, speculative decoding, parallelism strategies, and disaggregated serving
- Build instrumentation and profiling tools to identify bottlenecks
- Ensure the reliability of the inference pipeline through A/B launches, rollback, model versioning, and fault tolerance
- Collaborate with the Platform Engineering team to improve serving architecture based on performance findings
- Document and share knowledge, contributing to internal best practices and AI open-source projects whenever possible
Requirements
1 - Mandatory:
- At least 5 years of experience as a Software Engineer, Performance Engineer, or equivalent.
- Strong foundation in Software Engineering, Software Architecture, and Distributed Systems.
- Proficiency in at least one of the following languages: Python, Go, or C++.
- Experience developing or optimizing distributed systems, high-throughput backends, or large-scale serving systems.
- Experience with benchmarking, profiling, and performance tuning in production environments.
- Ability to analyze CPU, Memory, Network, or Storage bottlenecks.
- Strong systems thinking, Root Cause Analysis capabilities, and the ability to solve complex performance problems.
- Strong ownership mindset and the ability to work independently.
2 - Nice to Have:
- Experience with Linux internals, kernel tuning, or custom Linux kernel.
- Understanding of GPU Architecture or CUDA Programming.
- Experience with AI/ML Serving Systems or LLM Inference.- Have worked with one of the inference engines such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
- Understanding of batching, KV Cache, quantization, speculative decoding, tensor/pipeline parallelism, or disaggregated serving.
- Experience with the NVIDIA inference stack (TensorRT, Triton, CUTLASS, NCCL, cuBLAS, cuDNN).
- Experience with observability stacks such as Prometheus, Grafana, or OpenTelemetry.
- Open-source contributions or research related to AI Infrastructure, ML Systems, or Performance Optimization.