Sobre esta vaga de Inference Engineer na Weekday AI
𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀
𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟮𝟰𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟯𝟲𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟮𝟰-𝟯𝟲 𝗟𝗣𝗔)
Experience: 5+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are seeking a highly skilled Inference Engineer with strong expertise in Large Language Models (LLMs) and vLLM to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.
As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.
Requirements
Key Responsibilities
- Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.
- Build and optimize high-performance model serving pipelines using vLLM and other modern inference frameworks.
- Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.
- Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.
- Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.
- Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.
- Develop APIs, microservices, and deployment workflows for AI-powered applications.
- Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.
- Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.
- Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.
What Makes You a Great Fit
- 5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.
- Strong hands-on expertise with Large Language Models (LLMs) and vLLM for production-scale inference.
- Experience deploying and optimizing transformer-based models using modern inference frameworks.
- Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.
- Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.
- Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
- Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.
- Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.
- Experience implementing monitoring, benchmarking, and observability for AI inference workloads.
- Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.
- Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.
- Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.