Jobs โ€บ Companies โ€บ Weekday AI โ€บ Inference Engineer

About this Inference Engineer role at Weekday AI

Weekday AI ยท Onsite ยท Bengaluru, Karnataka, India

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿฐ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฏ๐Ÿฒ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿฎ๐Ÿฐ-๐Ÿฏ๐Ÿฒ ๐—Ÿ๐—ฃ๐—”)

Experience: 5+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are seeking a highly skilledย Inference Engineerย with strong expertise inย Large Language Models (LLMs)ย andย vLLMย to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.

As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.

Requirements

Key Responsibilities

  • Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.
  • Build and optimize high-performance model serving pipelines usingย vLLMย and other modern inference frameworks.
  • Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.
  • Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.
  • Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.
  • Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.
  • Develop APIs, microservices, and deployment workflows for AI-powered applications.
  • Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.
  • Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.
  • Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.

What Makes You a Great Fit

  • 5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.
  • Strong hands-on expertise withย Large Language Models (LLMs)ย andย vLLMย for production-scale inference.
  • Experience deploying and optimizing transformer-based models using modern inference frameworks.
  • Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.
  • Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.
  • Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
  • Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.
  • Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.
  • Experience implementing monitoring, benchmarking, and observability for AI inference workloads.
  • Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.
  • Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.
  • Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.
Ready to apply to Weekday AI?
Apply to Weekday AI

About Weekday AI

At Weekday (backed by YC; also Product Hunt #1 product of the day), we are building the next frontier in hiring. We have built the largest database of white collar talent in India and have built outreach tools on top of it to generate highest response rates.

See all jobs at Weekday AI โ†’

Similar jobs

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at Weekday AI

See all jobs at Weekday AI โ†’

Apply now
๐Ÿค–

Whoa โ€” hold up

JobsRadar was built for real people having a rough time in their job search โ€” not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back โ€” just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role โ€” salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel โ€” it's free