Jobs Companies Meesho Engineering Manager – AI Engineering

Sobre este puesto de Engineering Manager – AI Engineering en Meesho

Meesho · Presencial · Bangalore, Karnataka

About Meesho

Meesho is India's fastest-growing internet commerce company, on a mission to democratize e-commerce for everyone. We serve millions of customers and over 1.75 million sellers through technology-driven innovation, building the scalable systems that power Meesho's most critical surfaces — Search, Recommendations, Personalized Ranking, Logistics, Fraud Detection, and Image Match.

The AI Platform sits at the heart of this. It serves a peak of 1M+ real-time deep-learning model inferences per second on ordinary days, scaling 3x+ on sale days — with the reliability that scale demands. The team works at the frontier of applied AI and infrastructure — multi-region inference, novel embedding-search algorithms, and optimized open-weight LLM models — squeezing out every bit of computation and passing the cost savings straight back to customers.


About the Role

We are looking for an experienced Engineering Manager – AI Engineering to lead the development of scalable AI platforms and infrastructure while managing high-performing engineering teams. You will drive the design, delivery, and optimization of production-grade AI systems powering AI use cases across Meesho.



What You'll Do

  • Lead, mentor, and grow a team of AI engineers — setting technical direction, raising the engineering bar, and owning execution and delivery end to end.

  • Architect and scale Meesho's AI platform: cross-region model inference, multi-GPU fleet allocation and management, distributed training, and feature-engineering infrastructure.

  • Drive inference optimization across the full stack — GPU kernel tuning, quantization (including outlier/tail-distribution handling), and memory/IO-bandwidth optimization — while building agents that codify and delegate known optimization procedures.

  • Optimize open-weight models at both the model and inference-engine level — distillation, quantization, speculative decoding, KV-cache and serving-engine tuning.

  • Scale data-science productivity through autonomous, agent-driven workflows spanning feature engineering, model training, and rollout.

  • Push the frontier across MLOps, LLMOps, compute efficiency, and distributed ML systems.

  • Partner with Product, Data Science, and Platform teams to turn AI capabilities into production impact for millions of users.

  • Own the team's operating rhythm: hiring, performance management, sprint planning, and OKRs.

  • What You'll Need

  • Bachelor's or Master's in Computer Science or a related field.

  • 9+ years of software engineering experience, including 2+ years managing engineers.

  • Strong hands-on experience with the modern LLM inference stack — TensorRT-LLM, vLLM, SGLang — and with production, low-latency model serving at scale.

  • Depth in inference optimization: GPU kernel tuning, quantization, speculative decoding, KV-cache and memory/IO optimization. CUDA / GPU programming experience is a strong plus.

  • Experience with distributed training and the frameworks behind it — PyTorch FSDP, DeepSpeed, Megatron, or Ray.

  • Experience running GPU fleets in production — Kubernetes (ideally GKE), GPU scheduling and allocation, and multi-region/multi-cluster deployment.

  • Familiarity with building LLM-powered agents and agentic workflows, and a point of view on where autonomy can replace manual engineering toil.

  • Experience with big-data and streaming stacks — Spark, Flink, or similar.

  • Proficiency in Python; systems-level fluency (C++ / Go / Rust) for performance-critical paths.

  • Strong leadership, problem-solving, and stakeholder-management skills.

  • Preferred

  • Open-source contributions to inference engines, training frameworks, or ML infra tooling.

  • Experience managing GPU cost/efficiency (FinOps) for a large fleet on Cloud and Neo-Clouds. 

  • Track record building platforms for high-scale consumer products (millions of users).

  • Familiarity with observability and reliability for ML systems (SLOs, autoscaling, incident response).

  • ¿Listo para postularte en Meesho?
    Postúlate en Meesho

    Empleos similares

    Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

    Más empleos en Meesho

    Ver todos los empleos en Meesho →

    Postúlate ahora
    🤖

    Un momento — para

    JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

    Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

    Catch your next role the second it’s posted.

    Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

    Create free account

    Free forever · takes 30 seconds · already have one?

    Toma ventaja en tu búsqueda de empleo.

    Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

    Únete al canal — es gratis