Jobs Companies Cast AI Senior ML Engineer - Kimchi (LLM Inference Optimization)

Sobre este puesto de Senior ML Engineer - Kimchi (LLM Inference Optimization) en Cast AI

Cast AI · Presencial · Austria; France; Germany; Italy; Netherlands; Poland; Spain; United Kingdom

Why Cast AI?

Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale. By embedding autonomous decision-making directly into Kubernetes and cloud environments, Cast AI continuously optimizes performance, reliability, and efficiency in production.
The old way doesn't work. As Kubernetes and AI environments grow, manual decisions don’t. Cast AI replaces tickets, alerts, and manual tuning with continuous automation that adapts infrastructure as conditions change. Efficiency and cost savings follow naturally from that automation.
Over 2,100 companies already rely on Cast AI, including Akamai, BMW, Cisco, FICO, HuggingFace, NielsenIQ, Swisscom, and TGS.

Global team, diverse perspectives

We're headquartered in Miami, but our impact is international. We take a global and intentional approach to diversity. Today, Cast AI operates across 34 countries spanning Europe, North America, Latin America, and APAC, bringing a wide range of perspectives into how we build and lead.

Unicorn momentum

In January 2026, we achieved unicorn status with a strategic investment from Pacific Alliance Ventures, the corporate venture arm of Shinsegae Group (a $50+ billion Korean conglomerate). Our valuation now exceeds $1 billion, and we're just getting started.

Join us as we build the future of autonomous infrastructure.

About the role

Throughput. Latency. KV cache utilization.

Move those three numbers in the right direction, and two things happen: customers get faster, cheaper inference, and our margins improve. That's the entire thesis of this role. Every kernel you tune, every quantization scheme you ship, every scheduler tweak you land shows up directly in a customer's p99 and on our P&L.
This is a high-impact seat. It is also a high-autonomy seat as you'll be given the room to lead the technical direction of inference optimization at Kimchi, not execute someone else's roadmap.

The problem: running LLMs in production is a moving target. The "right" model and serving configuration for a workload depend on traffic shape, sequence-length distribution, batch dynamics, GPU SKU, memory bandwidth, quantization tolerance, and a dozen other variables that shift week to week. Most teams pick a model once, over-provision GPUs, and absorb the cost. Kimchi is the system that makes that decision automatically - continuously matching workloads to the most cost-efficient, best-performing LLM and serving configuration on a customer's infrastructure. We're building the optimization layer between the model and the hardware, and we need engineers who understand both sides deeply.

Stack

Python; vLLM; SGLang; TensorRT-LLM; PyTorch; CUDA-adjacent tooling; Kubernetes; gRP; ClickHouse; PostgreSQL; GCP Pub/Sub; AWS / GCP / Azure; GitLab CI; ArgoCD; Prometheus; Grafana; Loki; Tempo.

Requirements:

  • 5+ years building real ML systems, with a portfolio that shows depth in inference or training infrastructure (not just model training notebooks).
  • Strong Python - production services, not scripts.
  • Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM, and a working mental model of why an inference engine performs the way it does on a given GPU.
  • Fluency with quantization tradeoffs - you've measured quality regressions, not just compression ratios.
  • Comfort with distributed systems: collective communication, sharding strategies, and the practical failure modes of multi-GPU and multi-node setups.
  • A bias toward measurement. You instrument before you optimize, and you can tell the difference between a real win and a benchmark artifact.
  • Self-direction. This role comes with a wide mandate; you should be excited by that, not unsettled by it.

Responsibilities:

  • Push throughput. Continuous batching, speculative decoding, chunked prefill, kernel-level tuning across vLLM, SGLang, and TensorRT-LLM. Find the ceiling on each GPU SKU, then raise it.
  • Cut latency. Attack TTFT and TPOT separately. Profile, identify the actual bottleneck (compute, memory bandwidth, scheduling, networking), and fix it - not the bottleneck you assumed.
  • Get more out of the KV cache. Paged attention, prefix caching, eviction policies, cache reuse across requests, quantized KV. This is where a lot of the unrealized throughput lives, and it's an area you'll own.
  • Quantize without regressing quality. INT8, INT4, FP8 across weights, activations, and KV. Empirical work: measure quality on real workloads, not just perplexity benchmarks.
  • Shrink cold starts and memory footprint. Faster init, smarter weight loading, tighter memory accounting - the difference between a model that scales and one that doesn't.
  • Scale across nodes. Distributed inference topologies, network-aware placement, checkpointing strategies that don't bottleneck on storage or interconnect.
  • Set the technical direction. Decide what we benchmark, what we adopt, and what we build ourselves. Bring the team along with strong writeups and reproducible experiments.

What’s in it for you?

  • Competitive salary (depending on the level of experience).
  • Enjoy a flexible, remote-first global environment.
  • Collaborate with a global team of cloud experts and innovators, passionate about pushing the boundaries of Kubernetes technology
  • Equity options.
  • Get quick feedback with a fast-paced workflow. Most feature projects are completed in 1 to 4 weeks.
  • Spend 10% of your work time on personal projects or self-improvement. 
  • Learning budget for professional and personal development - including access to international conferences and courses that elevate your skills.
  • Annual hackathon to spark new ideas and strengthen team bonds.
  • Team-building budget and company events to connect with your colleagues.
  • Equipment budget to ensure you have everything you need.
  • Extra days off to help maintain a healthy work-life balance.

Hiring process

  • Screening call with Recruiter
  • Hiring Manager interview
  • Technical interview (system design)
  • Live coding
  • Culture Check interview with an executive

As part of our standard hiring process, we would like to inform you that a background check may be conducted at the final stage of recruitment through our third-party provider, Checkr.
Please note that Cast AI does not provide any form of visa sponsorship/work permit.

#LI-Remote

¿Listo para postularte en Cast AI?
Postúlate en Cast AI

Sobre Cast AI

Why Cast AI?

Cast AI is the leading Application Performance Automation (APA) platform, enabling customers to cut cloud costs, improve performance, and boost productivity – automatically.

Built originally for Kubernetes, Cast AI goes beyond cost and observability by delivering real-time, autonomous optimization across any cloud environment. The platform continuously analyzes workloads, rightsizes resources, and rebalances clusters without manual intervention, ensuring applications run faster, more reliably, and more efficiently.

Headquartered in Miami, Florida, Cast AI has employees in more than 32 countries worldwide and supports some of the world’s most innovative teams running their applications on all major cloud, hybrid, and on-premises environments. Over 2,100 companies already rely on Cast - from BMW and Akamai to Hugging Face and NielsenIQ.

What’s next? Backed by our $108M Series C, we’re doubling down on making APA the new standard for DevOps and MLOps, and everything in between.

Ver todos los empleos en Cast AI →

Empleos similares

Redwood Materials
Software Engineer - ML/Computer Vision (Battery Sorting)
Redwood Materials
⚡ Postúlate pronto McCarran, NV; San Francisco, C... Presencial $152,500–$287,500
● Nuevo 👁 Visto ✓ Postulado hace 47m
Stitch Fix
ML Platform Engineer
Stitch Fix
⚡ Postúlate pronto Remote, USA · restringido por ubicación $136,000–$167,000
● Nuevo 👁 Visto ✓ Postulado hace 1h
Fireworks AI
Applied Machine Learning Engineer, Singapore
Fireworks AI
⚡ Postúlate pronto Singapore, Singapore Presencial SGD 180,000–SGD 250,000
● Nuevo 👁 Visto ✓ Postulado hace 1h
Rackner
AI/ML Engineer — Generative AI Mission Systems
Rackner
⚡ Postúlate pronto Remote · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 1h
Fireworks AI
Applied Machine Learning Engineer, EMEA
Fireworks AI
⚡ Postúlate pronto London, UK Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1h
Coupang
Staff ML Infra Engineer, Search & Discovery
Coupang
⚡ Postúlate pronto Mountain View, USA Presencial
● Nuevo 👁 Visto ✓ Postulado hace 2h
CI
Staff Machine Learning Engineer, Ads
Coupang Internal
⚡ Postúlate pronto Mountain View, USA Presencial
● Nuevo 👁 Visto ✓ Postulado hace 2h
Coupang
Staff Machine Learning Engineer, Ads
Coupang
⚡ Postúlate pronto Mountain View, USA Presencial $164,000–$164,000
● Nuevo 👁 Visto ✓ Postulado hace 2h
XPENG
Senior Machine Learning Data Curation Engineer
XPENG
⚡ Postúlate pronto Santa Clara, CA Presencial $174,720–$295,680
● Nuevo 👁 Visto ✓ Postulado hace 3h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Cast AI

Ver todos los empleos en Cast AI →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis