Jobs Companies Featherless AI Machine Learning Engineer — Inference Optimization

Sobre este puesto de Machine Learning Engineer — Inference Optimization en Featherless AI

Featherless AI · Remote (world)

About the Role

We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users.

This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains.

What You’ll Do

  • Optimize inference latency, throughput, and cost for large-scale ML models in production

  • Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO)

  • Implement and tune techniques such as:

    • Quantization (fp16, bf16, int8, fp8)

    • KV-cache optimization & reuse

    • Speculative decoding, batching, and streaming

    • Model pruning or architectural simplifications for inference

  • Collaborate with research engineers to productionize new model architectures

  • Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks)

  • Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups

  • Improve system reliability, observability, and cost efficiency under real workloads

What We’re Looking For

  • Strong experience in ML inference optimization or high-performance ML systems

  • Solid understanding of deep learning internals (attention, memory layout, compute graphs)

  • Hands-on experience with PyTorch (or similar) and model deployment

  • Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations)

  • Experience scaling inference for real users (not just research benchmarks)

  • Comfortable working in fast-moving startup environments with ownership and ambiguity

Nice to Have

  • Experience with LLM or long-context model inference

  • Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton)

  • Experience optimizing across different hardware vendors

  • Open-source contributions in ML systems or inference tooling

  • Background in distributed systems or low-latency services

Why Join Us

  • Real ownership over performance-critical systems

  • Direct impact on product reliability and unit economics

  • Close collaboration with research, infra, and product

  • Competitive compensation + meaningful equity at Series A

  • A team that cares about engineering quality, not hype

¿Listo para postularte en Featherless AI?
Postúlate en Featherless AI

Empleos similares

Featherless AI
Machine Learning Engineer — AI Architecture Research
Featherless AI
⚡ Postúlate pronto Remote (world) · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 5 meses
Featherless AI
Machine Learning Engineer — Distillation
Featherless AI
⚡ Postúlate pronto Remote (world) · restringido por ubicación ⚠ 6+ meses
● Nuevo 👁 Visto ✓ Postulado hace 6 meses
Featherless AI
Machine Learning Engineer — Training Optimization
Featherless AI
⚡ Postúlate pronto Remote (world) · restringido por ubicación ⚠ 6+ meses
● Nuevo 👁 Visto ✓ Postulado hace 6 meses
Vibe
Senior Machine Learning Engineer - Identity & Audience
Vibe
⚡ Postúlate pronto Paris Remoto
● Nuevo 👁 Visto ✓ Postulado hace 3h
Paraform
Applied ML Engineer
Paraform
⚡ Postúlate pronto San Francisco Presencial $190,000–$320,000
● Nuevo 👁 Visto ✓ Postulado hace 3h
Redwood Materials
Software Engineer - ML/Computer Vision (Battery Sorting)
Redwood Materials
⚡ Postúlate pronto McCarran, NV; San Francisco, C... Presencial $152,500–$287,500
● Nuevo 👁 Visto ✓ Postulado hace 5h
Stitch Fix
ML Platform Engineer
Stitch Fix
⚡ Postúlate pronto Remote, USA · restringido por ubicación $136,000–$167,000
● Nuevo 👁 Visto ✓ Postulado hace 5h
Roblox
Distinguished Machine Learning Engineer - Safety
Roblox
⚡ Postúlate pronto San Mateo, CA, United States Presencial $399,420–$457,970
● Nuevo 👁 Visto ✓ Postulado hace 5h
Roblox
Distinguished Engineer, Machine Learning Systems – Economy
Roblox
⚡ Postúlate pronto San Mateo, CA, United States Presencial $397,460–$455,720
● Nuevo 👁 Visto ✓ Postulado hace 5h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Featherless AI

Ver todos los empleos en Featherless AI →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis