Jobs Companies Plaud Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

Sobre este puesto de Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco en Plaud

Plaud · Híbrido · San Francisco, CA

About Plaud Inc.

Plaud is building the real-world AI interface for professionals to amplify intelligence, elevate productivity and performance, loved by over 2,000,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud captures, structures, and compounds the intelligence generated in conversations — so humans can think better, decide faster, and execute with clarity.

 

Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human–AI intelligence through a hardware–software combination. With full ISO 27001, ISO 27701, SOC 2, GDPR, EN18031, and HIPAA compliances, Plaud is committed to the highest standards of data security and privacy protection.

To learn more about Plaud, please visit https://www.Plaud.ai and follow along on Instagram, X, Facebook, LinkedIn, and YouTube

 

Why You Should Join Us

Plaud is building the next generation intelligence infrastructure and interfaces to capture, extract, and utilize intelligence from what people say, hear, see, and think.

  • Plaud is a bootstrapped, skyrocketing, profitable company with a $250M revenue run rate achieved in just three years.

  • Define the next-gen paradigm for human-AI interaction.

  • Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion.

  • Work with passionate teammates who value innovation, collaboration, and customer success.

  • Grow your career in a culture that champions continuous learning and fast career development.

  • Market-competitive compensation, global exposure, and a vibrant, creativity-fueled work atmosphere.

 

You may be a good fit if you:

  • Have hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models.

  • Understand the intricate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments.

  • Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI.

  • Possess a deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and the memory hierarchy, allowing you to identify and eliminate hardware bottlenecks.

  • Communicate clearly and collaborate effectively, as you will sit at the critical intersection between the core ML training team and the backend infrastructure team.

  • Thrive in fast-moving environments and genuinely enjoy the systems-engineering challenge of squeezing every last drop of performance out of a cluster of GPUs.

  • Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity.

 

Strong candidates may also have experience with:

  • Frontier Serving Frameworks: Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories).

  • Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize conversational latency.

  • Advanced Inference Techniques: Implementing cutting-edge generation algorithms such as speculative decoding, lookahead decoding, or chunked prefill.

  • Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy.

  • Large-Scale Distributed Systems: Deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes.

 

What We Offer

  • Founding Team Initiative: Opportunity to be an early, foundational member of our core SpeechLLM lab, with meaningful ownership and impact on a fast-growing startup.

  • Competitive Compensation: $170K - $320K base salary + performance bonus + Equity.

  • Comprehensive Benefits: Top-tier healthcare for employees and dependents, including dental and vision, and a generous employer subsidy.

  • Retirement Planning: 401(k) plan for full-time employees with company matching.

  • Paid Time Off: Unlimited PTO, plus 13 paid holidays.

  • New Parent Leave: 12 weeks of paid time off to spend time with your new family, regardless of gender.

  • Hybrid Office: Minimum of 3x in-office per week to foster highly collaborative, fast-paced research.

  • Gear & Perks: Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office.

 

Plaud is and will continue to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristics.

¿Listo para postularte en Plaud?
Postúlate en Plaud

Cómo se compara este salario de ML Engineer

Este puesto paga $245,000/yren línea con el rango típico para los puestos de ML Engineer.

$170,000 la mediana de $252,000 $350,250

Rango típico $204,500–$296,200/yr, a partir de 93 ofertas comparables de ML Engineer en JobsRadar (salario anualizado en USD). Ver datos salariales de ML Engineer →

Empleos similares

Redwood Materials
Software Engineer - ML/Computer Vision (Battery Sorting)
Redwood Materials
⚡ Postúlate pronto McCarran, NV; San Francisco, C... Presencial $152,500–$287,500
● Nuevo 👁 Visto ✓ Postulado hace 3h
Stitch Fix
ML Platform Engineer
Stitch Fix
⚡ Postúlate pronto Remote, USA · restringido por ubicación $136,000–$167,000
● Nuevo 👁 Visto ✓ Postulado hace 3h
DoorDash USA
Software Engineer, Machine Learning Infrastructure - Generative AI
DoorDash USA
⚡ Postúlate pronto San Francisco, CA; Sunnyvale,... Presencial $137,100–$201,600
● Nuevo 👁 Visto ✓ Postulado hace 5h
DoorDash USA
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
DoorDash USA
⚡ Postúlate pronto San Francisco, CA; Sunnyvale,... Presencial $137,100–$201,600
● Nuevo 👁 Visto ✓ Postulado hace 5h
Nextdoor
Senior/Staff Machine Learning Engineer - Ads
Nextdoor
⚡ Postúlate pronto US Remote · restringido por ubicación $205,000–$355,000
● Nuevo 👁 Visto ✓ Postulado hace 7h
Discord
Senior Software Engineer, Machine Learning (Ads)
Discord
⚡ Postúlate pronto San Francisco Bay Area Presencial
● Nuevo 👁 Visto ✓ Postulado hace 16h
Figma
Software Engineer, Machine Learning
Figma
⚡ Postúlate pronto San Francisco, CA • New York,... Presencial $153,000–$376,000
● Nuevo 👁 Visto ✓ Postulado hace 20h
Liftoff
Software Engineer, ML Data
Liftoff
⚡ Postúlate pronto San Francisco Bay Area Presencial $180,000–$230,000
● Nuevo 👁 Visto ✓ Postulado hace 1d
Anthropic
Machine Learning Infrastructure Engineer, Safeguards Research
Anthropic
⚡ Postúlate pronto San Francisco, CA | New York C... Presencial $350,000–$500,000
● Nuevo 👁 Visto ✓ Postulado hace 1d

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Plaud

Ver todos los empleos en Plaud →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis