Jobs Companies FuriosaAI Sr. Software Engineer - Inference Engine (Platform Software)

À propos de ce poste Sr. Software Engineer - Inference Engine (Platform Software) chez FuriosaAI

FuriosaAI · Sur site · Seoul HQ

About the Job

Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.

In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential.

Responsibilities

  • Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency.

  • Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine.

  • Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake.

  • Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.

  • Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine.

Minimum Qualifications

  • BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience

  • Proficiency in Rust or C++ programming skill

  • Knowledge and passion of deep learning, LLM, and/or generative AI models

  • Excellent problem-solving and data analysis skills.

  • Strong communication and collaboration skills.

Preferred Qualifications

  • Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing.

  • A deep understanding of performance optimization systems.

  • Proficiency in C++/CUDA or Triton kernel development

  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM.

Contact

Prêt à postuler chez FuriosaAI ?
Postuler chez FuriosaAI

Emplois similaires

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez FuriosaAI

Voir tous les emplois chez FuriosaAI →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit