Jobs Companies FuriosaAI Sr. Software Engineer - Inference Engine (Platform Software)

Über diese Sr. Software Engineer - Inference Engine (Platform Software) Stelle bei FuriosaAI

FuriosaAI · Vor Ort · Seoul HQ

About the Job

Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.

In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential.

Responsibilities

  • Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency.

  • Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine.

  • Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake.

  • Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.

  • Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine.

Minimum Qualifications

  • BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience

  • Proficiency in Rust or C++ programming skill

  • Knowledge and passion of deep learning, LLM, and/or generative AI models

  • Excellent problem-solving and data analysis skills.

  • Strong communication and collaboration skills.

Preferred Qualifications

  • Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing.

  • A deep understanding of performance optimization systems.

  • Proficiency in C++/CUDA or Triton kernel development

  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM.

Contact

Bereit, sich bei FuriosaAI zu bewerben?
Bei FuriosaAI bewerben

Ähnliche Jobs

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei FuriosaAI

Alle Jobs bei FuriosaAI ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos