Jobs Companies Inferact Member of Technical Staff, Exceptional Generalist (Remote)

Über diese Member of Technical Staff, Exceptional Generalist (Remote) Stelle bei Inferact

Inferact · 🌍 Weltweit remote · Remote

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

This is a globally remote opportunity. We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems. This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance.

You'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure. You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after. The work you do will directly impact how the world runs AI inference.

Potential focus areas include:

  • Inference Runtime: Push the boundaries of LLM and diffusion model serving. Work at the core of vLLM to optimize how models execute across diverse hardware and architectures.

  • Kernel Engineering: Write the low-level kernels and optimizations that make vLLM the fastest inference engine in the world, running on hundreds of accelerator types.

  • Performance & Scale: Build the distributed systems that power inference at global scale—design foundational layers enabling vLLM to serve models across thousands of accelerators with minimal latency.

  • Cloud Orchestration: Build the operational backbone for cluster management, deployment automation, and production monitoring that enables teams worldwide to serve AI models without friction.

What We're Looking For

We're looking for engineers who thrive with autonomy. You should be able to take a vague problem statement and turn it into shipped code with minimal supervision. You communicate proactively, over-communicate context across time zones, and know when to ask for help versus when to push forward independently.

Core Requirements:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar

  • Demonstrated ability to work autonomously and drive projects to completion without close supervision

  • Excellent asynchronous communication skills and ability to collaborate effectively across time zones

  • Strong track record of shipping high-impact work in complex technical environments

  • Deep expertise in at least one of: systems programming, GPU/accelerator programming, distributed systems, or ML infrastructure

Technical Depth (strong in at least two):

  • CUDA kernels or equivalent (Triton, TileLang, Pallas) with deep understanding of GPU architecture

  • High-performance distributed systems in Rust, Go, or C++

  • Python with PyTorch internals and LLM inference systems (vLLM, TensorRT-LLM, SGLang)

  • Kubernetes, container orchestration, and infrastructure-as-code at scale

  • Transformer architectures, KV-cache memory management, and model serving

Preferred Qualifications:

  • Contributions to vLLM or other major open-source ML/systems projects

  • Experience with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel)

  • Knowledge of quantization techniques, ML-specific kernel optimization, or compiler technologies

  • Track record of improving system reliability and performance at scale

  • Written widely-shared technical blogs or impactful side projects in the ML infrastructure space

Logistics

  • Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.

  • Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.

Bereit, sich bei Inferact zu bewerben?
Bei Inferact bewerben

Ähnliche Jobs

Inferact
Product Marketing Manager
Inferact
⚡ Früh bewerben San Francisco Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Tg.
Inferact
HR & Recruiting Lead, APAC
Inferact
⚡ Früh bewerben Singapore Vor Ort SGD 150,000–SGD 300,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Inferact
Member of Technical Staff, CI/CD Infrastructure
Inferact
⚡ Früh bewerben San Francisco Vor Ort $200,000–$400,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Wo.
Inferact
Member of Technical Staff, AMD GPU Performance Engineering
Inferact
⚡ Früh bewerben Singapore Vor Ort SGD 200,000–SGD 400,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
Inferact
Member of Technical Staff, TPU Performance Engineering
Inferact
⚡ Früh bewerben San Francisco Vor Ort $200,000–$400,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
Inferact
Member of Technical Staff, AMD GPU Performance Engineering
Inferact
⚡ Früh bewerben San Francisco Vor Ort $200,000–$400,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
Inferact
Member of Technical Staff, TPU Performance Engineering
Inferact
⚡ Früh bewerben Singapore Vor Ort SGD 200,000–SGD 400,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
Inferact
Member of Technical Staff, Performance and Scale
Inferact
⚡ Früh bewerben Singapore Vor Ort SGD 200,000–SGD 400,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
Inferact
Member of Technical Staff, Kernel Engineering
Inferact
⚡ Früh bewerben Singapore Vor Ort $200,000–$400,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Inferact

Alle Jobs bei Inferact ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos