Jobs Companies LM Studio Software Engineer, Inference Runtime

About this Software Engineer, Inference Runtime role at LM Studio

LM Studio · Hybrid · New York City

LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family.

As a team, we work with high technical intensity and personal responsibility. We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software.

The Role

We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on.

Qualifications

  • Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure

  • Strong programming ability in Python and C++

  • Deep understanding of transformer architectures and the mechanics of model inference

  • Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement

  • Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

  • Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution

  • Takes personal responsibility for the correctness and performance of their work

Bonus Qualifications

  • Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

Responsibilities

  • Maintain and push forward our inference stack on-device and in the cloud

  • Bring up new model architectures and multimodal models

  • Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes

  • Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution

  • Benchmark and diagnose correctness and performance problems across the inference stack

  • Contribute upstream to open-source projects such as llama.cpp and MLX

Benefits

  • Competitive salary and equity grants

  • Great medical, vision, dental healthcare plans

  • Catered team lunch / expensed dinners in the office

  • Flexible PTO

  • Flexible WFH

  • Sun-drenched office in SoHo in NYC

Ready to apply to LM Studio?
Apply to LM Studio

How this Software Engineer salary compares

This role pays $250,000/yrin line with the typical range for Software Engineer roles.

$162,450 median $230,000 $402,500

Typical range $193,997–$301,250/yr, from 454 comparable Software Engineer listings on JobsRadar (pay annualized to USD). See Software Engineer salary insights →

Similar jobs

Nectar
Staff Software Engineer
Nectar
⚡ Apply early New York City, NY - Remote · location restricted $240,000–$400,000
● New 👁 Seen ✓ Applied 2h ago
VE
Software Engineer, Agent
Vercel
⚡ Apply early Hybrid - New York City Hybrid $232,000–$348,000
● New 👁 Seen ✓ Applied 5h ago
VE
Software Engineer - Next.js
Vercel
⚡ Apply early Hybrid - San Francisco, New Yo... Hybrid $196,000–$294,000
● New 👁 Seen ✓ Applied 5h ago
SA
Software Engineer — Security
Snorkel AI
⚡ Apply early New York City, NY (Hybrid); Sa... Hybrid $220,000–$300,000
● New 👁 Seen ✓ Applied 5h ago
VE
Software Engineer, Trust & Safety
Vercel
⚡ Apply early Hybrid - San Francisco, New Yo... Hybrid $196,000–$294,000
● New 👁 Seen ✓ Applied 6h ago
Forge Global
Senior Software Engineer (Tech Lead), Customer Domain Engineering
Forge Global
⚡ Apply early San Francisco, California, Uni... Hybrid
● New 👁 Seen ✓ Applied 9h ago
Forge Global
Senior Software Engineer (Tech Lead), Customer Domain Engineering
Forge Global
⚡ Apply early New York, New York, United Sta... Hybrid $209,000–$240,000
● New 👁 Seen ✓ Applied 9h ago
WRITER
Software engineer, generative AI
WRITER
⚡ Apply early San Francisco, CA Hybrid $132,000–$240,000
● New 👁 Seen ✓ Applied 9h ago
WRITER
Software engineer, connectors & MCP
WRITER
⚡ Apply early San Francisco, CA Hybrid $132,000–$240,000
● New 👁 Seen ✓ Applied 9h ago

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at LM Studio

See all jobs at LM Studio →

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free