Jobs Companies Gimlet Labs Member of Technical Staff - Infrastructure

Sobre esta vaga de Member of Technical Staff - Infrastructure na Gimlet Labs

Gimlet Labs · Presencial · San Francisco, CA

About Us

Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them.

The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.

Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.

We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.

About this Role

We are looking for an Infrastructure platform Engineer to design, build, and operate the cluster infrastructure behind Gimlet’s heterogeneous inference cloud. Unlike traditional cloud platforms built around a single hardware ecosystem, Gimlet's infrastructure spans multiple accelerator vendors and architectures. Infrastructure engineers play a key role in bringing new hardware platforms online, building the operational abstractions that make heterogeneous infrastructure manageable at scale, and ensuring new silicon can serve production workloads reliably from day one.

This role is highly hands-on. You will work across bare metal, Linux, Kubernetes or cluster schedulers, high-speed networking, observability, provisioning, and incident response. You will partner closely with distributed systems, runtime, compiler, and hardware teams to ensure Gimlet’s infrastructure can support demanding AI workloads at production scale.

What you will work on

  • Design, deploy, and operate large-scale CPU, GPU, and accelerator clusters powering production AI inference.

  • Build automation for provisioning, configuration, upgrades, validation, and lifecycle management.

  • Design and scale provisioning systems for heterogeneous bare-metal infrastructure across multiple datacenters and hardware vendors.Operate cluster scheduling, resource allocation, isolation, quotas, and utilization systems.

  • Debug complex production issues across Linux, networking, storage, drivers, firmware, and orchestration layers.

  • Build and operate high-performance networking infrastructure, including RDMA-enabled environments and accelerator interconnects.

  • Build observability for cluster health, capacity, performance, failures, and workload behavior.

  • Improve reliability, availability, and recovery across multi-node production systems.

  • Work with distributed systems and runtime teams to support low-latency, high-throughput inference workloads.

  • Evaluate and integrate new hardware platforms, accelerators, networking technologies, and datacenter designs.

  • Create runbooks, operational standards, and incident response practices as the fleet scales.

You may be a good fit if

  • Experience in infrastructure, cluster engineering, platform engineering, SRE, HPC, or distributed systems.

  • Deep Linux systems experience, including debugging performance, networking, storage, processes, and kernel-level issues.

  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration and scheduling systems.

  • Strong automation skills using tools such as Terraform, Ansible, Helm, Python, Go, or equivalent.

  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA/ROCm stacks, or hardware validation.

  • Familiarity with high-performance networking such as InfiniBand, RoCE, high-speed Ethernet, or datacenter fabrics.

  • Strong operational judgment: you know how to build systems that are observable, recoverable, and boring in production.

  • Comfort working in a fast-moving startup environment with high ownership and ambiguity.

Strong candidates may also have

  • Experience building or operating AI inference, training, HPC, or neocloud infrastructure.

  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up.

  • Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting.

  • Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks.

  • Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana, or similar tooling.

  • Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators.

Pronto para se candidatar à Gimlet Labs?
Candidatar-se à Gimlet Labs

Vagas semelhantes

GL
Senior Technical Project Manager
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $240,000–$330,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
Head of Security and Compliance
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $270,000–$330,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
General Counsel
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $275,000–$335,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
Member of Technical Staff - Compilers
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $180,000–$400,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
Member of Technical Staff - Distributed Systems
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $150,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
Member of Technical Staff - AI Research
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $150,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
Member of Technical Staff - ML Systems & Inference
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $150,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
Network Engineer
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $250,000–$320,000
● Nova 👁 Vista ✓ Candidatada há 12h
GL
Member of Technical Staff - Kernels & GPU Performance
Gimlet Labs
⚡ Candidate-se cedo San Francisco, CA Presencial $150,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 12h

Cadastre-se para receber sugestões sob medida com base nas vagas que você abre e nas buscas que você salva.

Mais vagas na Gimlet Labs

Ver todas as vagas na Gimlet Labs →

Candidatar-se agora
🤖

Opa — calma aí

A JobsRadar foi feita para pessoas de verdade passando por um momento difícil na busca por emprego — não para requisições automatizadas. Você está clicando rápido demais e agora está temporariamente bloqueado.

Volte mais tarde. Se você está mesmo procurando emprego, estamos com você — apenas aja como um ser humano.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Ganhe vantagem na sua busca por emprego.

Entre no nosso canal do Telegram para o que ajuda você a conseguir a vaga — referências salariais, o pulso semanal do mercado e avisos de novos recursos. Sem spam, só sinal.

Entre no canal — é grátis