Jobs › Companies › Clockwork.io › Senior HPC Developer - RDMA Networking

About this Senior HPC Developer - RDMA Networking role at Clockwork.io

Clockwork.io · Hybrid · On Site, Palo Alto, California

About Clockwork Systems

Clockwork.io – Software Driven Fabrics to increase GPU cluster utilization

Clockwork Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the foundations of distributed computing. As AI workloads grow increasingly complex, traditional infrastructure struggles to meet the demands of performance, reliability, and precise coordination. Clockwork is pioneering a software-driven approach to AI fabrics by delivering cross-stack observability to catch and quickly resolve problems, workload fault tolerance to keep jobs running through failures, and performance acceleration that dynamically routes and paces traffic to avoid congestion.

To learn more, visit www.clockwork.io.

About the Role

We’re looking for a Senior Systems Engineer – RDMA Networking & Storage to help build and optimize high-performance infrastructure for AI and HPC systems.
You’ll work across RDMA networking and storage data paths, spanning compute hosts, NICs, DPUs, smart switches, storage systems, and the Linux software stack. The work includes RoCE networking, NVMe-over-Fabrics, congestion control, performance optimization, programmable network hardware, and debugging complex interactions across hardware and software.
This role is ideal for an experienced systems engineer who enjoys working close to the hardware, diagnosing difficult performance problems, and building reliable systems where latency, throughput, and network behavior matter.

What You’ll Do
  • Design, build, and optimize high-performance RDMA networking and storage systems
  • Work with RoCE, RDMA Verbs, and NVMe-over-Fabrics
  • Develop systems that leverage DPUs, SmartNICs, and programmable switches for packet processing, telemetry, control, and data-path acceleration
  • Develop and optimize software using queue pairs, completion queues, memory registration, zero-copy I/O, and other RDMA primitives
  • Debug performance and reliability issues across applications, Linux kernel, drivers, NICs, DPUs, switches, and storage systems
  • Analyze network and storage behavior including latency, throughput, queueing, congestion, packet loss, retransmissions, and flow control
  • Develop mechanisms to improve network performance and congestion behavior in high-bandwidth RDMA fabrics
  • Build tooling and instrumentation to understand end-to-end performance across compute, network, and storage
  • Profile and tune systems across PCIe, NUMA, CPU, memory, NIC, DPU, network, and storage boundaries
  • Own major components from design through implementation, performance validation, and deployment
What You Bring
  • 5+ years of experience in systems, networking, storage, HPC, or other performance-critical software development
  • Strong proficiency in low-level C/C++
  • Strong understanding of RDMA networking, including RoCE and RDMA Verbs
  • Strong understanding of networking fundamentals, including packet processing, queues, congestion, and flow control
  • Experience debugging low-level Linux networking, storage, or driver behavior
  • Familiarity with concepts such as DMA, memory registration, queue pairs, completion queues, interrupts/polling, and zero-copy I/O
  • Experience profiling and optimizing latency- and throughput-sensitive systems
  • Ability to reason across hardware and software boundaries and independently debug complex system-level problems
  • Strong ownership and comfort working in a small, fast-moving engineering team
Bonus Points
  • Experience with DPUs, SmartNICs, or programmable network hardware
  • Experience with platforms such as NVIDIA BlueField or programmable switch environments
  • Experience with P4, switch SDKs, or in-network packet processing
  • Experience with NVMe, NVMe-oF, NVMe/RDMA, or NVMe/TCP
  • Experience with SPDK, DPDK, or other userspace I/O frameworks
  • Experience with Linux NVMe and RDMA subsystems
  • Experience with congestion control, including DCQCN, ECN, PFC, or related mechanisms
  • Experience tuning NVIDIA/Mellanox ConnectX NICs and RoCE fabrics
  • Experience working with high-performance switches and data-center network fabrics
  • Experience with PCIe, NUMA, huge pages, DMA, or CPU affinity optimization
  • Experience with performance profiling, tracing, or observability tooling
  • Background in AI infrastructure, HPC clusters, distributed storage, or large-scale data-center systems
Enjoy
  • Challenging, technically deep projects
  • A friendly and inclusive workplace culture
  • Competitive compensation
  • A great benefits package
  • Catered lunch
Compensation for this position will vary based on the skills and experience you bring, as well as internal equity considerations. For candidates hired at the posted level, the expected base salary range is $150,000–$230,000. The offered compensation package may also include stock options or other equity awards, subject to Clockwork’s equity program and applicable approvals.

Clockwork Systems is an equal opportunity employer. We are committed to building world-class teams by welcoming bright, passionate individuals from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, religion, age, sex, sexual orientation, gender identity or expression, national origin, disability, or protected veteran status. We believe diversity drives innovation, and we grow stronger together.

Ready to apply to Clockwork.io?
Apply to Clockwork.io

Similar jobs

Clockwork.io
Senior Human Resources Manager
Clockwork.io
⚡ Apply early Palo Alto, CA Onsite $140,000–$180,000
● New 👁 Seen ✓ Applied 1w ago
Clockwork.io
Software Engineer
Clockwork.io
⚡ Apply early On Site, Palo Alto, California Hybrid $140,000–$200,000
● New 👁 Seen ✓ Applied 2w ago
Clockwork.io
Software Engineer Intern
Clockwork.io
⚡ Apply early Palo Alto, CA Onsite
● New 👁 Seen ✓ Applied 2w ago
Clockwork.io
Senior Software Engineer – Network Observability
Clockwork.io
⚡ Apply early Onsite Palo Alto, California Hybrid $140,000–$210,000
● New 👁 Seen ✓ Applied 3w ago
Clockwork.io
Tech Lead - Full Stack
Clockwork.io
⚡ Apply early Onsite - Palo Alto, California Hybrid $180,000–$260,000
● New 👁 Seen ✓ Applied 1mo ago
Clockwork.io
Software Development Engineer in Test
Clockwork.io
⚡ Apply early On Site, Palo Alto, California Hybrid $130,000–$175,000
● New 👁 Seen ✓ Applied 1mo ago
Clockwork.io
Senior Software Engineer
Clockwork.io
⚡ Apply early On Site, Palo Alto, California Hybrid $140,000–$210,000
● New 👁 Seen ✓ Applied 1mo ago
Clockwork.io
Senior Account Executive - UK
Clockwork.io
⚡ Apply early London, England, United Kingdo... Onsite
● New 👁 Seen ✓ Applied 1mo ago
Clockwork.io
Senior Systems Engineer - AI Infrastructure
Clockwork.io
⚡ Apply early On Site, Palo Alto, California Hybrid $150,000–$230,000
● New 👁 Seen ✓ Applied 1mo ago

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at Clockwork.io

See all jobs at Clockwork.io →

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free