Sobre esta vaga de Senior Systems Software Engineer (Linux & HPC) - m/f/d na Langdock
Where Europe's enterprises adopt AI
Langdock is the AI platform used by more than 10,000 companies to give employees secure access to leading AI models, build agents, and automate workflows. We have grown past $40M ARR while remaining a small team. We’re building the AI platform for European enterprises while preserving their control over data, model providers, and deployment environments.
As a Systems Software Engineer, you build the foundations behind our agent execution, model serving, and shared services. You investigate how systems behave under load and failure, then write the software that makes them faster, safer, and more reliable.
This role requires deep Linux knowledge and hands-on experience with high-performance computing, distributed systems, or low-latency networking. You should be comfortable following a problem through application code, runtimes, operating-system behavior, and hardware constraints.
You work closely with Platform Engineering: your focus is the underlying execution, isolation, scheduling, and data-handling mechanisms; Platform Engineering owns the surrounding service contracts and long-term lifecycle. You operate what you build and remain accountable for its behavior in production.
What you might work on
The systems agenda includes:
Secure agent execution. Build runtime and isolation mechanisms for untrusted code using Linux, containers, and microVMs. Improve startup, persistent filesystems, network restrictions, suspension, and resumption.
Inference and HPC. Improve model-serving throughput, latency, accelerator utilization, and cost. Measure bottlenecks across serving software, scheduling, caching, memory, and hardware.
Linux performance and resource scheduling. Investigate CPU, memory, filesystem, and network behavior. Improve resource allocation and workload isolation so one customer’s workload cannot degrade another’s.
Distributed execution. Keep concurrent, long-running work correct through retries, crashes, and recovery. Define explicit guarantees for ordering, idempotency, and tenant isolation.
Storage and data systems. Improve how customer and agent-generated data are stored, versioned, moved, and recovered while preserving consistency and predictable performance.
Usage metering. Build accurate accounting for model tokens and compute time despite concurrent reporting, delayed events, duplicate requests, and capacity reservations.
Tech stack
Go and TypeScript in one Bazel monorepo, with implementation choices driven by the system
Linux, containers, microVMs, filesystems, and networking
Protobuf and gRPC for service contracts
Kubernetes across GCP, AWS, Azure, and on-premises deployments
Terraform and Terragrunt for infrastructure orchestration
PostgreSQL and Redis where durable metadata or coordination requires them
Open-source model serving and accelerator infrastructure
You do not need prior experience with every item. You do need enough systems depth to enter an unfamiliar part of the stack, understand its behavior, and make consequential changes safely.
How we work
We operate with high trust and autonomy in squads of 3 to 4 engineers. A squad owns its roadmap, prioritization, technical decisions, and operation in production. Engineers are expected to find the context they need, ask for input when it improves the outcome, and move work forward without waiting for every next step to be assigned.
We align asynchronously before scheduling a meeting. Product requirement documents (PRDs) define the user problem, intended outcome, and constraints. Design documents make architectural boundaries, tradeoffs, failure modes, migrations, and rollouts explicit. People read and challenge the thinking asynchronously; once the context is shared, a short in-office discussion or whiteboard session usually resolves the remaining questions quickly.
We optimize for leverage. Engineers choose the AI tools that work for them, supported by clear ticket context, focused branches, automated tests, and AI review before human review. We also invest in observability, migration tooling, automated recovery, and runbooks so recurring product maintenance does not depend on someone remembering a manual step.
The engineer who ships a change owns it in production. If something breaks, you lead the fix.
We’re looking for someone who
Has built systems software. You have personally designed and implemented difficult production systems in areas such as networking, storage, runtimes, databases, distributed computing, or inference.
Understands Linux deeply. You can reason about processes, scheduling, virtual memory, I/O, networking, and isolation, and investigate behavior down to the kernel when needed.
Measures before optimizing. You form hypotheses, use profiling and tracing, and can explain both the bottleneck and the measured effect of your changes.
Reasons carefully about failure. You understand concurrency, recovery, resource contention, and the tradeoffs between performance, reliability, and complexity.
Follows technical curiosity. You investigate unfamiliar systems independently and want to understand why they behave as they do.
Communicates clearly and takes ownership. You explain tradeoffs, make uncertainty visible, ship changes, and operate them in production.
Experience with eBPF, RDMA, confidential computing, accelerator infrastructure, or systems-level open-source projects is particularly relevant. You don’t need every specialty.
Your background might include cloud providers, networking companies, database teams, HPC, or trading infrastructure. We value demonstrated engineering depth, including through self-directed work and substantial open-source contributions. This is an experienced individual-contributor role; no particular degree is required.This is an experienced individual-contributor role. You should be able to independently investigate an unfamiliar problem, design a solution, ship it, and operate it in production. We generally expect this judgment to come from at least three to six six of full-time experience working on production systems.