Über diese DevOps / Platform Engineer Stelle bei Yassir
Context & Vision
In 2026, writing code is no longer the primary bottleneck; managing its complexity and ensuring its reliability is. We are building a highly resilient advertising platform with a very lean, senior internal team, and we intend to keep it that way.
To achieve scale without the overhead of a large engineering department, we rely on an AI-first development paradigm and a philosophy borrowed from the best large-scale open-source projects. Our code is not public, but our governance is theirs: asynchronous communication, exhaustive written documentation, explicit rules, and uncompromising quality gates. It is the only way a small team of humans and AI agents ships serious systems without accumulating technical debt.
This role owns the ground the whole platform runs on: the clusters, the pipelines, and the guarantees that everything ships predictably.
The Role
You will own the platform and delivery infrastructure end to end — the clusters, the infrastructure-as-code, the delivery pipelines, and the operational guarantees behind them. In a lean team, reliability is not a separate department; it is a discipline you carry for everyone.
1. Infrastructure as Code. Own the cloud footprint through infrastructure-as-code and a GitOps workflow. Infrastructure is declared, reviewed and versioned like any other code — no click-ops, no undocumented state.
2. Delivery Pipelines.Own CI/CD. Builds are reproducible, deployments are predictable, and rollbacks are boring. You make shipping a non-event.
3. Reliability, Performance & DR. Own observability (metrics, logs, distributed traces), performance testing, and backup / disaster-recovery. You define the SLOs that matter for a real-time serving platform and you make them measurable.
4. Quality Gates & AI-First CI. Enforce quality at every passage point in the pipeline. Integrating AI into CI/CD — automated reviews, security and policy checks — is an open frontier here, and we expect you to study it and propose implementations. Everything you build is documented; if it is not written down, it does not exist.
The Tech Stack
- Orchestration: managed Kubernetes on Google Cloud Platform.
- IaC & GitOps: infrastructure-as-code and a GitOps workflow.
- CI/CD: modern delivery pipelines (and proposals to evolve them).
- Observability: metrics, logs, distributed tracing.
- Runtime context: Rust services, a React frontend, streaming, relational databases.
- Cloud:Google Cloud Platform.
Profile & Requirements
We are looking for a platform engineer who treats infrastructure as a product, owns reliability for the whole team, and is comfortable in a lean, high-quality, AI-first environment. This role is not suited to someone who wants to run a ticket queue inside a large ops team.
Essential Experience
- 4+ years in DevOps / platform / SRE roles, running production Kubernetes** for real workloads (GCP preferred).
- Infrastructure-as-code and a GitOps workflow in production; infrastructure declared and reviewed as code.
- Solid CI/CD ownership and a real observability practice (metrics, logs, distributed tracing)
- Experience defining and defending SLOs, performance and disaster-recovery for latency-sensitive systems.
Core Competencies & Mindset
- AI Development Lifecycle: comfort integrating AI into the delivery workflow, and a point of view on automating quality and security gates in CI.
- Uncompromising Reliability: a maniacal focus on predictability, recoverability and security. Surprises in production are the enemy.
- Written & Asynchronous: exceptional written communication; runbooks and ADRs are part of the job, not a favour. Effective in a distributed, async environment (Paris timezone +/- 3h).
- Ownership:** you carry reliability for the whole team and raise risks before they become incidents.