Über diese Data Engineer - DMP, Streaming & Analytics Stelle bei Yassir
Context & Vision
In 2026, writing code is no longer the primary bottleneck; managing its complexity and ensuring its reliability is. We are building a highly resilient advertising platform with a very lean, senior internal team, and we intend to keep it that way.
To achieve scale without the overhead of a large engineering department, we rely on an AI-first development paradigm and a philosophy borrowed from the best large-scale open-source projects. Our code is not public, but our governance is theirs: asynchronous communication, exhaustive written documentation, explicit rules, and uncompromising quality gates. It is the only way a small team of humans and AI agents ships serious systems without accumulating technical debt.
This role owns the data engine underneath the platform: the streams, segments and audiences that make advertising intelligent.
The Role
You will own the data and DMP layer end to end — the event streams flowing in, the segment and audience computation, and the analytical stores that feed targeting and reporting. What we hire for is judgment: the problems here are hard in the interesting way, and the interesting part is the tradeoffs you choose.
1. Hard problems, at real-time scale.Computing audiences and measurement over very large, high-velocity event streams, in near-real-time, where being right is not always the same as being exact and every choice has a real cost. We describe the problems; we expect you to bring the approach.
2. Stream-First Pipeline.Own the ingestion and processing pipeline. It is stream-first by design; you keep it that way rather than falling back to batch.
3. Segments & Audiences.Own how raw signals become targetable audiences — how they are computed, stored, refreshed, and exposed to the serving path. This is the intelligence layer of the product.
4. AI-First Delivery & Governance. We expect AI to write most of the code; you are the editor, not the typist. You specify work as explicit, self-sufficient issues, direct agentic tooling to implement, and gatekeep every diff. Schemas, data contracts and Architecture Decision Records are documented. If it is not written down, it does not exist.
The Tech Environment
- Data: event streaming and stream processing, relational and analytical stores.
- Language: the codebase is a compiled systems language; we care more about how you reason about data than about the language you arrive with.
- Domain: DMP, segments, audiences, identity, measurement.
- Cloud: Google Cloud Platform, Kubernetes.
Profile & Requirements
We are looking for a data-minded engineer who frames a problem before reaching for a tool, owns correctness, and thrives in a lean, high-quality environment. This role is not suited to someone who wants a narrow, hand-fed scope inside a large team.
Essential Experience
- 5+ years building data or backend systems in production, with genuine ownership of a pipeline's reliability and the correctness of its output.
- Depth on hard data problems— you have built things where scale, real time, or accuracy made the design non-obvious, and you can talk through the tradeoffs you chose and why.
- Streaming experience and a real grasp of stream-vs-batch tradeoffs, event modelling, and schema evolution.
- Comfort in a compiled systems language (Rust, Go, C++, Scala or similar). The codebase is Rust; the language is the tool, the data thinking is the skill.
- Working knowledge of an analytical data store and SQL. DMP / ad-data or audience/segmentation experience is a strong plus.
Core Competencies & Mindset
- Judgment over recipes: you frame a data problem and its constraints before choosing an approach, and you can say when an approximate answer is the right engineering call and when it is not.
- Distributed & complex systems: knowledge of, and a taste for, distributed and complex systems is very welcome. Tell us about the ones you have built or operated — what made them hard.
- AI Development Lifecycle: genuine proficiency editing and validating AI-produced code, and decomposing ambiguous goals into specifications an agent can execute unattended.
- Uncompromising Quality: a maniacal focus on data correctness, freshness, and reliability. A wrong number that ships is a wrong number the business trusts.
- Written & Asynchronous: exceptional written communication; effective in a distributed, async environment (Paris timezone +/- 3h).
- Ownership: you own your pipeline and its documentation completely, and you surface problems before they are asked about.