Jobs Companies Statista Data Engineer - Data Platform & Ontology (m/f/d)

Über diese Data Engineer - Data Platform & Ontology (m/f/d) Stelle bei Statista

Statista · Hybrid · Hamburg or Berlin

At Statista, we’re all about facts and data, for we are the world's leading business data platform. By providing reliable and easy-to-use data as well as various data analytics products and services, we empower people worldwide to make fact-based decisions.

Founded in Hamburg in 2007, we have quickly grown into a global company with offices in major cities such as London, New York, Berlin and Tokyo. And we still have a lot of plans. Our constant growth does not only prove our success, but also keeps creating new development and career opportunities for our employees.

We value and celebrate our diverse culture. You are welcome here for who you are, no matter where you come from, what you look like, or whether you prefer bar graphs to pie charts. Your story matters – keep writing it as part of our team.

Are you ready to join us?

About the Role

As Data Engineer on our Healthcare Platform, you will own the foundational data ingestion, entity-resolution, and platform infrastructure end-to-end. You will design and operate scalable batch/streaming data pipelines, harden platform orchestration, and ensure system reliability, security, and cost efficiency.

A central challenge of this role is building and operationalizing our unified healthcare data ontology—a consistent semantic model covering hospitals, departments, specialties, metrics, and classification standards. Working closely with Analytics Engineers, Data Scientists, and Methodology experts, you will turn heterogeneous, multi-country hospital data into a coherent, highly queryable data asset.

Key Responsibilities

  • Pipeline Infrastructure & Orchestration: Build, optimize, and operate reliable ELT pipelines (using Python, SQL, and Prefect/Airflow) to ingest data from heterogeneous international sources, APIs, databases, and lakehouse storage (S3, Apache Iceberg).

  • Healthcare Ontology & Entity Resolution: Drive the implementation of entity resolution and master data management (MDM) for international hospital entities, mapping raw source data to canonical structures and maintaining standardized vocabularies (e.g., ICD/OPS, specialty taxonomies).

  • Data Contracts & Schema Governance: Establish strict data contracts (Pydantic, dbt contracts) and schema management, ensuring dataset reproducibility, data lineage tracking, and automated validation across all platform pipelines.

  • Platform Efficiency & Cloud Infrastructure: Optimize data storage, query execution, and compute costs across AWS and Snowflake, keeping data assets performant, secure, and cost-effective.

  • Automation & CI/CD: Implement automated testing and deployment workflows for data pipelines using GitHub Actions and Infrastructure as Code (Terraform).

  • Cross-Functional Data Enablement: Partner directly with Analytics Engineers, Data Scientists, and domain experts to deliver documented, research-grade, and production-ready datasets.

Qualifications

Core Requirements (Must-Haves)

  • Data Ingestion & Pipeline Orchestration: Advanced Python and analytical SQL for complex data ingestion across diverse file formats, REST APIs, databases, and cloud lakes (S3/Iceberg). Hands-on experience with modern orchestrators (Prefect, Airflow, or Dagster).

  • Entity Resolution & Data Governance: Practical experience with entity resolution/record linkage frameworks (e.g., Splink, dedupe, recordlinkage) and schema management/data contracts (Pydantic, dbt contracts, or JSON Schema).

  • Cloud Platform & Warehouse Infrastructure: Deep hands-on experience in an AWS production environment (S3, ECS/EC2) combined with cloud data warehouses (Snowflake).

  • Automated CI/CD & Workflow Automation: Proven track record of automating data pipeline deployments, integration tests, and validation workflows via GitHub Actions.

Nice-to-Haves (What Will Make You Stand Out)

  • Knowledge Graphs & Healthcare Terminologies: Exposure to ontology/semantic frameworks (RDF/OWL, SKOS, Neo4j, LinkML) or international medical classifications/vocabularies (SNOMED CT, ICD/OPS, FHIR).

  • Metadata & Lineage Tooling: Experience operating metadata registries and lineage catalogs (e.g., OpenMetadata, DataHub, dbt docs).

  • Infrastructure as Code (IaC): Proficiency in using Terraform to declaratively manage cloud resources and environments.

Your Profile

  • Degree: Bachelor's or Master's in Computer Science, Data Science, Software Engineering, or a related quantitative field.

  • Experience: 3+ years in data engineering building production pipelines and data platforms; including a sustained period within one organization seeing a core platform or product through build → launch → iteration.

  • Domain Knowledge: Healthcare domain experience is a plus (basic understanding of healthcare KPIs, quality metrics, or benchmarking concepts; familiarity with hospital structures and medical classification systems like ICD/OPS is especially valuable).

  • Mindset: Strong analytical and systems mindset, with a proven ability to transform messy, heterogeneous international data into a clean, well-governed, and highly structured data asset.

  • Languages: Fluent in English, German is a plus.

  • Working Style: Highly structured, curious, detail-oriented, and motivated to collaborate closely with analytics engineers, data scientists, and methodology experts in an international environment.

What we offer

In addition to our great team, culture, and our shared goal of empowering people with data, there are many other things that make Statista a great place to work! Join us and benefit from:

  • Work from abroad up to 30 calendar days a year

  • Hybrid work and flex-time

  • International team and social events

  • Subsidized urban mobility and access to fitness and wellness options

  • Free access to Langdock and all its amazing functionalities

  • Career & training opportunities

  • Attractive locations and modern offices

  • Mental health support with OpenUp

Some of the benefits listed here apply only to the German entity and to Junior-level roles or above.

Bereit, sich bei Statista zu bewerben?
Bei Statista bewerben

Ähnliche Jobs

Alcemy
Backend Engineer - Data (m/f/d)
Alcemy
⚡ Früh bewerben Berlin, Berlin, Germany Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
praxipal
(Senior) Software Engineer - Data Integration Focus (m/f/d) Remote EU
praxipal
⚡ Früh bewerben Berlin · standortgebunden €70,000–€100,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
HelloFresh
Senior Data Engineer, Analytical Data Platform, Inteligent Platforms (all genders)
HelloFresh
⚡ Früh bewerben Berlin, Berlin, Germany Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
MA
Oliver Wyman - Forward Deployed Engineer — Data & Analytics
Marsh
⚡ Früh bewerben Newcastle - Bank Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Almedia
AI Engineer - Data
Almedia
⚡ Früh bewerben Berlin Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
LawZero
Senior Data Platform Engineer
LawZero
⚡ Früh bewerben Berlin; Montreal Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
SM
Senior Data Platform Engineer
Smartly
⚡ Früh bewerben Berlin, Berlin, Germany; Helsi... Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
SumUp
Data Platform Engineer
SumUp
⚡ Früh bewerben Berlin, Germany Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
A1
Senior Data Engineer (m/f/d)
A11
⚡ Früh bewerben Berlin Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Statista

Alle Jobs bei Statista ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos