Sobre este puesto de Data Engineer, Data Cloud International en Zeta Global
The Role
Zeta Global is seeking a Data Engineer to help accelerate the operationalization of Zeta’s Data Cloud across international markets.
Sitting within the Data Cloud International / commercial data partnerships function, this position will work closely with Business Operations, Applications, Product, Engineering, Compliance, and external partner technical teams.
The focus is practical data engineering: turning high-potential data partnerships into usable, repeatable, revenue-generating Data Cloud assets that support product innovation, market expansion, and commercial growth through disciplined evaluation, ingestion, transformation, validation, and documentation.
You will work primarily across AWS-based data workflows, using tools such as S3, Athena, Glue, SQL, Python, APIs, and orchestration frameworks such as Airflow. This is a hands-on, delivery-focused position suited to someone who enjoys turning technical data requirements into practical workflows.
You will also use AI-assisted tools and automation techniques to improve data evaluation, documentation, workflow generation, and internal productivity.
The position is based in Copenhagen, Denmark. Candidates should be able to commute regularly to the Copenhagen office, located in Copenhagen or near Copenhagen Central Station.
Roles & Responsibilities
• Build and automate data workflows for new and existing data partnerships.
• Create repeatable ingestion, transformation, and replication processes for partner data feeds.
• Evaluate incoming datasets for structure, usability, coverage, completeness, and quality.
• Write scripts to transform, normalize, move, and prepare data for analysis or downstream use.
• Use SQL and Athena to query large datasets, validate outputs, and generate reporting tables or extracts.
• Build lightweight validation checks for files, schemas, counts, formats, and expected values.
• Produce clear technical documentation, including field mappings, data dictionaries, process notes, data flow summaries, and partner integration documentation.
• Use AI-assisted tools where appropriate to accelerate data investigation, documentation, code generation, workflow prototyping, and repeatable analysis.
• Communicate with technical contacts at data partners via email and occasional technical calls.
• Help translate partner data delivery requirements into practical ingestion and automation workflows.
• Work with internal teams across Operations, Product, Engineering, Compliance, and Data Cloud.
• Contribute to reusable templates, naming conventions, and documentation standards for partner data onboarding.
Required Qualifications
• Strong written and spoken English is required, as the role works across international teams and external data partners.
• 3–5 years of relevant experience, or equivalent practical experience, in data engineering, analytics engineering, technical data operations, cloud data automation, or a similar role.
• Practical experience writing SQL to query, validate, and transform data.
• Hands-on experience with Python for scripting, automation, or data manipulation.
• Familiarity with AWS data services, especially S3 and Athena; experience with Glue is strongly preferred.
• Understanding of ETL/ELT concepts and how data moves between systems.
• Experience working with structured and semi-structured data formats, including large delimited files, JSON, Parquet, or ORC.
• Comfortable reading technical documentation and working with API-based data sources.
• Ability to work with large datasets and investigate issues in schemas, counts, formats, or transformation outputs.
• Comfortable collaborating with internal technical teams and external partner technical contacts.
• Able to work independently on defined tasks while escalating ambiguity, blockers, or risks appropriately.
Required Skills
• Business-level English communication
• SQL
• Python
• AWS S3, Athena, and Glue / Glue Data Catalog
• ETL/ELT workflows
• Data ingestion and replication
• APIs and file-based data exchange
• Airflow, Prefect, or similar orchestration tools
• Snowflake, Databricks, Redshift, Hive, Presto, or similar data platforms
• Git or similar version control
• Data validation and quality checks
• Data dictionaries, field mappings, and technical documentation
• Structured and semi-structured data formats, including JSON, Parquet, ORC, and large delimited files
• Working knowledge of S3 policies, IAM permissions, and secure data access patterns
• Practical use of AI-assisted development, documentation, or data analysis tools