Sobre este puesto de Data, Engineer en Master-Works
Master works is seeking a Data Engineer to join our team and play a key role in developing, operating, and maintaining the organization’s data pipelines and data platform infrastructure. The role will be responsible for ensuring reliable data processing, platform availability, security, performance, and compliance with pipeline SLAs.
Key Responsibilities
- Design, develop, and maintain ELT/ETL pipelines to produce clean, structured, and reliable datasets.
- Build and maintain data orchestration workflows using Apache Airflow.
- Develop and test data transformations using SQL, Python, and PySpark.
- Monitor pipeline execution, troubleshoot processing issues, and proactively address SLA risks.
- Administer and maintain lakehouse, object storage, Apache Spark, and Trino environments.
- Build and maintain CI/CD pipelines using Bitbucket and Jenkins.
- Manage user access, roles, permissions, and resource quotas across data platforms.
- Monitor platform performance, respond to incidents, and perform system upgrades and patching.
- Optimize storage configurations, including partitioning, data formats, and data retention.
- Collaborate with Data Integration, Data Engineering, BI, and Analytics Engineering teams to support data initiatives.
- Maintain infrastructure documentation, operational procedures, and technical runbooks.
- Participate in capacity planning, platform optimization, and cost-management initiatives.
- Provide coverage and support for Data Integration activities during team absences or periods of increased workload.
Requirements
Qualifications
- Bachelor’s degree in Computer Science, Information Systems, Software Engineering, or a related field.
- Minimum 5 years of relevant experience in Data Engineering or a related field.
- Strong proficiency in SQL, Python, PySpark, Apache Spark, Trino, and Apache Airflow.
- Hands-on experience with lakehouse architecture, object storage, and ELT/ETL development.
- Practical experience with Bitbucket, Jenkins, and CI/CD practices.
- Knowledge of access management, platform monitoring, incident response, and performance optimization.
- Experience working with enterprise data platforms and supporting production data environments.
Key Skills & Experience
- ELT/ETL Pipeline Development
- Apache Airflow Orchestration
- SQL, Python & PySpark
- Apache Spark & Trino
- Lakehouse Architecture & Object Storage
- CI/CD using Bitbucket & Jenkins
- Data Platform Administration
- Access, Role, Permission & Quota Management
- Pipeline Monitoring & Incident Resolution
- Performance & Storage Optimization
Location: Client Site