About this Data Engineer role at Weekday AI
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ญ๐ด๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฎ๐ฏ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ญ๐ด-๐ฎ๐ฏ ๐๐ฃ๐)
Experience: 3+ yrs
Location: Mumbai, Maharashtra, India
Job Type: Full-time
We are looking for an experiencedย Data Engineerย to design, develop, and maintain scalable data platforms and pipelines usingย AWS, Apache Spark, Python, SQL, Kafka, and modern data engineering frameworks.
The role will focus on building reliable data solutions for high-volume batch and real-time workloads, including data ingestion, transformation, processing, orchestration, storage, and delivery. The ideal candidate will have strong hands-on experience with AWS data services and distributed data processing, along with a solid understanding of data modelling, streaming architectures, data quality, and pipeline optimisation.
Requirements
KEY RESPONSIBILITIES
- Design, develop, and maintain end-to-endย data pipelinesย for high-volume data ingestion, transformation, processing, and delivery.
- Build scalableย Spark-based ETL/ELT workflowsย for both batch and real-time data processing.
- Develop data ingestion solutions usingย Kafka, Amazon Kinesis, and other streaming technologies.
- Build and manage data lakes, warehouses, and lakehouse solutions usingย AWS S3, Glue, Redshift, Athena, and EMR.
- Develop efficient data models using dimensional modelling, star schemas, partitioning, and other data engineering practices.
- Implement data quality checks, validation rules, monitoring, and error-handling mechanisms.
- Develop automated workflows usingย Airflow, MWAA, AWS Step Functions, or similar orchestration tools.
- Collaborate with Data Analysts and Data Scientists to deliver clean, structured, and analytics-ready datasets.
- Optimise data pipelines forย performance, scalability, reliability, and AWS cost efficiency.
- Integrate data from multiple internal and external systems while maintaining data consistency and reliability.
- Develop Python-based automation and data-processing solutions.
- Monitor production pipelines, troubleshoot failures, and perform root-cause analysis.
- Follow modern software engineering practices includingย Git, CI/CD, testing, documentation, and code reviews.
- Contribute to data platform architecture, engineering standards, and continuous improvement initiatives.
- Support data governance, cataloguing, lineage, and metadata management practices where required.
- Work with Linux/Unix environments and efficiently process large datasets.
WHAT MAKES YOU A GREAT FIT
- 3+ years of professional experienceย in Data Engineering, Big Data, Analytics Engineering, or a related field.
- Strong hands-on programming experience withย Pythonย for data processing, automation, and pipeline development.
- Strong expertise inย Apache Spark, particularly PySpark and/or Spark SQL.
- Deep working knowledge of theย AWS data ecosystem, including S3, Glue, Redshift, Athena, EMR, Kinesis, Lambda, and IAM.
- Hands-on experience withย Kafka, Kinesis, Flink, or similar real-time streaming technologies.
- Strong command ofย SQLย and experience with data modelling, dimensional modelling, partitioning, and large-scale data processing.
- Experience withย Airflow, MWAA, Step Functions, or comparable workflow orchestration tools.
- Strong understanding of batch and real-time data processing architectures.
- Experience working with high-volume datasets and distributed data processing environments.
- Familiarity withย Git, CI/CD, testing, and modern software development practices.
- Comfortable working inย Linux/Unix environments.
- Strong troubleshooting, analytical, and problem-solving skills.
- Experience withย Delta Lake, Apache Iceberg, Hudi, or other lakehouse technologies is an advantage.
- Knowledge of data governance, cataloguing, metadata, and lineage tools such asย Glue Data Catalog, DataHub, or Amundsenย is a plus.
- Familiarity withย Docker, ECS, or EKSย and containerised deployments is desirable.
- Basic understanding ofย AI/ML data requirements and workflowsย is an advantage.
- Bachelor's degree inย Computer Science, Information Technology, Engineering, or a related technical disciplineis preferred.