About this Data Engineer - ETL/PySpark (Banking Domain) role at GSSTech Group
Role Summary
We are looking for a hands-on Data Engineer with strong PySpark and Python expertise to build and maintain data marts and ETL pipelines within a banking environment. The ideal candidate owns the full SDLC, from build through UAT, bug fixing, production deployment, and postproduction support.
Key Responsibilities
- Design, build, and maintain ETL pipelines and data marts using PySpark and Python
- Write clean, maintainable, and robust production-grade code
- Own end to end SDLC activities: build, UAT, UAT bug fixes, production deployment, and postproduction support
- Perform Oracle query analysis and PySpark code debugging
- Work across structured, semi structured, and unstructured data sources
- Apply software engineering best practices to production pipelines
- Support CI/CD processes and data testing/validation activities
Required Skills & Experience
- 5+ years commercial experience in a data-driven role
- Hands-on experience building data marts and ETL pipelines
- Expert level Python for ETL scripting
- Strong PySpark experience
- Analytical expertise in Oracle SQL and data analysis
- Understanding of software engineering concepts and best practices for production pipelines
- Banking client or banking domain knowledge
- Strong Data Warehousing fundamentals
- Familiarity with query languages and both SQL and NoSQL database technologies
- CI/CD exposure, including testing and validation of data pipelines
Daily Tech Stack
- Python
- Spark / PySpark
- Jupyter
- SQL and NoSQL DBMS
- Hadoop / MapReduce / Hive
- Pandas