About this Sr. Data Engineer - (RealTime Streaming) role at GSSTech Group
We are looking for an experienced Senior Data Engineer with strong expertise in real-time streaming technologies and large-scale data engineering solutions. The ideal candidate will have hands-on experience designing, building, and managing highly scalable and fault-tolerant streaming data platforms using Kafka, Flink, Java, and PySpark.
The candidate will be responsible for developing high-throughput, low-latency data pipelines, managing Kafka clusters, implementing cloud-native data solutions, and ensuring system reliability, security, and scalability in enterprise environments.
Requirements
Key Responsibilities
- Design, develop, implement, and manage Kafka-based real-time streaming architectures capable of handling high-volume and low-latency workloads.
- Build and maintain scalable streaming data pipelines using Kafka ecosystem components, including:
- Kafka Connect
- ksqlDB
- Schema Registry
- Develop and optimize real-time data processing applications using:
- Apache Flink
- Java
- PySpark
- Perform Kafka cluster setup, administration, configuration, tuning, monitoring, and performance optimization.
- Ensure Kafka clusters are highly available and implement disaster recovery strategies and best practices.
- Design and implement fault-tolerant, scalable, and resilient streaming systems for enterprise-grade applications.
- Work with cloud-native applications and modern data engineering frameworks to build scalable solutions.
- Automate infrastructure provisioning and deployment using Infrastructure-as-Code (IaC) tools such as Terraform.
- Implement and follow GitOps practices for deployment automation and infrastructure management.
- Apply robust security standards and best practices, including:
- SSL / mTLS
- SASL authentication
- ACL management
- Collaborate with cross-functional teams to deliver real-time data engineering solutions aligned with business and analytics requirements.
Required Skills & Technologies
Mandatory Skills
- Apache Kafka
- Apache Flink
- Java
- PySpark
- Real-Time Streaming Architecture
- Kafka Cluster Management
- Kafka Connect
- ksqlDB
- Schema Registry
- Streaming Data Pipelines
- Cloud-Native Applications
- Terraform
- Infrastructure-as-Code (IaC)
- GitOps
- High Availability & Disaster Recovery
- Performance Tuning & Optimization
- Fault Tolerance & Scalability
- SSL / mTLS
- SASL
- ACLs
Preferred Experience
- Strong experience in building enterprise-grade real-time data platforms.
- Hands-on experience working in high-throughput and low-latency environments.
- Experience with large-scale distributed data processing systems.
- Exposure to cloud platforms and modern DevOps practices will be an added advantage.