About this Data Engineer (Data Center) role at Coherent Solutions
Company Background
Our client is a leading global data center provider delivering hyperscale and edge infrastructure solutions across the Americas, EMEA, and Asia-Pacific. With 80+ data centers in 20+ countries, they partner with industry leaders such as Google, Oracle, NVIDIA, and Microsoft Azure to power the world’s digital infrastructure. Recognized as a USA TODAY Top Workplace for four consecutive years, the company continues to expand its global footprint and customer ecosystem.
Project Description
The project focuses on building a cloud-based data platform for processing high-volume sensor and telemetry data. The specialist will help develop and operate real-time and analytical data capabilities that support dashboards, applications, and downstream data consumers, while contributing to the platform’s migration and evolution in AWS.
Technologies
- Apache Kafka, Apache Flink
- Java, Python, SQL
- Apache Iceberg / Delta Lake / Hudi
- Parquet, Avro, Schema Registry
- Amazon Athena / Spark SQL
- ClickHouse / Druid / Pinot
- SQL Server / PostgreSQL
- Kubernetes (Amazon EKS)
- Helm, Argo CD, GitOps
- Terraform, GitHub Actions
- AWS (Amazon MSK, EMR Serverless, MWAA, S3, Athena, IAM)
- LLM, Embeddings, Vector Search
What You'll Do
- Design, build, and operate real-time streaming pipelines with Apache Kafka (Amazon MSK) and Apache Flink for high-throughput sensor and telemetry data;
- Define and manage streaming data contracts, including Avro schemas and schema evolution through a schema registry;
- Build and maintain analytical serving layers using ClickHouse or similar columnar OLAP databases, and develop REST APIs for dashboards, applications, and downstream teams;
- Develop and operate batch and scheduled data workflows with Apache Airflow (Amazon MWAA);
- Build and operate a data lakehouse based on Apache Iceberg and Amazon S3, using PySpark on EMR Serverless and Amazon Athena;
- Deploy and operate containerized data workloads on Kubernetes (Amazon EKS) using Argo CD, Helm, and GitOps practices;
- Manage cloud infrastructure as code with Terraform and support CI/CD automation with GitHub Actions;
- Monitor production data platforms with Prometheus and Grafana, troubleshoot issues, and participate in incident response;
- Contribute to cloud migration initiatives by porting data pipelines from existing platforms, including Azure-based data platforms, to AWS;
- Develop production-grade Java, Python, and SQL code with automated testing;
- Use AI-assisted development tools to accelerate analysis and implementation while maintaining code quality and architectural integrity;
- Keep technical documentation and operational runbooks current;
Job Requirements
- 7+ years of experience in data or software engineering, including production experience with streaming systems;
- Ability to work independently in ambiguous and fast-changing environments;
- Deep hands-on experience with Apache Kafka and Apache Flink or an equivalent stream-processing framework, including Java;
- Strong Python and SQL skills across transactional and analytical databases;
- Experience with Apache Iceberg, Delta Lake, or Hudi; Parquet, Avro, schema registries, Athena, or Spark SQL;
- Production experience with ClickHouse, Druid, Pinot, or similar analytical databases, as well as SQL Server or PostgreSQL;
- Experience with Kubernetes, Amazon EKS, Helm, Argo CD, and GitOps;
- Experience with Terraform and GitHub Actions;
- Strong knowledge of AWS services, including MSK, EMR Serverless, MWAA, S3, Athena, and IAM;
- Knowledge of ML fundamentals, including feature engineering, model training and evaluation, and ML data requirements;
- Familiarity with LLMs, prompt-based workflows, embeddings, vector search, anomaly detection, and forecasting;
- Experience with observability, alerting, troubleshooting, and incident response;
- Ability to communicate technical decisions clearly to technical and non-technical stakeholders;
Nice to Have
- Working knowledge of Azure Event Hubs, Data Factory, and Synapse;
- Experience with OT or industrial telemetry, including OPC UA, BMS/EPMS, or time-series sensor data;
- Experience with cloud-to-cloud or on-premises-to-cloud migrations;
- Experience with data center or other critical-infrastructure operations;
- Familiarity with data quality frameworks, data contracts, and metadata management;
What Do We Offer
The global benefits package includes:
- Technical and non-technical training for professional and personal growth;
- Internal conferences and meetups to learn from industry experts;
- Support and mentorship from an experienced employee to help you professional grow and development;
- Health insurance;
- Sports activities to promote a healthy lifestyle;
- Flexible work options, including remote and hybrid opportunities;
- Referral program for bringing in new talent;
- Work anniversary program and additional vacation days.