About this Data Engineer role at Coretek Services
Coretek is looking for a Data Engineer to build and operate the pipelines and data models that the rest of the business runs on. You'll own ingestion from source systems through to curated, well-documented datasets that analysts, data scientists, and application teams depend on. This is a hands-on engineering role: you'll write production code, design schemas, and be accountable for the reliability and cost of what you ship.
Responsibilities
- Design, build, and maintain batch and streaming data pipelines that are idempotent, observable, and recoverable.
- Model data for analytics (dimensional models, semantic layers, and curated marts), balancing query performance against maintainability.
- Integrate data from operational databases, SaaS APIs, files, and event streams, including handling schema drift and late-arriving data.
- Build data quality checks (freshness, volume, uniqueness, referential integrity) into pipelines rather than bolting them on afterward, and define how failures alert and escalate.
- Own pipelines in production: monitoring, on-call rotation for data incidents, root-cause analysis, and backfills.
- Tune performance and cost (partitioning, clustering, file sizing, warehouse and cluster sizing) and make the tradeoffs explicit.
- Apply engineering discipline to data: version control, code review, CI/CD, automated testing, and infrastructure as code.
- Implement access controls, PII handling, retention, and lineage and audit requirements in partnership with security and compliance.
- Partner with analysts, data scientists, and product engineers to turn ambiguous requirements into durable data contracts.
- Maintain data dictionaries, lineage, and pipeline runbooks so consumers can find a dataset, understand what each field means and how current it is, and use it correctly without having to ask the team that built it.
Requirements
- 5+ years building production data pipelines.
- Strong hands-on Python development for data engineering, with real testing, packaging, and code review practice, not scripting alone.
- Working knowledge of PySpark: DataFrame and SQL APIs, joins and aggregations at scale, partitioning and shuffle behavior, and the ability to read a Spark UI to diagnose a slow or failing job.
- Strong SQL: window functions, query plans, and performance tuning, not just SELECTs.
- Hands-on experience with the Azure data platform: Data Factory, Databricks, Synapse/Fabric, and ADLS.
- Solid data modeling fundamentals: normalization, star schemas, slowly changing dimensions.
- Git-based workflow and experience shipping through CI/CD.
- Excellent communication skills, with the ability to debug a failing pipeline end to end and articulate the impact to diverse audiences, including non-technical stakeholders.
- Exceptional analytical and problem-solving skills, with the judgment to find the root cause of a data issue rather than patching the symptom.
- Strong knowledge and experience in working with customers in a consultative approach in a technical environment.
Additional Qualifications
- Streaming experience (Kafka, Event Hubs).
- Lakehouse formats: Delta Lake, Iceberg.
- Infrastructure as code (Terraform, Bicep) and containerization (Docker, Kubernetes).
- Experience in a regulated environment (HIPAA, SOC 2, PCI, GDPR): auditability, encryption, data residency.
- Experience building data platforms for ML or supporting feature pipelines.
- Proven ability to manage multiple client projects and deliver high-quality results on time.
- Experience in Azure DevOps or GitHub for source control and pipelines.