Sobre esta vaga de Senior Azure Data Engineer - #35291 na Manila Recruitment
Key Accountabilities/BAU Objectives:
- Pipelines from Source to Reporting
- Design, build and maintain the pipelines across the medallion layers, from raw ingestion through to the datasets reporting depends on. You own yours end to end, including how they behave when they fail.
- Pipelines delivered end to end without needing a rebuild
- Refresh schedules met, with failures caught by monitoring rather than by a user
- Pipelines built to the team standard so someone else can pick them up
Engineering Depth and Data Design
- Transformations are production code: readable, tested and maintainable. Idempotency and safe replay, correct incremental logic, modelling for how the data is queried, and the judgement to see that generated SQL is the wrong shape even when it returns the right number today.
- Transformation logic readable, tested and reusable rather than one off
- Pipelines that replay and backfill correctly by design
- Work another engineer can pick up without a handover conversation
Ingestion from Product Sources
- Bring data in from operational systems through gateways or equivalent connectors, handling incremental loads, schema drift and late arriving data without silent loss.
- Source data landed completely and repeatedly, reconciled against the source
- Incremental loads correct on replay and on backfill
- Schema changes detected and handled rather than discovered downstream
Data Quality and Schema Consistency
- Put quality checks in at each layer and keep schemas consistent across them. Where something is wrong, find where it entered rather than patching the layer it surfaced in.
- Validation at each layer boundary, with failures visible and owned
- Defects traced to the layer they entered and fixed there
- A quality check added for every data issue that reached a report
Reporting and the Semantic Layer
- Build and tune the datasets, models and measures that reporting runs on, and work with analysts and application developers to get them in front of the people who act on them.
- Reporting backed by fast refresh and metrics that reconcile
- Shared datasets reused rather than duplicated per report
- Business logic implemented once and shared, not repeated per report
Source Control, Deployment and Support
- Work reaches production through source control and a pipeline: small changes, reviewed before merge, nothing edited in place. Monitor what you shipped and support it, including a share of cover for the pipelines your team owns
- Changes deployed from source control rather than edited in the portal
- Pipeline failures alerted, triaged and closed to root cause
- Repeat manual interventions automated away rather than absorbed
Building with Agents, Checked Against Real Data
- Specify the transformation and the result it must produce, direct coding agents to write it, then verify the output against real data before you trust it. Keep the repository context and the checks current so the next person gets the same leverage
- Specifications and acceptance criteria written before the build
- Generated logic verified against known data before release, with the check kept as a test
- Repository context and checks current, and derived from real failures
Requirements
- At least 5 years of experience in Data Engineering
- Understands data architecture. Real delivery against a Medallion or equivalent layered design, with dimensional modelling judgement and an understanding of data warehousing and ETL or ELT design
- Can build the technology that enables it. Production grade pipelines, advanced SQL including joins, aggregations and incremental loads, and Python or Spark for transformations. Microsoft Fabric is what we run; Databricks, Snowflake or Synapse experience is equally credible.
- Has run what they built. Pipelines you supported in production, with monitoring, refresh reliability and recovery, and the habit of writing them to be replayed safely
- Reporting and governance. Power BI dataset modelling, DAX and performance tuning, with metrics that reconcile; data governance, lineage and role based access; and CI/CD for data using Azure DevOps, GitHub Actions or Fabric Git integration.
- Knowledge and awareness of agentic engineering: what these tools are, where they add value and where their output has to be checked against real data. Hands-on experience is beneficial rather than required.