Über diese Data Engineer Stelle bei CipherHealth
About Us
CipherHealth is an award winning software company committed to enhancing care coordination and outcomes across the continuum. Since 2009, CipherHealth’s automated, scalable platform has empowered healthcare organizations to engage patients and care teams at every touchpoint, streamlining workflows and improving experiences. With tailored communication solutions powered by AI and deep integrations, CipherHealth drives better clinical results, operational efficiency, and financial sustainability, transforming healthcare one interaction at a time.
About the Role
We're looking for a Data Engineer to own the pipelines, models, and data platform that make trustworthy analytics and AI possible. This is an individual contributor role: you will personally design, build, and operate the data backbone—ingestion, transformation, orchestration, quality, and the semantic foundations that analysts, agents, and GTM systems depend on—while partnering closely with Product, Engineering, Analytics, Marketing, Sales, and Customer Success.
What makes this role different is how we do data engineering. You will build, operate, and continuously improve an agent-augmented data platform where AI agents handle much of the monitoring, diagnosis, documentation, and first-pass remediation—freeing you to focus on architecture, reliability, and high-leverage systems design. The ideal candidate pairs deep, traditional data engineering craft with an agentic mindset: someone who thinks about the data platform as a system of humans and agents working in a loop, and who is as comfortable hardening a production pipeline or evolving a warehouse model as designing an agent workflow that turns pipeline noise into prioritized engineering action.
Core Data Engineering Responsibilities
- Own the data platform. Design, build, and operate the warehouse, transformation layer, orchestration, and serving surfaces that power analytics, product insight, and AI agents.
- Build reliable pipelines. Deliver robust ingestion and ELT/ETL for product events, CRM, marketing, finance, support, and other critical sources—with clear contracts, retries, and observability.
- Model data for reuse. Create well-tested, documented dimensional and semantic models so metrics and entities stay consistent across dashboards, reverse ETL, and agents.
- Enforce data quality. Implement tests, anomaly detection, lineage, SLAs, and incident response so freshness and correctness are operational, not aspirational.
- Instrument and contract with product. Partner with Engineering on event taxonomy, tracking plans, and production data contracts so product changes don't silently break downstream systems.
- Enable analytics and AI. Provide stable, governed datasets and interfaces that analysts, MarOps, and agents can query safely—without becoming a ticket bottleneck.
- Secure and govern access. Apply least-privilege access, PII handling standards, and documentation that keep the platform usable and compliant.
- Scale the operating model. Document architecture and runbooks; automate toil; continuously raise reliability and developer experience—without building a team under you.
What You'll Own
- Build and Operate the Agentic Data Platform Loop
You will personally stand up and continuously improve a fleet of data-platform agents that augment engineering and reliability work. These agents should be able to:
- Monitor pipeline health and flag freshness, volume, schema drift, and quality anomalies before stakeholders lose trust
- Diagnose failed jobs with first-pass root-cause summaries—logs, upstream changes, and likely blast radius
- Detect warehouse cost and performance regressions and propose concrete optimization candidates
- Draft first-pass dbt/SQL models, tests, and documentation so humans start from a strong scaffold rather than a blank page
- Summarize schema and contract changes across sources and recommend migration or compatibility steps
- Generate weekly platform health reports that synthesize reliability, backlog risk, and prioritized engineering bets
- Watch data contracts and event instrumentation for breaks, missing properties, and taxonomy inconsistencies
- Propose and draft remediation PRs or runbook steps for recurring classes of incidents
- Answer routine platform and lineage questions with grounded references to approved models and docs—and escalate ambiguity
You own the quality, reliability, and evolution of this "Data Platform" agent layer. You treat it as a product in its own right—defining requirements, measuring accuracy and adoption, and iterating on the leverage it creates. The agentic loop doesn't replace data engineering craft; it accelerates detection, diagnosis, and scaffolding, and gives you more time for architecture, hard trade-offs, and cross-functional partnership.
- Run the Data Platform Operating System
Data platform work at our company depends on a crisp operating model, and you are the steward of that system—as the owner and operator, not as a people manager.
- Own the recurring cadence for pipeline reviews, incident retros, SLA reporting, and intake prioritization across Product, Analytics, and GTM.
- Ensure every major data product has clear ownership, tests, documentation, and rollback paths grounded in platform standards and the outputs of the agentic loop.
- Keep platform decisions evidence-based and priorities laddered up to company strategy—reliability, speed of insight, and AI readiness.
- Own Core Platform Surfaces
You are accountable for the reliability, cost, and usability of the company's core data systems.
- Own core warehouse models, transformations, orchestration, and the semantic/serving layer that powers dashboards, reverse ETL, and agents.
- Partner with Engineering on instrumentation, event taxonomy, and production data contracts so product changes don't silently break measurement or AI workflows.
- Make the hard trade-offs about where investment creates maximum return—pipeline reliability, modeling quality, self-serve tooling, cost efficiency, or agent quality.
How You Work
This is a high-leverage individual contributor role. You do the work yourself and multiply impact through systems, agents, and clear operating rhythms—not through direct reports.
- You own outcomes end to end — pipelines, models, quality, and the platform backbone stakeholders and agents trust.
- You build and run the agentic data-platform loop owning the agents, their outputs, and how they plug into engineering and decision forums.
- You enable partners through craft and process — clear contracts, strong models, working alongside agents, and helping teams build on a stable foundation.
- You have no people-management responsibility — influence comes from expertise, systems design, and trusted delivery.
What Success Looks Like
- A trusted, well-operated data platform that Product, Analytics, GTM, and leadership rely on without second-guessing.
- Pipelines and models that meet freshness and quality SLAs, with fast detection and recovery when they don't.
- Clear data contracts and instrumentation practices that keep product and platform changes safe.
- The agentic data-platform loop reliably shortens the path from signal to diagnosis to fix.
- Lower operational toil and better cost/performance as usage and data volume grow.
- A scalable operating model that raises platform quality as the company grows—without requiring a larger data headcount to keep pace.
What We're Looking For
- Proven data engineering craft. 5+ years building and operating production data pipelines and warehouse models as a hands-on IC; comfort spanning ingestion, transformation, orchestration, and serving.
- Modern data stack fluency. Strong SQL; experience with cloud warehouses (e.g., BigQuery/Snowflake/Redshift), transformation tooling (e.g., dbt or equivalent), orchestration (e.g., Airflow/Dagster/Prefect), and Python for platform automation.
- Reliability mindset. You design for observability, testing, SLAs, incident response, and cost control—not just happy-path delivery.
- Systems thinking. You design architectures and models that hold up under growth, schema change, and cross-functional pressure.
- Agentic mindset. You believe AI agents change how data platform work gets done, and you have hands-on instincts for designing agent workflows—defining requirements, evaluating outputs, grounding answers in lineage and docs, and iterating on quality.
- Cross-functional influence. You partner with Product Engineering, Analytics, and GTM and hold the platform bar high—without formal authority over them.
- Builder and multiplier. You improve the systems and the partners around you—documenting, automating, and raising the standard of how work gets done. You thrive as a senior IC, not by managing a team.
How We Invest In You
- Healthcare that begins on your first day:
- Generous company-funding of our health, vision, and dental plans
- HSA/FSA plans
- Short and Long-Term Disability
- Life and Personal Accident Insurance
- $40 monthly wellness stipend you can use towards any wellness, fitness, and wellbeing purchases
- Employee Assistance Program (EAP)
- Adoption Assistance
- Retirement: 401(k) at three months of employment — with a match upon enrollment!
- Time away:
- Discretionary PTO + 13 paid holidays
- Parenthood: Competitive paid parental leave and flexible return to work policy
- Recognition:
- Generous Employee Referral Program - earn cash for each employee referral that is hired
- Yearly Cipher-versary stipend
- Ci-Phives - receive public kudos and gift cards from peers and managers
- Culture:
- CARE2 Values
- Monthly All Teams Meetings
- Employee Resource Groups such as Rainbow Room and BIPOC Group
- Internal Webinars and robust onboarding / training programs
- Remote-first team: $50 per month reimbursement in your check for WFH expenses
- You’ll receive a new Macbook laptop, other hardware, and company swag upon hire