Jobs › Companies › Howard Hughes Medical Institute › AI Data Engineer

Über diese AI Data Engineer Stelle bei Howard Hughes Medical Institute

Howard Hughes Medical Institute · Hybrid · Headquarters
Primary Work Address: 4000 Jones Bridge Road, Chevy Chase, MD, 20815

Current HHMI Employees, click here to apply via your Workday account.

HHMI is focused on supporting and moving science forward in a variety of different ways ranging from conducting basic biomedical research, empowering educators, inspiring students, developing the next generation of scientists – even stretching into film and media production.  Our Headquarters is in the greater Washington, DC metro area and is home to over 300 employees with expertise in investments, communications, digital production, biomedical sciences, and everything in between.  The work housed here supports and augments the groundbreaking research conducted in HHMI labs across the nation.  As HHMI scientists continue to push boundaries in laboratories and classrooms, you can be sure that your contributions while working here are making a difference.

The AI Accelerator exists to turn AI into daily reality across HHMI’s administrative and operational functions. This role builds the data-engineering foundation every AI application and knowledge layer at HHMI runs on.


The work is hands-on. This person implements the pipelines, transformation patterns, and orchestration framework that turn HHMI’s institutional content into governed, AI-ready data. They own the day-to-day execution of the data-engineering side of the AI fabric — pipeline development, medallion curation, workflow-orchestration framework, and governance implementation — under the design authority of the Principal Knowledge & Data Architect.


This role works closely with the Principal Knowledge & Data Architect (who owns the knowledge and retrieval layer’s design), the AI Developer (who consumes what this role builds), and Operations Capabilities’ Data Integration Engineer (who lands source-system data at the interface). The AI Data Engineer is the seat that connects those layers into working data flow.


HHMI’s Principal K&D Architect designs the knowledge and data foundation for institutional AI, but the foundation only comes to life when someone builds and operates the pipelines that carry data through it. Without this role, our AI systems either wait on data that isn’t ready or consume content whose quality can’t be guaranteed. This role is what makes the K&D Architect’s design real day-to-day — and it’s the same role that keeps it working over years, as content evolves and models change.


This position works on a hybrid schedule, reporting in-person three days a week to our headquarters in Chevy Chase, Maryland. 


We encourage qualified candidates who are eligible to work in the United States to apply. Please note, we are not able to sponsor a visa for this position at this time.


What you will do

  • Build the AI-facing data pipelines: ingestion from the landing zone Operations Capabilities delivers, transformation through raw → bronze → silver → gold, and serving of governed, AI-ready content for downstream consumption.
  • Implement the medallion architecture: design patterns from the K&D Architect become working pipelines, tables, and materialization schedules. Delta Lake tables designed with partitioning, optimization, and evolution in mind.
  • Own the workflow-orchestration framework: Databricks Workflows, Delta Live Tables, retry policies, alerting routes, run history, cost tags.
  • Implement governance patterns: Unity Catalog structure, sensitivity classification, access control, and audit for AI-facing data assets. Design comes from the K&D Architect; day-to-day implementation lives here.
  • Build and operate retrieval-supporting infrastructure: embedding pipelines, vector store maintenance, reindexing when models upgrade, retrieval evaluation frameworks.
  • Partner with Operations Capabilities on source-system contracts: define what the AI Fabric consumes at the landing zone — schema, cadence, SLA, quality thresholds. Own the platform-side of that contract.
  • Design and operate data quality: data-quality checks, freshness monitoring, drift detection, and the alerting that surfaces issues before they hit AI users.
  • Support AI Developer velocity: when AI Developers deploy into product teams, this role is the data engineer they turn to when a use case needs specific data.
  • Contribute to and consume the reference-pattern library: reusable pipeline patterns, code templates, and standards live in the shared platform layer. This role uses them, contributes new ones, and evolves them as we learn.

What We Are Looking For

  • Hands-on production data engineering: at least four years designing, building, and operating production data pipelines. Not a Databricks-course track record — real production experience where you owned the pipeline through breakage, iteration, and recovery.
  • Databricks and Spark depth: Delta Lake, medallion architecture, Delta Live Tables, Workflows, Databricks SQL, Unity Catalog. Comfortable at the layer where code meets platform.
  • Python and SQL fluency: PySpark, ETL patterns, and SQL that runs at scale. Version control (Git), CI/CD for data pipelines, and infrastructure-as-code (Terraform) as working tools.
  • AI-adjacent data engineering: real experience building the data foundation for AI use cases — embedding pipelines, vector stores, chunking strategies, retrieval evaluation. Not required to be an ML researcher; required to have built the plumbing.
  • Workflow orchestration: Databricks Workflows, or Airflow in production. Retry semantics, dependency management, failure handling — as working discipline, not concepts.
  • Data quality and observability: Great Expectations, Databricks data-quality monitors, or equivalent. Treats data quality as a first-class engineering concern.
  • Governance discipline: works with Unity Catalog structures, understands sensitivity classification, and designs pipelines with access control and audit in mind from the first commit.
  • AWS foundations: IAM, S3, KMS at the level needed to work in a Databricks-on-AWS environment. Not required to be a cloud architect; required to be productive.
  • Communication: works productively with the K&D Architect on design, AI Developer on integration, and Operations Capabilities on contracts. Explains data-engineering trade-offs to non-engineers.
  • Education and experience: bachelor’s degree or equivalent, plus at least four years of hands-on data-engineering experience with meaningful exposure to AI or knowledge-management use cases.

Nice to Have

  • Prior experience with knowledge graphs (Neo4j or comparable), entity resolution, or semantic data models.
  • Experience with the modern data stack alongside Databricks-native tooling.
  • Familiarity with LLM-based extraction, chunking, and evaluation frameworks.
  • Background in research, academic, or mission-driven institutional environments.
  • Experience with cross-platform data engineering (Snowflake, BigQuery) — cross-platform judgment is useful even when Databricks is the primary tool.

What This Role Is Not

  • Not a Data Platform Architect: design authority for the Databricks-on-AWS platform sits with Platform Architects. This role builds AI-facing data pipelines on that platform.
  • Not the Knowledge & Data Architect: the K&D Architect owns the design of the knowledge and retrieval layer; this role implements against that design. Required to be a strong hands-on builder, not an architect.
  • Not an AI/ML engineer: the AI Developer owns model access, orchestration, and application delivery. This role builds the data layer AI Developers consume.
  • Not a warehouse or BI data engineer: traditional analytics warehouses and BI-facing pipelines belong to Operations Capabilities. This role is AI-facing.

Practical Details

This role is hybrid, with three days per week in-person at HHMI’s offices in Chevy Chase, Maryland. It reports to the Director of AI Enablement. HHMI is not able to sponsor a visa for this position at this time.


Physical Requirements

Remaining in a normal seated or standing position for extended periods of time; reaching and grasping by extending hand(s) or arm(s); dexterity to manipulate objects with fingers, for example using a keyboard; communication skills using the spoken word; ability to see and hear within normal parameters; ability to move about workspace. The position requires mobility, including the ability to move materials weighing up to several pounds (such as a laptop computer or tablet).

 

Persons with disabilities may be able to perform the essential duties of this position with reasonable accommodation. Requests for reasonable accommodation will be evaluated on an individual basis.

Please Note:

This job description sets forth the job’s principal duties, responsibilities, and requirements; it should not be construed as an exhaustive statement, however. Unless they begin with the word “may,” the Essential Duties and Responsibilities described above are “essential functions” of the job, as defined by the Americans with Disabilities Act.

#LI-RK1

Compensation and Benefits

Our employees are compensated from a total rewards perspective in many ways for their contributions to our mission, including competitive pay, exceptional health benefits, retirement plans, time off, and a range of recognition and wellness programs. Visit our Benefits at HHMI site to learn more. 

Hiring Pay Range

$128,816.80 - $161,021.00

Pay Type:

Annual

The posted range reflects HHMI’s good faith estimate of the anticipated hiring salary range for this role at the time of posting. Actual hiring compensation is determined by a candidate’s qualifications, experience, and internal equity.



HHMI is an Equal Opportunity Employer

We use E-Verify to confirm the identity and employment eligibility of all new hires.

Bereit, sich bei Howard Hughes Medical Institute zu bewerben?
Bei Howard Hughes Medical Institute bewerben

Wie sich dieses Gehalt für Data Engineer vergleicht

Diese Stelle zahlt $144,919/yr — im Einklang mit der üblichen Spanne für Data Engineer Stellen.

$136,250 dem Median $150,334 $209,000

Übliche Spanne $143,105–$178,063/yr, aus 6 vergleichbaren Data Engineer Anzeigen auf JobsRadar (Vergütung auf USD hochgerechnet). Gehaltseinblicke für Data Engineer ansehen →

Über Howard Hughes Medical Institute

Howard Hughes Medical Institute (HHMI) is an independent, ever-evolving philanthropy that supports basic biomedical scientists and educators with the potential for transformative impact. We make long-term investments in people, not just projects, because we believe in the power of individuals to make breakthroughs over time. Why HHMI To move science forward we need a diverse collection of talents, expertise, and backgrounds in scientific research and science education, as well as communications, finance, human resources, information technology, investments, law, and operations. At HHMI, we encourage collaborative and results-driven working styles and offer an adaptable environment where empl

Alle Jobs bei Howard Hughes Medical Institute ansehen →

Ähnliche Jobs

New Balance
Principal Data Engineer
New Balance
⚡ Früh bewerben Boston, MA Headquarters - (NB) Hybrid $138,500–$173,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
DA
Data Engineer II
DailyPay
⚡ Früh bewerben NYC Headquarters Hybrid $120,000–$140,000
● Neu 👁 Gesehen ✓ Beworben vor 5 Tg.
Garner Health
Data Engineer III
Garner Health
⚡ Früh bewerben New York City, New York Hybrid $166,000–$205,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Garner Health
Senior Data Engineer
Garner Health
⚡ Früh bewerben New York City, New York Hybrid $220,000–$245,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Fischer Homes
SENIOR DATA PLATFORM ENGINEER
Fischer Homes
⚡ Früh bewerben Erlanger, KY Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Wo.
Bitvavo
Senior Data Engineer - Storage & Messaging
Bitvavo
⚡ Früh bewerben Headquarters Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Wo.
Bitvavo
Senior Data Engineer - Data Platform
Bitvavo
⚡ Früh bewerben Headquarters Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Wo.
Singapore Public Service
Data Engineer - Integrated E-Services [ITE Headquarters] – 2 Yr Cr
Singapore Public Service
⚡ Früh bewerben ITE-HQ (Headquarters) Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Wo.
Singapore Public Service
Data & Analytics Engineer - Learning Technologies Devt [ITE Headquarters]
Singapore Public Service
⚡ Früh bewerben ITE-HQ (Headquarters) Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Wo.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Howard Hughes Medical Institute

Alle Jobs bei Howard Hughes Medical Institute ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos