Jobs Companies Formation Bio Senior Data Engineer

À propos de ce poste Senior Data Engineer chez Formation Bio

Formation Bio · Hybride · New York, NY; Boston, MA; San Francisco, CA

About Formation Bio

Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. 

Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress because of the high cost and time of clinical trials. Recognizing that this development bottleneck may ultimately limit the number of new medicines that can reach patients, Formation Bio, founded in 2016 as TrialSpark Inc., has built technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials. Formation Bio partners, acquires, or in-licenses drugs from pharma companies, research organizations, and biotechs to develop programs past clinical proof of concept and beyond, ultimately helping to bring new medicines to patients. The company is backed by investors across pharma and tech, including a16z, Sequoia, Sanofi, Thrive Capital, John Doerr, Spark Capital, SV Angel Growth, and others. 

You can read more at the following links:

At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry. Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently.

About the Position

As a Senior Data Engineer at Formation Bio, you will build the trusted data systems that support clinical operations, drug asset evaluation, business development, analytics, and machine learning through AI Enabled Employees and Agents. You will work across clinical, operational, and third-party data sources to design and operate reliable ingestion pipelines, transformations, data models, and data products.

This role sits at the intersection of Product Engineering, Data Engineering, and Data Science. You will help shape how application data is modeled and exposed, own shared data models and production data products, and prioritize data platform work against clinical and business needs. You will work closely with Product Engineering on application data models and contracts, with Data Science on training datasets and ML use cases, and with human and AI consumers of the data platform.

A data product is not complete merely because it is technically correct or available in a warehouse. It should be understandable to people, usable by applications, useful to Data Science, and structured so AI Enabled Employees and Agents can access it reliably and safely.

Responsibilities

  • Design and operate production data systems that ingest clinical, operational, and third-party vendor data into reliable, queryable data products.
  • Own shared and canonical data models, data contracts, transformations, orchestration, warehouse models, and downstream interfaces.
  • Partner with Product Engineering on application data models, source-system contracts, APIs, events, and data access patterns.
  • Partner with Data Science on productionized training datasets, feature pipelines, data interfaces, and ML use cases.
  • Turn recurring data cleaning, normalization, and transformation work into versioned, tested, observable, and maintainable production pipelines.
  • Build data products for clinical operations, asset evaluation, Business Development, analytics, machine learning, and AI Enabled Employees and Agents.
  • Design data products that are semantically clear, discoverable, machine-readable, permission-aware, traceable, and safe to query.
  • Establish strong data quality, testing, freshness, completeness, lineage, documentation, and observability practices.
  • Own data governance practices for sensitive and regulated data, including access controls, auditability, traceability, and appropriate data handling.
  • Participate in support and incident response for data platform issues, including diagnosis, stakeholder communication, remediation, and prevention of recurrence.
  • Use AI tools, including LLMs and agentic coding systems, to accelerate pipeline development, data modeling, debugging, documentation, and data quality investigation while validating their output.
  • Contribute to architecture and design reviews, mentor other engineers, and improve the engineering practices used across the organization.

About You

  • 5+ years of relevant data engineering experience building and operating production data systems.
  • Experience with pharmaceutical, biology, HIPPA or other regulated data core to BioTech is required.
  • Strong Python and SQL skills, with deep experience in data modeling and warehouse systems, especially Snowflake.
  • Experience with Dagster as an orchestration systems (or equivalent) transformation tooling such as dbt or an equivalent approach.
  • Experience with data contracts, schema evolution, data quality testing, observability, lineage, and production incident response.
  • Experience integrating messy clinical, operational, vendor, or otherwise complex source data.
  • Working knowledge of Docker, GitHub, and Terraform or OpenTofu sufficient to partner effectively with SRE.
  • Experience building data products and access patterns for applications, Data Science, analytics, human users, and AI Enabled Employees and Agents.
  • Strong judgment about when to build reusable platform capabilities versus one-off stakeholder solutions.
  • Daily fluency with AI tools and the ability to validate generated code, transformations, and data-modeling decisions.
  • Exceptional collaboration and communication skills across Product Engineering, Data Science, Clinical Operations, Data Management, Business Development, and other non-technical partners.
  • Experience working within and building validated computerized systems (CSV) is a plus.

Total Compensation Range: $185,500 - $232,000

Compensation Individual compensation is determined by several factors, including role scope, geographic location, and skills & experience. Your offer will reflect where you fall within the range based on these considerations. In addition to base salary, we offer equity, comprehensive benefits, and generous perks. If the posted range doesn't match your expectations, we still encourage you to apply!

Where We Hire Formation Bio is prioritizing hiring in key hubs, primarily the New York City and Boston metro areas, with a hybrid model requiring 3 days per week in office. Applicants from the Research Triangle (NC) and San Francisco Bay Area may also be considered. Please apply only if you reside in these locations or are willing to relocate.

Equal Opportunity Formation Bio is committed to building a diverse and inclusive team. We are an equal opportunity employer and welcome candidates from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, national origin, ancestry, sex (including pregnancy, childbirth, breastfeeding, and related medical conditions), gender identity or expression, sexual orientation, age, disability, genetic information, marital status, military or veteran status, or any other characteristic protected by federal, state, or local law.

Prêt à postuler chez Formation Bio ?
Postuler chez Formation Bio

Comment se compare ce salaire pour Data Engineer

Ce poste paie $208,750/yrau-dessus de la fourchette habituelle pour les postes Data Engineer.

$97,184 la médiane $157,500 $242,340

Fourchette typique $125,000–$197,500/yr, à partir de 1,635 annonces Data Engineer comparables sur JobsRadar (rémunération annualisée en USD). Voir les aperçus de salaire pour Data Engineer →

Emplois similaires

DP
Data Engineer
DLA Piper
⚡ Postuler tôt Reston, VA Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
PwC
Data Engineer - Senior Manager
PwC
⚡ Postuler tôt IL-Chicago Sur site $124,000–$280,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 mois
Credence
Data Engineer
Credence
⚡ Postuler tôt McLean, Virginia, United State... Sur site $115,000–$155,000
● Nouveau 👁 Vu ✓ Postulé il y a 15 min
payabl.
Data Engineer
payabl.
⚡ Postuler tôt Poland · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 26 min
Vanta
Staff Software Engineer, AI and Data Risk
Vanta
⚡ Postuler tôt London, UK Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 37 min
Stand Insurance
Member of the Technical Staff - Data Engineer
Stand Insurance
⚡ Postuler tôt San Francisco Hybride $240,000–$295,000
● Nouveau 👁 Vu ✓ Postulé il y a 39 min
Omni
Software Engineer, Growth Data Platform
Omni
⚡ Postuler tôt San Francisco, CA · lieu restreint $170,000–$250,000
● Nouveau 👁 Vu ✓ Postulé il y a 43 min
Infinity
Senior Data Engineer - Paradox Machines
Infinity
⚡ Postuler tôt United States · lieu restreint $140,000–$170,000
● Nouveau 👁 Vu ✓ Postulé il y a 45 min
Hilbert's AI
Forward Deployed Data Engineer - LATAM
Hilbert's AI
⚡ Postuler tôt LATAM · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 46 min

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Formation Bio

Voir tous les emplois chez Formation Bio →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit