Jobs Companies Roche Data Engineer - Pharma R&D

Über diese Data Engineer - Pharma R&D Stelle bei Roche

Roche · Vor Ort · Hyderabad

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections,  where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

Job Description:

The Data Engineer – Clinical Study Design sits at the intersection of data architecture, clinical science, and technology delivery, helping build the data foundation behind Study Designer, a digital product transforming how studies are designed. This role combines strong data engineering skills, an understanding of clinical study data and workflows, and technical curiosity to translate complex clinical data structures into reliable, scalable pipelines and models that power the platform's insights. The ideal candidate is curious, collaborative, AI-minded, and passionate about using data to enable smarter, faster, and more effective clinical study design.

Key Responsibilities

  • Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization

  • Design and implement data models in Amazon Aurora (relational) and GraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships

  • Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population

  • Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores

  • Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity (NCT IDs, EudraCT numbers, MeSH terms)

  • Build data serving APIs (Python/FastAPI) that expose curated datasets to the Angular frontend and LangGraph agent layer

  • Set up data lineage tracking and audit trails to support regulatory traceability of AI-generated trial design recommendations

Preferred Qualifications:

  • Education: Bachelor's degree in Computer Science, Data Engineering, or a related discipline.

  • 5-8 years of experience building production grade data platforms and pipelines. 

  • Experience with biomedical knowledge graphs (e.g., linking drugs -> targets -> diseases -> trials)

  • Prior work with PubMed/MEDLINE data, ClinicalTrials.gov API, or EMA/CTIS data

  • Apache Spark or Databricks for batch processing of large document 

Required Skills

  • Python — Primary language; experience with PDF/document parsing libraries (PyMuPDF, pdfplumber, unstructured.io, or similar)

  • SQL — Advanced PostgreSQL-compatible SQL (Aurora); schema design, migrations, query optimization, indexing strategies for clinical data volumes

  • Graph Databases — Hands-on with Neptune, Neo4j, or similar; SPARQL or Cypher query language; ontology/knowledge graph modeling for biomedical entities

  • AWS — Aurora (PostgreSQL), S3, Lambda, Step Functions, SQS/SNS, IAM; infrastructure for data pipeline orchestration

  • NLP / Document Processing — Text extraction from PDFs, section classification, named entity recognition for clinical/biomedical text; familiarity with embedding models and vector stores (OpenSearch, pgvector, or Pinecone)

  • FastAPI — Building data serving endpoints; async patterns; integration with the application backend

  • AI/ML Data Infrastructure — Preparing data for LangChain/LangGraph consumption; RAG pipeline design (chunking, retrieval, reranking); prompt-data integration patterns

  • Pipeline Orchestration — Experience with workflow orchestration tools (Airflow, Prefect, Step Functions, or Temporal); designing DAGs for multi-stage data pipelines with dependency management, retry logic, and monitoring

  • CI/CD & IaC — Terraform or CDK, Docker, Git; automated pipeline testing and deployment on AWS

#Hyderabad2026

 

 

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.


Let’s build a healthier future, together.

Roche is an Equal Opportunity Employer.

Bereit, sich bei Roche zu bewerben?
Bei Roche bewerben

Über Roche

We believe it’s urgent to deliver medical solutions right now – even as we develop innovations for the future. We are passionate about transforming patients’ lives. We are courageous in both decision and action. And we believe that good business means a better world. That is why we come to work each day. We commit ourselves to scientific rigor, unassailable ethics, and access to medical innovations for all. We do this today to build a better tomorrow. We are proud of who we are, what we do, and how we do it. We are many, working as one across functions, across companies, and across the world. We are Roche.

Alle Jobs bei Roche ansehen →

Ähnliche Jobs

Capgemini
FBS Data Engineer
Capgemini
⚡ Früh bewerben Hyderabad, Telangana, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
Cohere Health
Senior Data Engineer
Cohere Health
⚡ Früh bewerben Hyderabad, Telangana, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Tg.
DXC Technology
Senior Azure Data Engineer at Hyderabad
DXC Technology
⚡ Früh bewerben IND - AP - HYDERABAD Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
Mindrift
Senior Python Data Scraping Engineer (Freelance)
Mindrift
⚡ Früh bewerben India · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
Lloyds Banking Group
Principal Data Engineer
Lloyds Banking Group
⚡ Früh bewerben Hyderabad Knowledge Park Tower... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
Wells Fargo
Senior Software Engineer- SQL, Data Engineer
Wells Fargo
⚡ Früh bewerben Hyderabad, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
Carrier
Senior Engineer-Data Analytics
Carrier
⚡ Früh bewerben Building No: 12C, Floor 9,10,1... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
DoorDash India
Software Engineer, Data Engineer II
DoorDash India
⚡ Früh bewerben Hyderabad, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.
Zinnia - Employee Referral
Data Engineer III
Zinnia - Employee Referral
⚡ Früh bewerben Gurugram, Haryana, India; Hyde... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Tg.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Roche

Alle Jobs bei Roche ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos