Jobs Companies Irth Solutions Data Engineer (Mid Level)

À propos de ce poste Data Engineer (Mid Level) chez Irth Solutions

Irth Solutions · Télétravail · United States

Data Engineer – Insights (AI/ML)

Location: Remote (US)
Department: Insights (AI/ML)
Reports to: Engineering Manager

About the Role

Irth is building a new AI-driven threat and risk management platform for pipeline asset integrity. The platform brings together three capabilities that have historically been separate at Irth:

  • A governed, cross-product data platform built on Databricks and Azure
  • An AI-powered ingestion layer that normalizes, repairs, and enriches customer data without services-heavy onboarding
  • A reusable analytical layer that runs industry-standard, Irth-developed, and customer-built risk models against the data

As a Data Engineer, you will build the pipelines that make this platform real. Pipeline operators hold their integrity data across inline inspection reports, GIS systems, maintenance records, spreadsheets, scanned documents, and enterprise systems of record. Getting that data ingested, cleaned, aligned, and made model-ready is one of the biggest obstacles to adoption in this market—and it is the problem this role exists to solve.

You will implement ingestion and transformation pipelines based on patterns established by the Data Architect, build the AI-assisted ingestion layer in partnership with data scientists, and help operationalize both. This role is well suited to a mid-level engineer who wants to deepen their expertise in Databricks, Spark, and modern lakehouse engineering while helping build a platform from the ground up.

Key Responsibilities

1. Pipeline Development — Primary Responsibility

  • Build and maintain ingestion pipelines for structured and semi-structured sources, including GIS, inline inspection data, SCADA, maintenance systems, and enterprise systems of record.
  • Implement batch and streaming ingestion using Databricks Workflows, Spark, PySpark, SQL, and declarative pipeline tooling.
  • Apply medallion architecture patterns (Bronze, Silver, Gold) for transformation, standardization, and enrichment.
  • Implement change data capture (CDC), slowly changing dimensions (SCD), schema evolution, and data-validation rules.
  • Normalize third-party and public data feeds, including weather history, soil characteristics, satellite-derived data, and one-call ticket data, into the shared data model.

2. AI-Assisted Ingestion Layer

  • Build pipelines that automate normalization of units, schemas, and semantics across inconsistent customer data.
  • Implement automated data-quality repair workflows, including gap filling, error correction, and reconciliation, with clear provenance for every synthesized value.
  • Work with data scientists to productionize document-extraction pipelines that parse reports, spreadsheets, and field records into the target schema.
  • Build human-in-the-loop review and exception workflows so low-confidence extractions are surfaced rather than propagated silently.

3. Platform & Storage Implementation

  • Configure and manage Delta Lake tables, partitioning strategies, and optimization routines.
  • Implement metadata, lineage, and cataloging standards using Unity Catalog.
  • Build and maintain connectors to customer systems of record with configurable refresh cadences.
  • Support geospatial data processing, including spatial joins and alignment of results to pipeline centerline geometry.

4. Governance, Quality & Compliance Enablement

  • Implement data-quality tests, profiling, and drift monitoring based on standards established by the Data Architect.
  • Apply access-control policies, security rules, and classification tags defined by the governance model.
  • Implement lineage capture sufficient to support regulatory traceability from ingestion through model output.

5. Orchestration, Automation & Operational Support

  • Build, schedule, and monitor data workflows, and own alerting and failure handling for the pipelines you develop.
  • Contribute to CI/CD for pipeline code, including version control, automated testing, and environment promotion.
  • Troubleshoot production incidents, recover failed pipeline runs, and optimize performance and infrastructure costs.

6. Collaboration & Documentation

  • Work closely with the Data Architect to translate architectural designs into production implementations and identify gaps or ambiguities in the design.
  • Participate in architecture, design, and code reviews.
  • Document pipelines, transformation logic, data dictionaries, job schedules, operational procedures, and runbooks.

Requirements

Required Qualifications

  • 3–5 years of experience in data engineering, ETL development, or cloud data platform engineering.
  • Hands-on experience with Databricks, Spark, PySpark, or comparable distributed data-processing technologies.
  • Strong SQL skills and experience with structured data transformation.
  • Experience with at least one major cloud platform; Azure experience preferred.
  • Familiarity with data modeling, data-quality practices, schema evolution, and pipeline troubleshooting.
  • Experience with workflow orchestration and scheduling frameworks.
  • Understanding of core data-security practices, including access control, encryption, and credential management.
  • Experience with Git-based development and comfort working within a code-reviewed engineering team.

Preferred Qualifications

  • Experience with Delta Lake, medallion architecture, and lakehouse engineering best practices.
  • Experience with Unity Catalog, Microsoft Purview, or comparable metadata and data-lineage tooling.
  • Experience building pipelines that ingest unstructured or semi-structured documents.
  • Experience with geospatial data processing and common GIS data formats.
  • CI/CD and DevOps experience for data workloads, including infrastructure as code (IaC).
  • Experience preparing and transforming data specifically for machine learning or probabilistic model consumption.
  • Cloud or Databricks certifications.
  • Experience using AI-assisted coding tools such as Cursor or GitHub Copilot and/or agentic coding tools such as Claude Code as part of a professional development workflow.

Nice to Have

  • Experience integrating oil and gas or utility asset data, including pipelines, facilities, and GIS assets, into a data platform.
  • Understanding of asset integrity concepts, including inspection data, risk scoring, corrosion, and defect tracking.
  • Familiarity with regulatory and compliance reporting requirements for pipeline or asset integrity data.
  • Experience migrating customers from legacy or spreadsheet-based systems to modern data platforms.

Success Metrics

Success in this role will be measured by:

  • Reliable, well-documented pipelines delivering consistent Bronze, Silver, and Gold data flows.
  • Measurable reduction in the manual effort required to onboard new customer data.
  • High data-quality pass rates, with failures identified and contained at ingestion rather than downstream.
  • Full compliance with established cataloging, lineage, security, and governance standards.
  • Low pipeline incident rates, with fast recovery and clear, actionable runbooks when incidents occur.
  • Effective collaboration with the Data Architect, data scientists, application engineers, and other cross-functional partners.
Prêt à postuler chez Irth Solutions ?
Postuler chez Irth Solutions

À propos de Irth Solutions

Above our heads and below our feet lie a spiderweb of cables, pipelines, electric lines, sewer systems, telecommunications networks and more. Many of the conveniences that we take for granted in our daily lives is thanks to this critical network infrastructure. From the assets themselves to the people that manage and maintain them, everything and everyone must work together seamlessly to ensure reliability, resiliency, safety and ongoing compliance with local and federal mandates.

For decades, Irth Solutions has worked with leaders across the energy, utility, municipality, telecommunication and media industries to provide industry-leading technology to help enhance the reliability and resiliency of their critical network infrastructure. Today, Irth Solutions helps protect millions of consumers and workers as well as billions of dollars in infrastructure for global leaders like Shell, ExxonMobil, T-Mobile, Verizon, COX and more.

We are committed to our goal of responding to customer needs and continuing to innovate to deliver impactful solutions for the future.

But we can’t do it alone. We need dedicated, hard-working professionals to help us achieve our objectives.

We’re expanding our team so we can continue to deliver the products and services our customers expect. Are you interested? Start by applying below.

Voir tous les emplois chez Irth Solutions →

Emplois similaires

Symmetrio
Data Engineer
Symmetrio
⚡ Postuler tôt Denver, Colorado, United State... Sur site $115,000–$140,000
● Nouveau 👁 Vu ✓ Postulé il y a 4 h
Leading Path Consulting
Backend Developer (Data Engineer) with Python/Node - TS/SCI w/ Poly required
Leading Path Consulting
⚡ Postuler tôt Chantilly, Virginia, United St... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 4 h
Clay Labs
Software Engineer, Data Products & Platform
Clay Labs
⚡ Postuler tôt New York Hybride $170,000–$300,000
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
Accordion
Cloud DevOps Engineer, Data & Analytics
Accordion
⚡ Postuler tôt Atlanta; Boston; Charlotte; Ch... Hybride $144,000–$250,000
● Nouveau 👁 Vu ✓ Postulé il y a 6 h
Airbnb
Staff Software Engineer, Data Catalog
Airbnb
⚡ Postuler tôt USA Sur site $212,000–$265,000
● Nouveau 👁 Vu ✓ Postulé il y a 6 h
Roblox
Senior Machine Learning Engineer, 3D Data
Roblox
⚡ Postuler tôt San Mateo, CA, United States Sur site $243,290–$295,250
● Nouveau 👁 Vu ✓ Postulé il y a 6 h
Stoke Space
Frontend Software Engineer, Data Platform (Svelte + Rust )
Stoke Space
⚡ Postuler tôt Kent, Washington Sur site $154,350–$231,525
● Nouveau 👁 Vu ✓ Postulé il y a 6 h
Vast
Senior Data Engineer
Vast
⚡ Postuler tôt Long Beach, California, United... Sur site $143,000–$203,000
● Nouveau 👁 Vu ✓ Postulé il y a 7 h
Samsara
Senior Data Engineer II
Samsara
⚡ Postuler tôt Remote - US Hybride $134,470–$203,400
● Nouveau 👁 Vu ✓ Postulé il y a 7 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Irth Solutions

Voir tous les emplois chez Irth Solutions →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit