À propos de ce poste Cloud Infrastructure & Platform Specialist (Databricks + AWS + Azure) (Mid-Level chez Irth Solutions
About Irth Solutions
Irth Solutions is a leading provider of cloud-based SaaS software for damage prevention, asset integrity, stakeholder engagement and land management, helping energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. With nearly three decades of industry experience, Irth serves customers across North America and continues to expand its platform with new data-driven and AI-powered capabilities.
Cloud Infrastructure & Platform Specialist – Insights (AI/ML)
Location: Remote – India
Department: Insights (AI/ML) – Platform Engineering
Reports to: Data Platform & Analytics Manager
Experience: 3–5 years
About the Role
Irth is building a governed, multi-cloud data estate on Databricks to power cross-product insights across Damage Prevention, Asset Integrity Pipeline, Land Management, and Stakeholder Management.
We are looking for a Cloud Infrastructure & Platform Specialist to own and operate the underlying cloud and Databricks infrastructure that enables our Data Engineering, Data Science, and MLOps teams to build and deliver data and AI solutions.
You will be the go-to engineer for infrastructure provisioning, identity and access management, Databricks administration, networking, security, and platform onboarding. You will help provision workspaces and catalogs, manage IAM roles and administrative access, establish secure platform guardrails, and resolve infrastructure bottlenecks so engineering teams can onboard and ship efficiently.
This role is ideal for a mid-level cloud/platform engineer who enjoys solving infrastructure problems, automating repetitive tasks, and serving as the technical enabler for other engineering teams.
Key Responsibilities
1. Cloud Infrastructure Provisioning & Management
- Provision, configure, and manage core cloud infrastructure across AWS and Azure supporting Irth’s Databricks-based data estate.
- Manage infrastructure components including:
- VPCs / VNets
- Subnets
- Route tables and network connectivity
- S3 buckets / Azure Storage Accounts
- Compute resources
- Private endpoints and related networking services
- Establish and maintain secure DEV, QA, and PROD environments with appropriate isolation and platform guardrails.
- Support cloud landing-zone implementation and ensure environments follow established security, governance, and cost policies.
- Use Infrastructure-as-Code (IaC) technologies such as:
- Terraform
- Azure ARM/Bicep
- AWS CloudFormation
- Version, review, test, and deploy infrastructure changes through controlled engineering workflows.
- Troubleshoot infrastructure, networking, and cloud-service issues affecting data and ML workloads.
2. Identity & Access Management (IAM)
- Provision and manage IAM roles, service principals, managed identities, groups, and access policies across AWS, Azure, and Databricks.
- Implement least-privilege access and role-based access control (RBAC) aligned with enterprise security standards.
- Configure and support identity federation and automated identity lifecycle management using technologies such as:
- SSO
- SCIM
- SAML
- OAuth
- Manage access across cloud resources, Databricks workspaces, Unity Catalog, storage, compute, and other platform services.
- Review and process elevated or administrative access requests through controlled and auditable procedures.
- Balance engineering-team productivity with security, compliance, and least-privilege requirements.
- Participate in periodic access reviews and remediation of excessive or inappropriate permissions.
3. Databricks Platform Administration & Team Onboarding
- Create, configure, and administer Databricks workspaces and associated platform resources.
- Configure and manage Unity Catalog metastores, catalogs, schemas, storage credentials, external locations, and permissions.
- Support onboarding of Data Engineering, Data Science, and MLOps projects by establishing the appropriate platform resources and access.
- Troubleshoot onboarding blockers, including:
- Missing administrative privileges
- Cluster-policy conflicts
- Workspace configuration issues
- Catalog and schema permission problems
- Storage-access failures
- Identity and authentication issues
- Define and maintain Databricks cluster policies to control compute configurations, security, and cost.
- Configure instance profiles, managed identities, storage credentials, and other mechanisms required for secure data access.
- Establish workspace-level guardrails for compute, networking, storage, and user access.
- Partner with Data Engineering, Data Science, and MLOps teams to ensure new projects can be onboarded efficiently with the appropriate catalogs, schemas, permissions, and infrastructure already available.
- Act as the primary escalation point for Databricks platform and infrastructure issues.
4. Security, Networking & Compliance
- Implement secure private connectivity between Databricks, cloud services, and enterprise resources.
- Configure and troubleshoot technologies such as:
- AWS PrivateLink / VPC endpoints
- Azure Private Link / private endpoints
- VNet/VPC peering
- Secure cluster connectivity
- Network security groups and firewall controls
- Egress restrictions
- Implement encryption and key-management practices using:
- AWS KMS
- AWS Secrets Manager
- Azure Key Vault
- Equivalent enterprise security services
- Establish appropriate network segmentation and isolation between environments.
- Support security reviews and enterprise compliance requirements.
- Maintain audit logging and provide evidence for:
- Access reviews
- Administrative activity
- Infrastructure changes
- Data access
- Retention controls
- Security and governance reviews
- Help ensure infrastructure and platform configurations align with applicable regulatory and organizational requirements.
5. Cost Management, Automation & Operations
- Implement and enforce cloud and Databricks tagging standards across infrastructure and workloads.
- Monitor infrastructure and Databricks spend and identify opportunities for cost optimization.
- Support showback/chargeback initiatives and cost reporting.
- Implement resource policies and controls that prevent unnecessary or unapproved cloud and Databricks spend.
- Automate recurring infrastructure and access-management activities, including:
- Workspace provisioning
- Catalog and schema setup
- IAM role assignment
- Access provisioning
- Environment configuration
- Standard platform onboarding
- Build reusable automation that reduces manual provisioning and improves onboarding turnaround time.
- Monitor platform health and respond to infrastructure and access-related incidents.
- Troubleshoot cloud, networking, identity, and Databricks issues affecting production workloads.
- Maintain clear operational runbooks, architecture documentation, access procedures, troubleshooting guides, and onboarding documentation.
Role Outcomes
Success in this role means that Irth’s data and ML teams can rely on a secure, available, well-governed, and easy-to-use platform.
Key outcomes include:
- Fast and repeatable provisioning of cloud and Databricks infrastructure.
- Efficient onboarding of Data Engineering, Data Science, and MLOps teams.
- Secure, auditable, least-privilege access across AWS, Azure, and Databricks.
- Stable and well-governed DEV, QA, and PROD environments.
- Reduced infrastructure and platform-related blockers for engineering teams.
- Increased automation of provisioning and access-management processes.
- Strong adherence to networking, security, governance, and compliance standards.
- Improved visibility and control over cloud and Databricks costs.
- Well-documented infrastructure and operational procedures.
- A platform that enables engineering teams to build and ship without unnecessary infrastructure friction.
Requirements
Qualifications
Required Qualifications
- 3–5 years of experience in cloud infrastructure, platform engineering, cloud administration, or a closely related role.
- Hands-on experience working with both AWS and Microsoft Azure, including:
- IAM and role provisioning
- Cloud networking
- Storage services
- Compute and platform resources
- Hands-on experience administering Databricks, including:
- Workspace provisioning and configuration
- Unity Catalog
- Catalogs and schemas
- Metastores
- Cluster policies
- Workspace and data-access controls
- Working knowledge of Infrastructure-as-Code (IaC) tools, with Terraform preferred. Experience with Azure ARM/Bicep or AWS CloudFormation is also valuable.
- Strong understanding of Identity and Access Management (IAM) concepts, including:
- RBAC
- Least-privilege access
- Service principals
- Managed identities
- SSO
- SCIM
- SAML
- Familiarity with cloud networking and security fundamentals, including:
- VPC/VNet architecture
- Subnets and routing
- Private connectivity
- Network segmentation
- Secrets management
- Encryption
- Key management
- Ability to quickly troubleshoot infrastructure, networking, identity, permission, and platform-access issues that may block engineering teams.
- Strong communication skills, with the ability to explain infrastructure and access issues clearly to Data Engineering, Data Science, MLOps, and other technical teams.
- Familiarity with enterprise compliance and security frameworks, including SOC 2, ISO 27001, and GDPR.
- Understanding of audit logging, access reviews, administrative activity tracking, and evidence collection.
Preferred Qualifications
- Experience supporting multiple engineering teams—including Data Engineering, Data Science, and MLOps—with workspace, catalog, schema, compute, and access provisioning.
- Experience with infrastructure and platform CI/CD, including:
- GitHub Actions
- Databricks Asset Bundles (DABs)
- Infrastructure deployment automation
- DEV → QA → PROD environment promotion
- Knowledge of FinOps and cloud cost-management practices, including:
- Resource tagging
- Budgets
- Cost monitoring
- Showback/chargeback
- Cost anomaly detection and alerting
- Experience implementing cloud and Databricks policies that balance security, reliability, developer productivity, and cost efficiency.
- Relevant industry certifications, such as:
- AWS Solutions Architect
- Microsoft Azure Solutions Architect
- Azure Administrator
- Databricks Platform Administrator
- Equivalent cloud or platform certifications
Nice-to-Have Qualifications
- Experience onboarding Asset Integrity Pipeline data, including inspection, risk, maintenance, and compliance datasets, into a governed enterprise data estate.
- Experience designing catalogs, schemas, storage structures, and access controls for asset-integrity or infrastructure datasets.
- Experience provisioning infrastructure and access for utility, energy, oil & gas, pipeline, or asset-integrity data sources with elevated security or compliance requirements.
- Familiarity with GIS and geospatial data storage, networking, access, and platform-integration patterns relevant to pipeline and infrastructure datasets.
- Understanding of regulatory and compliance reporting requirements associated with pipeline, utility, energy, or asset-integrity data.
- Experience working with infrastructure and data platforms where data residency, auditability, security, and controlled access are critical requirements.
Benefits
Benefits
- Competitive Salary – A competitive compensation package based on experience and qualifications.
- Medical, Dental, and Vision Insurance – Comprehensive insurance coverage to support you and your family.
- 401(k) Plan with Company Match.
- Generous Paid Time Off (PTO) – Time off to support work-life balance and personal needs.
- Company-Paid Holidays – Paid holidays throughout the year.
- Flexible Work Options – Work-from-home opportunities are available, depending on role and business needs.
- On-Call Compensation – Additional pay for eligible on-call shifts.