Jobs Companies CoorsTek Software Site Reliability Engineer

Über diese Software Site Reliability Engineer Stelle bei CoorsTek

CoorsTek · Vor Ort · Golden, CO

It's exciting to work for a company that makes the world measurably better.

We're committed to bringing safety, quality, and customer focus to the business of advanced ceramics manufacturing.

Job Title

Software Site Reliability Engineer

As the Site Reliability Engineer, you will support CoorsTek's Databricks application and data product strategy by ensuring solutions built, migrated, and deployed on Databricks are reliable, secure, observable, supportable, and cost-effective in production. This role is not solely focused on monitoring and operational support. In this role, you will actively develop automation, platform tooling, deployment pipelines, observability capabilities, and reliability solutions that reduce operational toil and improve the scalability of Databricks-hosted applications and data products.

This role sits within Data & Analytics and partners closely with Architecture, Cybersecurity, Infrastructure, Manufacturing IT/OT, Enterprise Applications, citizen developers, and business teams. As the Databricks Site Reliability Engineer, you will support production reliability for Databricks-hosted applications (pattern B), analytics products, workflows, and AI-enabled solutions.

In this role, you will help CoorsTek move quickly without creating unmanaged technical debt by contributing to and improving support patterns, monitoring standards, deployment practices, runbooks, incident response, and operational guardrails for Databricks solutions created by both business-enabled citizen development and IT delivery teams.

Roles and Responsibilities

  • Support production reliability, operational readiness, and lifecycle support for Databricks-hosted applications, data products, dashboards, notebooks, jobs, workflows, APIs, and AI-enabled solutions.

  • Support applications migrated to Databricks, built directly in Databricks, or promoted from citizen development and IT development into governed production patterns.

  • Execute intake, review, handoff, support, and release practices for Pattern B Databricks applications, including minimum requirements before production deployment.

  • Partner with citizen developers, IT developers, data engineers, enterprise architects, and business stakeholders to convert prototypes into reliable, monitored, documented, and supportable services.

  • Implement and maintain observability standards, including logging, alerting, health checks, SLIs/SLOs, lineage, usage monitoring, cost monitoring, and operational dashboards.

  • Respond to incidents, coordinate troubleshooting, participate in root cause analysis and support corrective actions for failed jobs, broken pipelines, access issues, performance issues, data refresh failures, and application outages.

  • Maintain and update runbooks, support procedures, escalation paths, ownership models, service catalogs, and knowledge articles for Databricks applications and data products.

  • Partner with Data & Analytics on Databricks workflows, Delta Lake, Unity Catalog, data lineage, permissions, SQL warehouses, jobs, clusters, serverless capabilities, and performance tuning.

  • Partner with Cybersecurity and Architecture to ensure Databricks solutions meet standards for identity, access, secrets management, logging, data classification, responsible AI, and least-privilege access.

  • Support CI/CD, testing, environment promotion, release controls, rollback procedures, and change management for Databricks applications and related Azure or integration components.

  • Identify recurring failure patterns and assist with automating manual support work, reducing operational toil, and creating reusable templates and standards.

  • Advise teams on production-ready design, including resiliency, scalability, maintainability, cost control, data quality checks, monitoring hooks, and clear ownership.

  • Collaborate with manufacturing, finance, supply chain, quality, and other business teams to understand impact, prioritize recovery, and maintain trust in critical Databricks-supported solutions.

  • Support governance for citizen-built solutions by ensuring business-created applications have appropriate documentation, testing evidence, security review, support model, and IT transition plan before broad use.

  • Monitor and problem solve service health, support metrics, incidents, problem records, platform risks, and improvement backlog items for Databricks applications and data products.

  • Design and develop automation, self-healing workflows, monitoring integrations, and operational tooling using Python and cloud-native technologies.

Job Requirements

Education:

  • Bachelor's degree in Computer Science, Information Technology, Data Engineering, Software Engineering, Systems Engineering, or a related field required.

  • Master's degree preferred.

Experience:

  • 5 or more years of progressive experience in site reliability engineering, data platform engineering, cloud operations, DevOps, software engineering, data engineering, or production application support.

  • 3 or more years supporting cloud, data, analytics, application, or platform services in production environments preferred.

  • Experience with Databricks, Delta Lake, Unity Catalog, SQL, Python, PySpark, notebooks, jobs/workflows, SQL warehouses, clusters, or lakehouse architecture.

  • Experience operating applications through incident management, problem management, change management, monitoring, release management, and production readiness practices.

  • Preferred experience with Azure, CI/CD pipelines, Git-based development, infrastructure patterns, logging, alerting, automation, and support runbooks.

  • Preferred experience supporting data pipelines, analytics products, dashboards, APIs, AI-enabled applications, or business-critical reporting environments.

Functional / Technical Knowledge, Skills & Abilities:

  • Strong understanding of SRE, DevOps, IT operations, and production support practices, including reliability, observability, automation, incident response, and operational excellence.

  • Working knowledge of Databricks platform capabilities, including Delta tables, notebooks, workflows/jobs, SQL, Unity Catalog, lineage, permissions, compute configuration, and governed access patterns.

  • Ability to troubleshoot Databricks jobs, pipelines, notebooks, SQL queries, permissions, data refreshes, performance issues, and environment or integration failures.

  • Ability to write and review SQL and Python; PySpark, scripting, API, and automation experience preferred.

  • Ability to define operational readiness standards for applications created by citizen developers, IT teams, consultants, and data engineering teams.

  • Strong understanding of monitoring, alerting, logging, service health, SLOs, runbooks, release controls, rollback planning, and root cause analysis.

  • Ability to balance speed, business enablement, cybersecurity, supportability, cost control, and long-term platform sustainability.

  • Ability to partner effectively with Data & Analytics, Cybersecurity, Architecture, Infrastructure, Enterprise Applications, Manufacturing IT/OT, and business stakeholders.

  • Strong documentation and communication skills, including support models, knowledge articles, architecture notes, production checklists, escalation paths, and operational dashboards.

  • Ability to manage multiple production priorities, operate calmly during incidents, drive follow-through on corrective actions, and influence teams without direct authority.

Preferred Certifications:

  • Relevant Databricks certifications, including Data Engineer, Data Analyst, Machine Learning, or Lakehouse Fundamentals preferred.

  • Relevant Microsoft Azure, DevOps, cloud engineering, cybersecurity, ITIL, SRE, observability, or data engineering certifications preferred.

  • ITIL Foundation, Azure Administrator, Azure Developer, GitHub, Terraform, Kubernetes, or related platform operations certifications are a plus.

Additional Position Details:

  • Location: Golden, CO (on-site)

  • Work Authorization: Requires U.S. Person status (U.S. citizen, a Green Card holder, or a protected refugee/asylee)

Target Hiring Range

Annual Salary: USD 115,000.00 - USD 155,000.00

Actual compensation is commensurate with experience, skills and education. CoorsTek strives to give all qualified applicants equal opportunity and to make selection decisions on job related factors. Do not provide any information on the application which will indicate your race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity, pregnancy, genetic information, veteran status, or any other status protected by law or regulation.

If you like working for a company that makes a real difference in the world, you'll enjoy your career with us!

Bereit, sich bei CoorsTek zu bewerben?
Bei CoorsTek bewerben

Wie sich dieses Gehalt für SRE vergleicht

Diese Stelle zahlt $135,000/yrim Einklang mit der üblichen Spanne für SRE Stellen.

$96,226 dem Median $155,054 $228,739

Übliche Spanne $121,000–$193,000/yr, aus 816 vergleichbaren SRE Anzeigen auf JobsRadar (Vergütung auf USD hochgerechnet). Gehaltseinblicke für SRE ansehen →

Über CoorsTek

About Us: Making the world measurably better...one person at a time. CoorsTek products and components touch people’s lives in unexpected, awe-inspiring ways. Though most people may not realize it, CoorsTek components enable technological solutions across virtually every industry—from simple items like beverage dispensers and light switches to high-tech items like cell phones and satellites. Chances are, you are using a CoorsTek-enabled technology every day without even knowing it! We create these amazing solutions by leveraging our expertise in materials science, deploying best-in-class R&D capabilities, and by creating efficient, safe work environments. Most importantly, we employ dedicated

Alle Jobs bei CoorsTek ansehen →

Ähnliche Jobs

Loftorbital
Site Reliability Engineer (Network)
Loftorbital
⚡ Früh bewerben San Francisco, CA Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
Tailor
SRE (Site Reliability Engineer)
Tailor
⚡ Früh bewerben Tokyo · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 7 Std.
Okta
Manager, Site Reliability Engineering
Okta
⚡ Früh bewerben Bellevue, Washington; Chicago,... Vor Ort $204,000–$306,000
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Okta
Manager, Site Reliability Engineering (Auth0)
Okta
⚡ Früh bewerben New York, New York; Washington... Vor Ort $182,000–$250,800
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Okta
Manager- Site Reliability Engineer
Okta
⚡ Früh bewerben Bengaluru, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Loftorbital
IT Platform Engineer
Loftorbital
⚡ Früh bewerben Golden, CO Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Loftorbital
Senior Cloud Infrastructure Engineer
Loftorbital
⚡ Früh bewerben Golden, CO or San Francisco Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 3 Wo.
CoorsTek
Sr Network Administrator
CoorsTek
⚡ Früh bewerben Golden, CO Vor Ort $112,767–$155,055
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
CoorsTek
Cloud Engineer II
CoorsTek
⚡ Früh bewerben Golden, CO Vor Ort $103,040–$136,013
● Neu 👁 Gesehen ✓ Beworben vor 2 Mon.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei CoorsTek

Alle Jobs bei CoorsTek ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos