À propos de ce poste Systems Development Engineer chez Union.ai
About Us
At Union, we are solving one of the hardest challenges in AI infrastructure today: enabling high-velocity iteration while maintaining seamless production-readiness for AI workloads at scale.
Flyte, the open-source project we steward, has emerged as the modern standard for data and AI orchestration, and is trusted by leading technology organizations including LinkedIn, Stripe, and Wayve to run millions of mission-critical workflows on the platform. These workflows comprise data preparation, model training, and scaled inference spanning thousands of GPUs, all major clouds, and on-premise infrastructure.
We have a technical founding team who created Flyte while at Lyft, a deep bench of infrastructure experts from top companies, and have raised from top investors like NEA and Nava Ventures.
About the Role
We are hiring a Systems Development Engineer to improve the reliability, operability, and customer experience of our production platform. This is a 50/50 operations and engineering role. Part of your time will be spent investigating customer-impacting production issues, and part will be spent building the tools, automation, system design, and engineering practices that prevent those issues from recurring.
For this role, production is the customer. You will work from real operational signals: customer issues, incidents, on-call pages, recurring support patterns, and gaps in observability or automation. You will sit in engineering and partner with customer-facing teams to turn those signals into durable platform improvements.
This is not a traditional support role. It is a systems engineering role for someone who can debug deeply, communicate clearly, and write software that reduces operational load. The full engineering team backs you on on-call.
This role is hybrid, based out of our Seattle office.
What You'll Do
Investigate and resolve customer-impacting production issues across cloud infrastructure, workflow execution, access control, storage, networking, deployment systems, and observability.
Identify patterns in customer issues and convert them into automation, product improvements, runbooks, tests, or design changes.
Build internal tools and diagnostics that make production issues easier to detect, understand, and resolve.
Improve platform observability, including logs, metrics, dashboards, alerts, and customer-visible debugging information.
Participate in design and development so systems are easier to operate, debug, and support, and guide engineering teams toward durable fixes.
Define and uphold operational engineering practices: production readiness, alert quality, runbook discipline, observability standards, regression prevention, and code quality.
Drive measurable reductions in on-call pages, recurring customer issues, manual operational work, and time-to-resolution.
What We're Looking For
Strong software engineering skills in Python, Go, Java, or a similar language.
Experience debugging production systems across multiple layers of the stack.
Practical knowledge of Kubernetes, Linux, cloud infrastructure, distributed systems, networking, storage, and IAM.
Experience with infrastructure-as-code, deployment systems, CI/CD, observability, and operational automation.
Ability to move from ambiguous customer symptoms to clear technical diagnosis and durable remediation.
Strong judgment about when to fix directly, automate, escalate, redesign, or build a broader platform improvement.
Clear written and verbal communication, especially around root cause analysis, technical recommendations, runbooks, and design feedback.
A bias toward reducing toil through engineering rather than repeatedly solving the same issue by hand.
Preferred Experience
Operating customer-facing SaaS, cloud infrastructure, self-hosted or on-prem deployments, or workflow orchestration systems.
Batch workloads, autoscaling, capacity management, identity and access systems, storage systems, or platform observability.
Improving on-call health, reducing ticket volume, or building production diagnostics.
Working across support, customer success, product, and engineering teams.
Success Looks Like
Customer issues are diagnosed faster and recur less often.
Engineering teams receive actionable feedback from production and customer pain.
Common operational problems become automated workflows, better diagnostics, clearer runbooks, or product fixes.
On-call pages trend toward roughly one per month.
The platform becomes easier to operate, easier to debug, and safer to change.
Customers experience fewer production surprises and faster resolution when issues do happen.
Benefits & Belonging
At Union.ai we know that employees who feel their best can build amazing things and we are proud to offer best in class benefits that will continually evolve and grow as the needs of our employees do. Benefits may vary based on country.
Excellent medical - We pay 100% of your premiums and 90% for your dependents
Generous dental and vision plans- We pay 90% of the premiums for you and your dependents
Meaningful equity in the form of options – all employees are owners here
Unlimited time off + 12 company holidays
401K match - Union.ai matches 100% of contributions up to the first 3%, and 50% up to 5%
12 weeks paid parental leave for primary and secondary caregivers
Flexible work schedule (some restrictions apply)
For in office employees: Lunch provided onsite and well stocked kitchen with snacks and drinks.
We believe that our differences are what bring us together to achieve truly special outcomes. We strive to be inclusive and focus on building teams that embody that quality too. Union.ai is an equal-opportunity employer and we encourage you to apply, even if your experience doesn’t align exactly with our job description.