Jobs › Companies › Smartsheet › Senior Software Engineer I (Observability Platform)

Sobre esta vaga de Senior Software Engineer I (Observability Platform) na Smartsheet

Smartsheet · Remoto · -REMOTE, USA-

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.

Smartsheet’s Observability Engineering team owns how the company sees itself: the collection, modeling, storage, and analysis of metrics, logs, distributed traces, and events across a global, multi-region infrastructure. As a Senior Software Engineer I on this team, you will be hands-on-keyboard building that platform (instrumentation libraries, telemetry pipelines, SLO and alerting systems, dashboards as code), and you will be the engineer who connects it to everything around it. Observability only pays off when signals flow into the systems where engineers actually work: CI/CD, the service catalog, incident management, ticketing, chat, and the automated remediation that closes the loop without waking anyone up.

We run observability as an internal product, and the engineers of Smartsheet are its users. That means published interfaces, versioned client libraries, honest deprecation paths, and SLOs on the platform itself. It also means we measure ourselves on adoption rather than on components shipped: a beautifully engineered tracing pipeline that three teams use is a failure. You will own a slice of that product end to end, including the golden path that makes correct instrumentation the easy choice, the documentation and hands-on sessions that get teams there, and the numbers that tell you which part to fix next.

This is a build role, not a configure role, and it is engineering first. You will write the services, the integrations, and the automation that turn a collection of separate tools into one coherent platform, and you will own the architecture decisions that make it hold up as it grows. Where tooling and teaching compete for the same problem, we prefer the tool: a CI check that rejects an unlabeled metric works for every team forever, while a workshop works for the people in the room. You will report to the Team Lead, Observability Engineering, and partner closely with the Principal Engineer setting telemetry data platform direction.


You Will

  • Architect and build the observability platform: Design and ship end-to-end capability for metrics, logs, distributed traces, and events across multiple regions and environments, owning components from collection through storage, query, and presentation
  • Run observability as an internal product: Build for the engineers who depend on you, with published interfaces, versioned client libraries, clear deprecation paths, and SLOs on the platform itself, so teams rely on it the way they rely on any production service
  • Drive instrumentation with OpenTelemetry: Build and maintain shared instrumentation libraries, collector deployments, semantic conventions, context propagation, and sampling strategies so service teams get correlated signals by default rather than by effort
  • Engineer telemetry pipelines at scale: Build high-volume collection, enrichment, redaction, and routing pipelines with the reliability, backpressure handling, and tiered retention that multi-region, high-cardinality traffic requires
  • Connect the platform end to end: Integrate the observability toolchain with CI/CD, service catalog, incident management, ticketing, chat, and feature-flag systems through REST APIs, webhooks, and event-driven services, so that a deploy, an alert, an incident, and a ticket form one continuous thread instead of four disconnected ones
  • Build automated and self-healing remediation: Design event-driven and agent-assisted workflows that detect, diagnose, and resolve known failure modes automatically, with human-in-the-loop approval gates as the safety mechanism for anything consequential
  • Make reliability measurable: Implement SLOs, error budgets, and golden-signal alerting as code, and drive down alert noise so that a page means something is genuinely wrong
  • Build the golden path: Own the paved road for instrumenting a new service, including scaffolding and templates, versioned Terraform modules, and CI checks that catch missing or malformed telemetry before merge. Make the correct path the easy path, so coverage comes from good defaults rather than from chasing teams
  • Instrument the platform itself: Track coverage, onboarding time, time to first useful dashboard, query performance, and cost per service, and build the guardrails that keep cardinality, sampling, and retention proportional to the value of the signal. Let those numbers decide what you build next instead of guessing
  • Build the enablement layer: Write documentation as code, reference architectures, and worked examples that scale past the conversations you can personally have, build the onboarding path that takes a team from zero to instrumented without a meeting, and run the workshops, office hours, and game days that exercise dashboards and alerts under realistic failure
  • Raise the technical bar: Lead code reviews and architecture discussions, author the instrumentation standards other teams build against, mentor engineers on signal design and cost-aware instrumentation, and grow a group of instrumentation champions who carry the practice inside their own teams
  • Apply AI where it earns its place: Use AI tooling to improve your own and the team’s efficiency across coding, testing, design, and troubleshooting, and help instrument Smartsheet’s AI and agentic systems so their behavior is as observable as any other service
  • Turn incidents into durable improvements: Join the team’s on-call rotation, drive root-cause analysis, and close every incident with an instrumentation change, an automation, or a documented lesson that reaches the teams who need it

You Have

  • 5+ years building and operating highly scalable, highly available distributed systems, platform services, or observability infrastructure
  • 5+ years programming in Go, Python, Java/Kotlin, or TypeScript/Node.js, with the ability to move fluently between backend services and the operational tooling around them
  • Experience building internal platforms, developer tooling, or an internal developer platform (e.g., Backstage), with a product mindset toward internal users
  • Hands-on depth across all three primary signals (metrics, logs, and distributed tracing) on a major commercial or open-source observability platform, including practical understanding of cardinality, sampling, and cost mechanics, and experience defining SLOs and alerting for production systems
  • Practical OpenTelemetry experience: collectors, instrumentation (auto and manual), semantic conventions, and context propagation across service boundaries
  • Experience building high-volume data or telemetry pipelines: log and metric shippers, streaming transport, transformation and enrichment, and search or time-series backends
  • Strong REST API and integration engineering skills, including OAuth and other authentication patterns, webhooks, and event-driven architectures connecting third-party SaaS platforms
  • Advanced AWS and Kubernetes expertise (EKS, ECS, EC2, Lambda, and the managed messaging and eventing services that tie them together), delivered through Terraform and GitOps-based workflows
  • Demonstrated technical teaching: workshops, onboarding curricula, internal courses, conference or meetup talks, or a track record of raising a team’s capability. You write things down, and what you write gets used
  • Experience measuring and driving adoption of a platform or tool, and comfort treating low adoption as a product problem rather than a user problem
  • 6+ months of professional experience leveraging AI to enhance engineering productivity, with a view on where it helps and where it does not
  • A track record of leading large projects autonomously: decomposing ambiguous problems, planning the work, and carrying it to production
  • Computer Science degree, Engineering degree, or equivalent practical experience
  • Legal eligibility to work in the U.S. on an ongoing basis

Nice to Have

  • Workflow automation or orchestration experience (managed state machines, workflow engines, or RPA platforms) with human-in-the-loop approval patterns
  • Experience instrumenting LLM or agentic systems, including token, latency, quality, and cost telemetry
  • Experience migrating or consolidating overlapping observability tooling without a coverage gap in between, especially where the hard part was the people rather than the technology
  • Docs-as-code toolchains and internal developer portals, including the discipline of keeping generated reference material accurate as the platform changes
  • Community building or public speaking: conference talks, meetup organizing, or an internal guild you started and kept alive
  • Multi-region and data residency experience, including telemetry redaction and handling of sensitive data across jurisdictions
  • Regulated or government environment experience (FedRAMP, GovCloud) and the instrumentation constraints that come with it
  • Prometheus, Grafana, and Alertmanager, or comparable open-source observability tooling
  • Frontend experience building operational UIs, dashboards, or internal tools

 

Current US Perks & Benefits:

  • Employer subsidized medical/vision and dental coverage for full-time employees
  • 401k Match to help you save for your future (50% of your contribution up to the first 6% of your eligible pay)
  • Monthly stipend to support your work and productivity
  • Flexible Time Away Program, plus Sick Time Off
  • US employees are automatically covered under Smartsheet-sponsored life insurance, short-term, and long-term disability plans
  • US employees receive 12 paid holidays per year
  • Up to 24 weeks of Parental Leave
  • Personal paid Volunteer Day to support our community
  • Opportunities for professional growth and development including access to Udemy online courses
  • Company Funded Perks, including a counseling membership, local retail discounts, and your own personal Smartsheet account
  • Teleworking options from any registered location in the U.S. (role specific)

Smartsheet provides a competitive base salary range for roles that may be hired in different geographic areas we are licensed to operate our business from. Actual compensation is determined by several factors including, but not limited to, level of professional, educational experience, skills, and specific candidate location. In addition, this role will be eligible for a market competitive incentive opportunity.

US Base Salary Pay Range
$161,250—$193,750 USD

 

Get to Know Us:

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.

Equal Opportunity Employer:

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information. 

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.

 

#LI-Remote

Pronto para se candidatar à Smartsheet?
Candidatar-se à Smartsheet

Como este salário de Software Engineer se compara

Esta vaga paga $177,500/yr — em linha com da faixa típica para vagas de Software Engineer.

$117,000 a mediana $168,938 $238,275

Faixa típica $140,000–$204,363/yr, com base em 571 vagas de Software Engineer comparáveis na JobsRadar (pagamento anualizado em USD). Ver insights salariais de Software Engineer →

Vagas semelhantes

Aeratechnology
Senior Software Engineer (CALC engine)
Aeratechnology
⚡ Candidate-se cedo San Francisco or Mountain View... Híbrido
● Nova 👁 Vista ✓ Candidatada há 4h
HU
Senior Software Engineer
Hudl
⚡ Candidate-se cedo Lexington, KY, United States;... Híbrido $112,000–$187,000
● Nova 👁 Vista ✓ Candidatada há 6h
Zscaler
Principal Software Development Engineer (Control-Plane)
Zscaler
⚡ Candidate-se cedo San Jose, California, USA; San... Híbrido $185,500–$265,000
● Nova 👁 Vista ✓ Candidatada há 6h
Zscaler
Staff Software Development Engineer (Microservices)
Zscaler
⚡ Candidate-se cedo San Jose, California, USA Híbrido $133,000–$190,000
● Nova 👁 Vista ✓ Candidatada há 6h
Wayve
Senior Software Engineer - Runtime Platform, Robot Software
Wayve
⚡ Candidate-se cedo London Presencial
● Nova 👁 Vista ✓ Candidatada há 7h
mthree Recruiting Portal
Junior Software Engineer
mthree Recruiting Portal
⚡ Candidate-se cedo USA Presencial $48,000–$66,000
● Nova 👁 Vista ✓ Candidatada há 7h
mthree Recruiting Portal
Junior Software Engineer
mthree Recruiting Portal
⚡ Candidate-se cedo USA Presencial $48,000–$66,300
● Nova 👁 Vista ✓ Candidatada há 7h
mthree Recruiting Portal
Junior Software Engineer
mthree Recruiting Portal
⚡ Candidate-se cedo USA Presencial $61,000–$67,000
● Nova 👁 Vista ✓ Candidatada há 7h
New Relic
Software Engineer - Container Fabric
New Relic
⚡ Candidate-se cedo Portland, Oregon, USA Presencial $126,000–$158,000
● Nova 👁 Vista ✓ Candidatada há 7h

Cadastre-se para receber sugestões sob medida com base nas vagas que você abre e nas buscas que você salva.

Mais vagas na Smartsheet

Ver todas as vagas na Smartsheet →

Candidatar-se agora
🤖

Opa — calma aí

A JobsRadar foi feita para pessoas de verdade passando por um momento difícil na busca por emprego — não para requisições automatizadas. Você está clicando rápido demais e agora está temporariamente bloqueado.

Volte mais tarde. Se você está mesmo procurando emprego, estamos com você — apenas aja como um ser humano.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Ganhe vantagem na sua busca por emprego.

Entre no nosso canal do Telegram para o que ajuda você a conseguir a vaga — referências salariais, o pulso semanal do mercado e avisos de novos recursos. Sem spam, só sinal.

Entre no canal — é grátis