Sobre esta vaga de Sr. Site Reliability Engineer na Filevine
embedded with a cross-functional team — and a primary driver of Observability excellence. You
bring broad, deep SRE expertise across the full software development lifecycle, and you bring
mastery in making complex systems visible, understandable, and actionable through best-in-class
monitoring, alerting, incident management, and platform tooling.
You will establish and own the observability posture for your team’s systems, serve as an
authoritative voice for reliability, and invest meaningfully in the engineers around you. As you ramp
up and gain context in the Filevine environment, you’ll grow into increasingly impactful work —
taking on mission-critical objectives that build out and improve our autonomous, observable
systems while cementing your reputation as an exceptional engineer who designs and maintains
systems that perform reliably at enterprise scale.
Responsibilities
monitoring, alerting, dashboards, distributed tracing, log aggregation, and SLI/SLO/SLA
frameworks.
blameless post-mortems that drive lasting improvements.
leadership, sound judgment, and a clear voice on reliability across the SDLC.
Filevine products with minimal human intervention.
toil and accelerate resolution time.
defending overall security posture.
knowledge proactively, and leaving people more capable than before.
documentation for the technologies in your domain and actively close gaps for fellow SREs.
communicate clearly with technical and management stakeholders at all levels.
for ongoing improvement.
Observability & Incident Management
candidate will bring hands-on, expert-level experience in these areas:
platform (e.g., Datadog, Dynatrace, Grafana/Prometheus) — including dashboarding, alerting,
APM, infrastructure monitoring, and log management.
• Proven ability to design and implement comprehensive observability strategies: defining
meaningful SLIs and SLOs, building actionable alert hierarchies, and instrumenting services
for full-stack visibility.
• Strong incident management discipline: structured on-call practices, runbooks, escalation
paths, stakeholder communication under pressure, and rigorous post-incident review.
• Experience integrating observability tooling into CI/CD pipelines to surface reliability signals
earlier in the development lifecycle.
• Ability to translate complex system behavior into clear, consumable signals for both technical
teams and non-technical stakeholders
Qualifications
operations roles, including a minimum of 5 years dedicated to Site Reliability Engineering.• Expert-level, well-rounded SRE skill set — proficient across monitoring/alerting, incident
response, capacity planning, performance optimization, CI/CD, and reliability engineering best
practices.
• Deep hands-on expertise with New Relic or a comparable observability platform; strong
preference for candidates who have led observability platform adoption or migration at scale.
• Demonstrated experience owning incident management programs: on-call processes,
escalation design, post-mortem culture, and measurable MTTR/MTTD improvement.
• Strong proficiency in Python, Bash, PowerShell, and other common SRE scripting and
automation technologies.
• Expert-level experience designing, building, and maintaining autonomous systems that handle
software build, deployment, testing, monitoring, and operations.
• Proficient hands-on experience with AWS (EC2, EKS/Kubernetes, CloudWatch, Lambda, S3,
IAM) and the broader cloud-native ecosystem.
• Strong communicator who proactively informs stakeholders, operates transparently, and can
bridge technical complexity for product and management audiences.
• Proven track record of mentoring engineers, leading initiatives to completion, and making
those around them measurably better.
• Bachelor's degree in Computer Science, Information Systems, or a related field; equivalent
certifications (e.g., AWS certifications, Google Cloud Professional); or substantial comparable
direct work experience