Sobre esta vaga de System Engineer na Vultr
Who We Are
Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by David Aninowsky and self-funded for over a decade, Vultr has grown to become the world’s largest privately-held cloud infrastructure company.
Vultr Cares
Medical Insurance stipend paid annually
9 Company-Paid Holidays
Generous Leave Policy + 1 month paid sabbatical every 5 years + Anniversary Bonus each year
Professional Development Reimbursement
Internet reimbursement
Fitness membership reimbursement
Company paid Wellable subscription
Join Vultr
Vultr is expanding its India presence and is building its first Global Integrated Operations Command Center (GIOC) in Chennai — the 24x7 nerve center for monitoring, triage, and first response across Vultr’s global cloud, systems, network, and security operations.
We are seeking Systems Engineers to serve as the first line of response for Vultr’s systems estate — compute, storage, operating systems, and virtualization. This is a frontline operations role for engineers who thrive in a fast-paced, alert-driven environment and want to grow into deeper SRE, systems, and platform engineering careers. You must be comfortable working on a rotational shift model (24x7) — including nights, weekends, and holidays, to ensure we have full coverage.
Role Overview
The Systems Engineer is the entry point of the GIOC incident lifecycle. You monitor health signals across Vultr’s systems estate, acknowledge and triage alerts within defined first-response SLAs, execute documented runbooks to resolve common issues, and escalate cleanly to Senior Systems Engineer when an issue falls outside L1 scope. The role ensures fast, accurate first response that protects customer experience and keeps incidents moving toward resolution.
You operate from a consolidated single-pane-of-glass dashboard, classifying and routing incidents by severity and tower while owning accurate, complete documentation of every action. Success is measured by first-response SLA adherence, triage accuracy, and runbook resolution rate. The role runs on a rotational schedule to sustain 24x7 coverage.
Key Responsibilities
Alert Monitoring & First Response
Monitor systems-estate health signals (compute, storage, OS, virtualization) from a consolidated single-pane-of-glass dashboard
Acknowledge alerts and incidents within first-response SLA targets
Perform initial review — read the alert, check recent changes, and review the CMDB before acting
Triage & Severity Classification
Classify each incident by severity and customer impact, confirming the owning tower (Systems / Network / Security / cross-tower)
Route and assign tickets accurately to the correct tower or escalation path
Sustain triage accuracy at or above the GIOC go-live standard
Reassess severity as impact evolves, and declare higher when in doubt
Runbook-Driven Resolution
Match alert signatures to documented runbooks and execute permitted L1 actions
Resolve common, well-understood systems issues end to end within L1 scope
Document every step taken, with supporting evidence, in the incident ticket
Escalation & Handoff
Escalate unresolved or out-of-scope issues to Senior Engineers with a clean, complete handoff (symptom · evidence collected · steps tried · current state)
Flag missing or ineffective runbooks as runbook gaps for review and continuous improvement
Initiate and support major-incident (Sev-0 / Sev-1) bridges per the escalation matrix
Track escalations to closure and confirm service is restored before resolving the ticket
Operational Documentation & Communication
Maintain accurate shift logs, ticket updates, and incident timelines
Provide clear, timely status communication to stakeholders and customers per severity
Produce thorough shift-handover notes for seamless continuity
Continuous Improvement
Identify recurring alerts and noise as candidates for tuning and automation
Contribute to and improve runbooks and knowledge-base articles
Participate in post-incident reviews (PIR) and trend reviews
Qualifications & Experience
Graduate/Engineer in a relevant field (B.E./B.Tech, or equivalent)
3–5 years of experience in systems administration, preferably in NOC/SOC, or IT operations
Good understanding of Linux/Windows
Working knowledge of operating systems — processes, services, and log analysis (Linux/Windows)
Familiarity with compute, storage, and virtualization concepts — VMs, Hypervisors, and basic networking
Understanding of incident-management fundamentals and severity-based triage
Exposure to monitoring/observability tools and ticketing systems (alerting consoles, ITSM)
Strong written communication for accurate documentation and clean escalation handoffs; willingness and ability to work a rotational 24x7 shift model, including nights, weekends, and holidays
Proficient in English verbal and written communication
Preferred Qualifications
ITIL V4 Foundation certification or equivalent ITSM knowledge
Linux certification (RHCSA / LFCS) or a cloud-fundamentals certification
Scripting familiarity (Bash, Python, or PowerShell) for routine automation
Exposure to cloud platforms and GPU / high-density infrastructure
Prior working experience in a 24x7 NOC, SOC or Command-Center environment
Experience in using tools such as JIRA, Confluence, PagerDuty, and observability platforms
Inclusion & Privacy
We are an equal opportunity employer and are committed to creating an inclusive environment for all employees. We welcome applications from individuals of all backgrounds and experiences, and we prohibit discrimination based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected status under applicable laws. Vultr will consider qualified applicants with arrest or conviction records in accordance with applicable laws and will not conduct a background check until after an offer of employment has been extended and accepted.
We also take your privacy seriously. We handle personal information responsibly and follow applicable laws, including U.S. privacy rules and India’s Digital Personal Data Protection Act, 2023. Your data is used only for legitimate business purposes and is protected with proper security measures.
Where allowed by law, applicants may request details about the data we collect, access or delete their information, withdraw consent for its use, and opt out of nonessential communications. For more details, please see our Privacy Policy.