Sobre esta vaga de Shift Lead (Linux) na Keyloop
Role Summary
Key Responsibilities
Shift Leadership
- Lead and coordinate the shift team, assigning work and prioritising tickets, alerts, and incidents.
- Run structured shift handovers, ensuring open incidents, risks, and pending changes are clearly documented.
- Act as the escalation point for L1/L2 engineers and make timely decisions during live incidents.
- Mentor and coach team members; support onboarding, training, and skills development.
- Monitor dashboards and alerts, triage events, and drive incidents through detection, resolution, and closure.
- Manage major incidents: coordinate bridge calls, assign roles, provide regular updates to stakeholders, and escalate to engineering or management as needed.
- Contribute to post-incident reviews and root cause analysis, and track follow-up actions to completion.
- Perform first- and second-line troubleshooting on Linux servers (RHEL, CentOS, Ubuntu, Amazon Linux), covering CPU, memory, disk, network, and service issues.
- Manage users, permissions, services (systemd), cron jobs, log analysis, and package updates.
- Support patching, reboots, and routine maintenance using approved change procedures.
- Write and maintain Bash scripts and runbooks to automate repetitive operational tasks.
- Work with monitoring tools (e.g., CloudWatch, Nagios, Zabbix, Prometheus/Grafana, Datadog) to tune alerts and reduce noise.
- Follow ITIL practices for incident, change, and problem management using a ticketing tool (e.g., ServiceNow, Jira Service Management).
- Track SLA/KPI performance and report on shift activity, trends, and improvement opportunities.
- Maintain and improve runbooks, SOPs, and knowledge base articles.
- Identify automation opportunities and work with engineering teams to reduce manual toil.
- Ensure compliance with security, access, and change control policies.
Incident and Problem Management
Linux Operations
Monitoring and Service Management
Process and Continuous Improvement
Required Skills and Experience
- 7+ years in IT operations, NOC, or infrastructure support, including experience in a 24/7 environment.
- Prior experience as a shift lead, team lead, or senior operations engineer.
- Strong hands-on Linux administration skills: command line, file systems, LVM, networking (DNS, TCP/IP, firewalls), SSH, systemd, and log analysis.
- Solid troubleshooting skills with the ability to stay calm and decisive under pressure.
- Experience handling major incidents and running bridge calls.
- Working knowledge of ITIL processes (incident, problem, change management).
- Experience with monitoring and alerting tools and ticketing systems.
- Scripting skills in Bash and/or Python.
- Excellent verbal and written communication, including clear stakeholder updates and handover documentation.
- Willingness to work rotational shifts, nights, weekends, and holidays.