Über diese NOC & Application Monitoring Manager Stelle bei Envision Employment Solutions
Envision Employment Solutions is currently looking for a NOC & Application Monitoring Manager for one of our partners, a leading Digital Bank!
Job Summary
The NOC & Application Monitoring Manager leads the Bank's 24/7 Network Operations Centre and application monitoring function, ensuring real-time visibility of system availability and performance across all production applications and infrastructure. This role owns incident detection, triage, and escalation, and drives the adoption of observability and Site Reliability Engineering (SRE) practices to strengthen system resilience.
Responsibilities
- Manage the 24/7 monitoring of all production systems, applications, and infrastructure using observability tooling.
- Lead and manage the Network Operations Centre (NOC) team to ensure real-time visibility of system availability and performance.
- Detect, triage, and escalate incidents in line with defined operational playbooks and runbooks, executing first-line response actions.
- Produce daily and weekly operational health reports, tracking Service Level Objective (SLO) breaches and recurring incident patterns.
- Establish and operate observability and Site Reliability Engineering (SRE) practices across applications and infrastructure.
- Drive root cause analysis and continuous improvement initiatives to enhance system reliability and reduce incident recurrence.
- Define and maintain monitoring dashboards, alerting thresholds, and escalation matrices in coordination with application and infrastructure teams.
- Manage relationships with monitoring tool vendors and ensure optimal use of observability platforms.
- Coach and develop the NOC team, ensuring appropriate shift coverage and readiness for major incidents.
Requirements
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- Minimum of 7 years of experience in NOC operations, application monitoring, or Site Reliability Engineering, ideally within banking or financial services.
- Hands-on experience with observability and monitoring tools (e.g., Splunk, Datadog, Dynatrace, Prometheus/Grafana, or similar).
- Strong understanding of ITIL incident and problem management practices.
- Experience managing 24/7 shift-based operations teams.
- Solid understanding of infrastructure, application, and data monitoring across hybrid/on-prem and cloud environments.
- Strong analytical, troubleshooting, and reporting skills.
- Excellent communication skills, with the ability to manage escalations under pressure.