About this Sr. Staff TechOps & Support Engineer role at Sumerge
Our Sr. Staff TechOps & Support Engineer provides technical leadership for enterprise production operations by ensuring the stability, availability, and continuous improvement of mission-critical environments while driving automation, operational excellence, and service reliability across the organization.
Responsibilities
- Ensure high availability, performance, and reliability of production systems and platforms
- Drive incident response, root cause analysis, problem management, and continuous service improvement initiatives to enhance platform reliability and reduce operational risk.
- Architect and optimize CI/CD pipelines, deployment strategies, automation frameworks, monitoring, and observability solutions to improve operational efficiency.
- Provide technical leadership for IBM Cloud Pak Stacks, Confluent, Elasticsearch, container platforms, and enterprise infrastructure while supporting complex production deployments.
- Establish operational standards, security controls, backup, disaster recovery, and service reliability practices to ensure compliance and business continuity.
- Mentor engineers, provide technical direction, and collaborate with Development, Platform Engineering, Security, QUALITY, and Infrastructure teams to resolve critical production issues.
- Maintain technical governance, operational documentation, and platform standards while driving innovation and continuous operational improvement.
Requirements
- Bachelor's degree or Diploma in Computer Science, Engineering, or a related field.
- 6–8 years of experience in Technical Operations, Site Reliability Engineering (SRE), DevOps, Infrastructure Operations, or a similar technical discipline.
- Extensive experience with Linux/Windows administration, Kubernetes, OpenShift, Docker, CI/CD platforms, automation, scripting, monitoring, and observability tools (Grafana & Prometheus).
- Strong expertise in IBM Cloud Pak Stacks, Confluent, Elasticsearch, Databases (SQL, DB2, MongoDB, etc.), JVM performance analysis, networking, and enterprise production environments.
- Advanced knowledge of security best practices, IAM, backup and disaster recovery, scalability, performance tuning, incident management, and SLA/SLO management.
- Demonstrated ability to lead complex troubleshooting efforts, mentor engineering teams, and drive operational excellence across enterprise platforms.
- Excellent leadership, stakeholder management, communication, decision-making, and strategic problem-solving skills.
Benefits