About this Staff TechOps & Support Engineer role at Sumerge
Our Staff TechOps & Support Engineer Ensures the stability, performance, and availability of production systems by managing infrastructure, monitoring environments, responding to incidents, automating operational tasks, supporting deployments, and maintaining security, backups, and documentation
Responsibilities
- Ensure the availability, performance, and resilience of production environments by proactively monitoring systems and resolving complex operational issues.
- Administer and optimize CI/CD pipelines, containerized platforms, IBM Cloud Pak Stacks, Confluent, Elasticsearch, and supporting infrastructure to improve operational efficiency.
- Lead incident response, perform root cause analysis, and implement preventive actions to reduce recurring issues and improve service reliability.
- Develop and enhance monitoring, logging, alerting, automation, and operational processes to improve system observability and reduce manual effort.
- Provide technical leadership and mentorship to junior engineers while supporting deployments, release management, and production readiness activities.
- Maintain system security, backup, disaster recovery, and compliance standards while ensuring accurate operational documentation and runbooks.
- Collaborate with Development, Platform Engineering, QUALITY, Network, and Security teams to resolve complex technical challenges and continuously improve production operations.
Requirements
- Bachelor's degree or Diploma in Computer Science, Engineering, or a related field.
- 4–6 years of experience in Technical Operations, Production Support, Site Reliability Engineering (SRE), DevOps, or System Administration.
- Strong experience with Linux/Windows administration, Docker, Kubernetes, OpenShift, CI/CD pipelines, automation tools, scripting (Bash, PowerShell, Python), and configuration management.
- Hands-on experience with IBM Cloud Pak Stacks (CP4BA, CP4I, CP4D), Confluent, Elasticsearch, Databases (SQL, DB2, MongoDB, etc.), JVM troubleshooting, monitoring, and observability platforms.
- Strong understanding of networking, security best practices, IAM, production support, backup, disaster recovery, scalability, performance tuning, and service reliability principles.
- Proven experience in incident management, troubleshooting, root cause analysis, and implementing operational improvements within enterprise production environments.
- Excellent technical leadership, analytical, communication, mentoring, and collaboration skills with the ability to manage multiple priorities effectively.