About this TechOps & Support Engineer role at Sumerge
Our TechOps & Support Engineer Ensures the stability, availability, and performance of production environments by monitoring systems, automating operational processes, supporting deployments, resolving incidents, and maintaining secure, reliable platform operations.
Responsibilities
- Monitor, maintain, and optimize production environments to ensure high availability, performance, and compliance with established service level objectives (SLAs/SLOs).
- Manage deployments, CI/CD pipelines, and release activities while supporting configuration management and environment consistency across production and non-production environments.
- Administer and troubleshoot containerized platforms, IBM Cloud Pak Stacks, Confluent, Elasticsearch, and supporting infrastructure to ensure reliable system operations.
- Investigate incidents, perform root cause analysis, implement corrective actions, and contribute to post-incident reviews to improve service reliability.
- Configure and maintain monitoring, logging, and alerting solutions to proactively identify performance issues and minimize service disruptions.
- Support backup, disaster recovery, security, patching, and operational compliance activities while maintaining accurate technical documentation and runbooks.
- Collaborate with Development, QUALITY, Platform Engineering, Network, and Support teams to resolve complex technical issues and continuously improve operational processes.
Requirements
- Bachelor's degree or Diploma in Computer Science, Engineering, or a related field.
- Around 2+ years of experience in Technical Operations, Production Support, Site Reliability Engineering (SRE), DevOps, or System Administration.
- Good experience with Linux/Windows administration, Docker, Kubernetes, OpenShift, CI/CD pipelines, automation tools (Jenkins, GitLab CI, GitHub Actions), and scripting (Bash, PowerShell, Python).
- Working knowledge of IBM Cloud Pak solutions (CP4BA, CP4I, CP4D), Confluent, Elasticsearch, Databases (SQL, DB2, MongoDB, etc.), and observability platforms such as Grafana and Prometheus.
- Good understanding of networking fundamentals, security best practices, IAM, backup and disaster recovery, configuration management, and production environment support.
- Strong troubleshooting, incident management, root cause analysis, communication, and collaboration skills with the ability to work effectively under pressure.
- Ability to participate in on-call support, prioritize operational issues, and deliver reliable production support while driving continuous service improvement