Über diese L2 DevOps Engineer Stelle bei Showpad
Summary: We are seeking an experienced L2 DevOps Engineer with over 5+ years of experience to design, implement, and manage infrastructure and deployment pipelines across hybrid environments, including on-premises data centers and cloud platforms. This role requires expertise in AWS, Azure, Kubernetes, IaC, CI/CD, and monitoring tools, along with strong troubleshooting and security practices. You will work closely with development, QA, and operations teams to ensure reliable, scalable, and secure delivery of enterprise-grade applications, while also mentoring junior team members and driving automation initiatives.
Responsibilities:
- Manage and support hybrid infrastructure, including physical data centers and AWS/Azure services.
- Design and implement Infrastructure-as-Code (IaC) using tools such as Terraform, Pulumi, CloudFormation, and Helm; manage GitOps deployments with ArgoCD/Flux.
- Configure and maintain Kubernetes clusters (EKS), ensuring scalability and high availability.
- Build, maintain, and optimize CI/CD pipelines with Jenkins, GitLab CI, Bitbucket, and GitHub.
- Automate deployments, scaling, and monitoring processes.
- Set up and maintain observability using Grafana, Prometheus, and other monitoring tools.
- Perform root cause analysis of performance issues and participate in on-call support.
- Manage SSL/TLS and code signing certificates and implement best practices for infrastructure security.
- Work with cross-functional teams to improve infrastructure reliability and delivery processes.
- Maintain clear documentation and contribute to knowledge sharing and incident post-mortems.
Required Skills & Experience:
- 5+ years of experience in DevOps, Site Reliability Engineering, or related roles.
- Strong hands-on expertise with AWS and Azure cloud services.
- Solid experience with Infrastructure-as-Code tools and GitOps.
- Proven experience with Kubernetes/EKS setup and operations.
- Proficiency in CI/CD pipelines and GitOps practices.
- Strong knowledge of monitoring and observability tools and incident/on-call management.
- Scripting and automation skills in Python, TypeScript, and Bash/shell.
- Good understanding of networking, Linux/Windows Server administration, and troubleshooting.
- Experience with data streaming and CDC pipelines and administering relational and in-memory data stores.
- Configuration management and server automation experience with SaltStack and/or Ansible.
- Experience managing SSL/code signing certificates and understanding of security best practices and compliance frameworks.
Soft Skills:
- Strong problem-solving and root cause analysis capabilities.
- Excellent communication and documentation skills.
- Ability to collaborate in weekly/bi-weekly team meetings and contribute to on-call rotations.
- Mentoring and knowledge-sharing mindset.