Sobre esta vaga de Technical Lead, Cloud QA na Rakuten
Job Description:
Job Summary:
We are looking for a highly skilled Site Reliability Engineer (SRE) / Cloud Operations Engineer responsible for operating, monitoring, troubleshooting, and maintaining cloud-native platforms and Kubernetes-based telecom infrastructure. The role focuses on platform reliability, availability, monitoring, incident management, and operational excellence.
Mandatory Skills:
Linux System Administration
Istio Service Mesh Operations
Prometheus, Grafana, and AlertManager
Kubernetes Networking and RBAC
CI/CD Operations
Ansible and Infrastructure Operations
Storage Administration (PV, PVC, CSI)
Observability and Monitoring
Incident Management and Root Cause Analysis
Roles & Responsibilities:
Perform day-to-day operations and support for cloud-native infrastructure platforms.
Diagnose and resolve issues related to pods, deployments, StatefulSets, networking, and storage.
Manage Istio service mesh configurations, traffic routing, mTLS, and observability.
Create and maintain monitoring dashboards, alerts, and SLO/SLA reporting using Prometheus and Grafana.
Perform incident response, troubleshooting, root cause analysis, and service recovery activities.
Support CI/CD deployment validation and operational readiness checks.
Manage Linux servers including patching, hardening, user management, and system optimization.
Handle backup, restore, upgrade, and lifecycle management activities for Kubernetes environments.
Collaborate with engineering teams to improve reliability, scalability, and operational best practices.
Job Requirement:
4+ years of experience in Site Reliability Engineering, Cloud Operations, or Platform Operations.
Strong hands-on Kubernetes administration and troubleshooting experience.
Solid Linux administration expertise (RHEL/CentOS/Ubuntu).
Experience with Istio Service Mesh and Kubernetes networking.
Hands-on experience with Prometheus, Grafana, and AlertManager.
Knowledge of Ansible and basic Terraform operations.
Experience supporting production-critical cloud-native environments.
Strong analytical and troubleshooting skills.
Excellent verbal and written communication skills.