Über diese Incident Manager Stelle bei Rakuten
Job Description:
Why should you choose us?
Rakuten Symphony is reimagining telecom, changing supply chain norms and disrupting outmoded thinking that threatens the industry’s pursuit of rapid innovation and growth. Based on proven modern infrastructure practices, its open interface platforms make it possible to launch and operate advanced mobile services in a fraction of the time and cost of conventional approaches, with no compromise to network quality or security.
Rakuten Symphony has operations in Japan, the United States, Singapore, India, South Korea, Europe, and the Middle East Africa region. For more information, visit: https://symphony.rakuten.com.
Building on the technology Rakuten used to launch Japan’s newest mobile network, we are taking our mobile offering global.
To support our ambitions to provide an innovative cloud-native telco platform for our customers, Rakuten Symphony is looking to recruit and develop top talent from around the globe. We are looking for individuals to join our team across all functional areas of our business – from sales to engineering, support functions to product development.
Let’s build the future of mobile telecommunications together!
About Rakuten Group, Inc. (TSE: 4755) is a global leader in internet services that empower individuals, communities, businesses and society. Founded in Tokyo in 1997 as an online marketplace, Rakuten has expanded to offer services in e-commerce, fintech, digital content and communications to 2 billion members around the world. The Rakuten Group has over 30,000 employees, and operations in 30 countries and regions. For more information visit https://global.rakuten.com/corp/.
About Department:
The Network Operations (SXC) Division at Rakuten Mobile is pioneering the future of telecommunications through a fully virtualized, cloud-native architecture. We are transitioning from traditional manual operations to an Autonomous Network Operations model. We are looking for an AI-Ops Manager to lead this transformation, running our incident management function through the deployment, supervision, and continuous refinement of Agentic AI models.
Role Purpose:
You will not just manage incidents; you will manage the AI agents that resolve them. Your mission is to minimize MTTR (Mean Time To Recovery) by overseeing an AI-first operational environment, ensuring our cloud-native network (Open RAN/Core) remains resilient through autonomous and semi-autonomous AI intervention.
Job Duties:
- Agentic AI Supervision & Incident Command: Serve as the lead incident commander, monitoring and steering Agentic AI models as they perform real-time diagnostics and remediation. You are responsible for the "Human-in-the-loop" (HITL) oversight of these agents, ensuring they operate within safety guardrails.
- AI-Orchestrated Operations: Manage the end-to-end incident lifecycle by orchestrating AI-driven workflows across CLOUD/CORE/IPTX/RAN/OSS/BSS. You will validate AI-generated root cause analysis (RCA) and approve autonomous remediation actions.
- AI Knowledge Hub Management: Own the evolution of our AI-based Knowledge Hub. You will ensure that AI models continuously ingest network telemetry and incident data to generate, update, and refine autonomous resolution articles, turning past incidents into future AI-driven preventative logic.
- Predictive Network Resilience: Leverage AI capabilities to shift from reactive to predictive operations. Use AI-driven observability to detect anomalies before they impact the end customer, proactively triggering agentic workflows to mitigate service degradation.
- Model Performance Tuning: Act as the primary owner of the AIOps roadmap. Regularly audit the decision-making logic of AI agents, refine their training data, and optimize their performance to ensure maximum accuracy in complex, cloud-native network environments.
- Continuous Improvement: Present AI-derived incident insights to senior management, utilizing data-backed trends to drive infrastructure improvements and architectural changes.
Minimum Qualifications:
- Work Experience: Minimum 7 years in Telecom/IT operations, with at least 2 years specifically focused on AIOps, machine learning operations (MLOps), or AI-driven network automation.
Core Technical Skills:
- Agentic AI Mastery: Deep expertise in managing autonomous agents. Ability to interpret agent decision-making paths, troubleshoot automation failures, and perform real-time steering during critical outages.
- Cloud-Native & Telecom Expertise: Expert-level knowledge of Open RAN, Kubernetes, containerized network functions (CNFs), and cloud-native architecture.
- AIOps Tooling: Proficiency in operating enterprise-grade monitoring systems and ticketing systems that integrate with AI-driven observability platforms.
- Knowledge Management: Experience managing AI-based knowledge ecosystems where documentation is dynamically generated and maintained by machine learning models.
- Process Mastery: Strong alignment with ITIL frameworks, adapted for a high-velocity, automated, AI-first environment
- Analytical Mindset: Exceptional ability to synthesize high-dimensional technical data into actionable strategic intelligence
Preferred Qualification:
- Advanced degree or certification in Artificial Intelligence, Data Science, or Machine Learning.
- Experience in Python or other scripting languages for the purpose of creating custom AI-agent integrations or automation hooks.
- Proven track record of leading a successful digital transformation from manual operations to an AI-automated or "Self-Healing" network model.
- Relevant certifications in Cloud (e.g., CKA/CKAD) and ITIL 4 (Managing Professional).
RAKUTEN SHUGI PRINCIPLES:
Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
- Always improve, always advance. Only be satisfied with complete success - Kaizen.
- Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.
- Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.
- Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.
- Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.