About this Site Reliability Engineer role at Technology & Product
Are you looking
- To make meaningful changes at the fastest-growing tech company in Japan?
- For a fair, highly flexible, inclusive working environment where anyone, regardless of their background, can make an impact?
- To build the next bold idea in customer acquisition and revenue growth through conversational commerce technology
- For a place where things are not just fast - but they happen now?
We Are Zeals - “Designing conversations. Driving conversions.”
Zeals is not another chatbot. Our automated chat solution and conversational technology change how consumers purchase or act online. Zeals converts customers who appear "unengaged" to become "valuable" through designed chat on Messaging Apps. Using AI-powered, human-designed conversation flows across chat platforms, Zeals enables brands to build personalized, automated conversations with their consumers and develop a deeper understanding of their preferences while creating a delightful, seamless shopping experience. Our last 3 years have been phenomenal. We have captured the heart of many businesses in Japan, with over 480 enterprises relying on our chat commerce offering to boost their mobile business.
The goal is simple: to “revolutionize hospitality on the Internet” by unlocking digital customer service experiences that were limited to physical stores previously.
Just Getting Started…
With more than 480 enterprise companies worldwide already utilizing our product, we are perfectly poised to seize market share and fuel exponential growth in the US market, with an eye on global expansion. Investors have taken note of our potential, leading us to recently secure nearly $40 million in funding, backed by Salesforce Ventures.
About the role
At ZEALS, we're building the infrastructure behind the next generation of conversational AI products.
Our conversational commerce platform processes millions of customer conversations every day on a high-volume, event-driven system built from dozens of microservices. In parallel, we're rapidly expanding Omakase.ai, our new AI agent platform, and developing deeper integrations between our products to unlock new capabilities and use cases.
Joining the SRE team means more than maintaining production systems—you'll help shape the infrastructure that powers multiple AI products. From improving reliability and performance to enabling new feature development and platform integrations, you'll have the opportunity to work on challenging, large-scale engineering problems while influencing the future architecture of our services.
Who you are
- You can discuss systems architecture as a whole and how its components interact before diving into detail, and you can confidently walk stakeholders and colleagues through how you reached a conclusion.
- You love to solve problems. You find the root cause of issues and stay on an error until you understand it, not just until it goes away or resort to guesswork.
- Calm under pressure in an incident, you follow a problem across the whole stack,\, rather than stopping at the first suspicious layer.
- Proactive in adding value where you can identify it. You spot toil and weak spots, validate the real need, and bring a proposed solution before executing. Being pragmatic is critical: your time is valuable, and we want what we release to have a tangible impact on our system and our developers.
- A team player: affable, empathetic, and understanding with our developers. Eager to learn and get involved, as well as holding others accountable.
What you'll do
- Keep critical customer journeys reliable and available, and respond first when they aren't.
- Join the on-call rotation for business critical services.
- Own alerting and observability: better coverage, fewer false positives, and catching issues before customers do.
- Build and operate in-house platform tooling: Kubernetes operators, custom Kubernetes APIs, GitOps automation, and deployment tooling that helps teams ship faster and safer.
- Eliminate toil by turning recurring developer requests into self-service, bringing an automation mindset to reduce complexity.
- Drive down platform cost without trading away reliability.
- Design infrastructure that scales with bursty, high-volume traffic, managed as code across many environments.
- Measure the platform against Well-Architected principles and act on what's weak, whether that's a reliability gap, an oversized resource, or a manual process worth automating.
What you'll need
- 3+ years in software engineering, SRE, infrastructure, or platform engineering on GCP or AWS.
- Hands-on Kubernetes experience: you can debug a workload, not just deploy one.
- Monitoring and observability experience (we use Grafana, Loki, Prometheus, and Thanos).
- Comfortable troubleshooting databases (we run MongoDB, Elasticsearch, and Postgres).
- CI/CD and GitOps experience (we use GitHub Actions, Helm, and Argo CD).
- Able to write and read code for tooling and automation (we use Go, Bash, and Python).
- Eligible to work in Japan, and business-level English.
What makes you stand out
- Terraform and infrastructure-as-code at scale.
- Go proficiency and experience building Kubernetes operators.
- Operating event-driven systems and high-throughput messaging at scale.
Conversational Japanese.