Jobs Companies CloudFactory Senior Site Reliability Engineer

Sobre esta vaga de Senior Site Reliability Engineer na CloudFactory

CloudFactory · Presencial · Kathmandu, Bagmati Province, Nepal

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale. 

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

Role Summary

As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly. You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security. This is an exciting opportunity to grow professionally while contributing to a mission-driven organization.

Responsibilities:

  • Design, build, and maintain core infrastructure for developer focus, platform resilience, and scalability.
  • Establish & maintain IaC best practices and templates.
  • Troubleshoot, rollback, and restore services, creating necessary dashboards for high availability.
  • Create and enhance run books for on-call issue resolution.
  • Deploy, support, and monitor new and existing services, platforms, and application stacks.
  • Manage environment capacity and performance.
  • Define availability/SLA for Platform products, ensuring necessary processes, tools, and people are in place.
  • Create and maintain production readiness standards (performance, availability, security, compliance) for all services/APIs, ensuring adherence before go-live.
  • Create, maintain, and enhance monitoring, alerting, and debugging tools.
  • Collaborate with engineering on performance improvements identified via tracking metrics (latency, CPU, etc.).

Requirements

Knowledge:

  • Must Have skills (required)
    • Cloud Architecture: Expertise in AWS cloud infrastructure and microservice solutions (serverless/containers).
    • Infrastructure as Code (IaC): Proficient in provisioning and managing infrastructure via code.
    • CI/CD & DevSecOps: Knowledge of CI/CD, web security, and DevSecOps.
    • Operational Excellence: Understanding of monitoring, alerting, incident management, and 24x7 operational support.
  • Nice To Have skills (Preferred)
    • MLOps / ML tool chain operations
    • Broader web security principles beyond standard DevSecOps.

Skills and Experience:

  • Must Have skills (required)
    • AWS Services: Extensive experience with AWS - EC2, CloudFormation, ECS Fargate, Lambda, SQS, SNS, S3, ECR, RDS, and Route 53.
    • IaC Tools: Hands-on experience with Terraform, CloudFormation, and Serverless Framework and scripting language - Bash/Python/Go. 
    • Monitoring & Logging: Experience with monitoring stack - any of Grafana, ELK stack, CloudWatch, and Prometheus.
    • Containerization & Scripting: Proficiency with Docker and Shell scripting.
    • CI/CD Tools: Experience with GitHub Actions.
  • Nice To Have skills (Preferred)
    • Programming: Experience with Go, Node.js, or Python.
    • Troubleshooting: Advanced troubleshooting skills for complex customer issues.

Other General requirements

  • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
  • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Availability: Willingness to support processes for 24x7 operational support.

Benefits

At CloudFactory, we believe that work should be more than just a job—it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

Join us today and be part of our mission to connect people and technology for a better world! Apply now and bring your whole, authentic self to work—we can’t wait to meet you!

Pronto para se candidatar à CloudFactory?
Candidatar-se à CloudFactory

Sobre a CloudFactory

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale.

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do, as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things, together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!


After submitting your application, all of our communication will be via email, so please check your inbox and spam folders regularly. CloudFactory will at no stage of this process ask candidates to make payments or pay fees of any kind.

Ver todas as vagas na CloudFactory →

Vagas semelhantes

Verisign
Site Reliability Engineer - IBM AIX
Verisign
⚡ Candidate-se cedo Reston,Virginia,United States Híbrido $135,800–$183,800
● Nova 👁 Vista ✓ Candidatada há 34min
Verisign
SRE - Linux
Verisign
⚡ Candidate-se cedo Reston,Virginia,United States Híbrido $135,800–$183,800
● Nova 👁 Vista ✓ Candidatada há 34min
Verisign
Site Reliability Engineer
Verisign
⚡ Candidate-se cedo Reston,Virginia,United States Híbrido $135,800–$183,800
● Nova 👁 Vista ✓ Candidatada há 34min
Okta
Staff Site Reliability Engineer (FedRAMP)
Okta
⚡ Candidate-se cedo Bellevue, Washington; Chicago,... Presencial $194,000–$267,000
● Nova 👁 Vista ✓ Candidatada há 54min
Roblox
Senior Machine Learning Engineer, Reliability
Roblox
⚡ Candidate-se cedo San Mateo, CA, United States Presencial $196,750–$243,290
● Nova 👁 Vista ✓ Candidatada há 1h
Roblox
Senior Site Reliability Engineer, Compute
Roblox
⚡ Candidate-se cedo San Mateo, CA, United States Presencial $243,290–$295,250
● Nova 👁 Vista ✓ Candidatada há 1h
Job Board
Lead Cloud Infrastructure Engineer / Site Reliability Engineer (SRE)
Job Board
⚡ Candidate-se cedo North America Presencial $172,000–$219,000
● Nova 👁 Vista ✓ Candidatada há 1h
Zipline
Hardware Reliability Engineer
Zipline
⚡ Candidate-se cedo South San Francisco, Californi... Presencial
● Nova 👁 Vista ✓ Candidatada há 1h
The New York Times
Technical Product Manager II, Site Reliability Engineering
The New York Times
⚡ Candidate-se cedo New York, NY Híbrido $120,000–$142,000
● Nova 👁 Vista ✓ Candidatada há 2h

Cadastre-se para receber sugestões sob medida com base nas vagas que você abre e nas buscas que você salva.

Mais vagas na CloudFactory

Ver todas as vagas na CloudFactory →

Candidatar-se agora
🤖

Opa — calma aí

A JobsRadar foi feita para pessoas de verdade passando por um momento difícil na busca por emprego — não para requisições automatizadas. Você está clicando rápido demais e agora está temporariamente bloqueado.

Volte mais tarde. Se você está mesmo procurando emprego, estamos com você — apenas aja como um ser humano.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Ganhe vantagem na sua busca por emprego.

Entre no nosso canal do Telegram para o que ajuda você a conseguir a vaga — referências salariais, o pulso semanal do mercado e avisos de novos recursos. Sem spam, só sinal.

Entre no canal — é grátis