Jobs Companies Bosonai Site Reliability Engineer

À propos de ce poste Site Reliability Engineer chez Bosonai

Bosonai · Télétravail · Toronto
About The Role
 
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.
 
Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from “it works” to dependable, observable, and scalable.
 
You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area—networking, cluster scheduling, storage, GPU systems, or AI infrastructure— and the curiosity and judgment to collaborate across the rest.
 

Responsibilities

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads
  • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms
  • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads
  • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements
  • Improve provisioning, configuration management, testing, and deployment automation
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
  • Minimum Qualifications

  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role
  • Strong hands-on expertise in at least one of the following:
  • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand
  • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms
  • Distributed storage, particularly Ceph
  • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting
  • AI training or model-serving infrastructure
  • Experience operating production systems with a focus on availability, performance, security, and automation
  • Strong Linux administration and scripting skills
  • A systematic approach to troubleshooting across multiple layers of a complex system
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team
  • Preferred Qualifications

  • Experience supporting GPU-intensive AI or HPC environments
  • Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet
  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling
  • Experience operating or tuning Ceph clusters
  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
  • Experience with hardware provisioning, firmware management, and bare-metal automation
  • Experience running large-scale distributed training or high-throughput inference workloads
  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure

  • Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we’d love to hear from you.
    Prêt à postuler chez Bosonai ?
    Postuler chez Bosonai

    Emplois similaires

    MongoDB
    Staff Site Reliability Engineer, Fabric
    MongoDB
    ⚡ Postuler tôt Toronto; Vancouver Sur site CA$144,000–CA$200,000
    ● Nouveau 👁 Vu ✓ Postulé il y a 5 h
    MongoDB
    Site Reliability Engineer (Senior or Staff), Deployments
    MongoDB
    ⚡ Postuler tôt Boston; Miami; New Jersey; New... Sur site $127,000–$249,000
    ● Nouveau 👁 Vu ✓ Postulé il y a 5 h
    MongoDB
    Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
    MongoDB
    ⚡ Postuler tôt Montreal; Toronto Sur site CA$144,000–CA$200,000
    ● Nouveau 👁 Vu ✓ Postulé il y a 5 h
    Baseten
    Site Reliability Engineer
    Baseten
    ⚡ Postuler tôt San Francisco Hybride $165,000–$330,000
    ● Nouveau 👁 Vu ✓ Postulé il y a 5 j
    Movable Ink
    Senior Site Reliability Engineer
    Movable Ink
    ⚡ Postuler tôt Movable Ink - Toronto (Remote) · lieu restreint $140,000–$182,000
    ● Nouveau 👁 Vu ✓ Postulé il y a 6 j
    Movable Ink
    Lead Site Reliability Engineer
    Movable Ink
    ⚡ Postuler tôt New York, NY, United States Sur site $184,200–$240,000
    ● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
    mthree Recruiting Portal
    Junior Production Support/SRE Analyst
    mthree Recruiting Portal
    ⚡ Postuler tôt Toronto, Ontario, Canada Sur site
    ● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
    mthree Recruiting Portal
    Junior Production Support/SRE Analyst
    mthree Recruiting Portal
    ⚡ Postuler tôt Toronto, Ontario, Canada Sur site
    ● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
    Smiledigitalhealth
    Cloud Performance Engineering - Site Reliability Engineer
    Smiledigitalhealth
    ⚡ Postuler tôt Toronto, Ontario Télétravail
    ● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.

    Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

    Plus d’emplois chez Bosonai

    Voir tous les emplois chez Bosonai →

    Postuler maintenant
    🤖

    Doucement — un instant

    JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

    Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

    Catch your next role the second it’s posted.

    Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

    Create free account

    Free forever · takes 30 seconds · already have one?

    Prenez une longueur d’avance dans votre recherche d’emploi.

    Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

    Rejoindre le canal — c’est gratuit