Sobre este puesto de Software Development Team Lead, Infrastructure Platform Team en Spare
About Us
Spare is a fast-growing, successful startup. We thrive on innovation, rapid execution, and delivering top-tier products that make a real impact in the transit industry. We are now seeking a Software Development Team Lead to join our Infrastructure Platform Team and help build and operate the cloud platform that every Spare product runs on.
About This Role
As leader of the team responsible for Spare's core infrastructure platform, you will play a critical role in designing and building the reliable, secure, and cost-efficient foundation that powers our customers' on-demand transit systems. Working closely with others on the team, you will own the correct and efficient operation of the platform across a variety of real-world production scenarios.
In this role, you'll split your time 50/50 between hands-on technical contribution and people leadership, working in an autonomous environment where you'll solve interesting technical challenges while growing and mentoring a high-performing team. Currently there are four software developers reporting to this role.
Given the nature of our business, this role requires someone who can balance technical expertise with strong product sensibilities, creating solutions that are both technically sound and accessible to end users. The role includes some travel as part of the job responsibilities – specifically, up to four customer site visits per year to gain firsthand insights, plus participation in our biannual software development hackathons in Vancouver.
Key Responsibilities
Own the design and development of core infrastructure platform capabilities from inception to launch
Build and evolve tooling, automation, and platform services that make every engineering team at Spare faster and safer
Architect and implement high-performance, scalable distributed systems on GCP and Kubernetes
Drive improvements in cluster reliability, application resilience, and internal access security
Operate and maintain Spare's Redis and PostgreSQL databases — availability, performance tuning, scaling, upgrades, backups, and disaster recovery
Drive Spare toward AI SRE practices — embed AI agents into incident detection, alert triage, and operational workflows to make reliability work faster and more proactive
Manage and continuously improve the SRE on-call rotation — healthy schedules, clear escalation paths, and blameless post-mortems that turn incidents into systemic improvements
Drive FinOps practices across the organization — own cloud spend visibility, right-sizing, and cost optimization initiatives that deliver measurable savings
Use AI agentic tooling daily to accelerate your own and your team's engineering output, and coach the team to do the same
Actively mentor software developers of all levels and uplift team capacity
Collaborate cross-functionally with product managers, designers, and other software developers
Ensure 99.99% uptime and maintain exceptional system performance
Participate in team agile rituals and help improve software development processes
Who You Are
A highly productive software developer with a proven track record of delivering high-quality code in complex environments
Highly proficient with modern AI development tools and agentic workflows — you use AI to move faster and think better, not as a crutch
A passionate mentor and technical leader who enjoys helping others grow
Passionate about distributed systems, cloud infrastructure, and platform engineering
Cost-conscious by default — you treat cloud spend as an engineering problem and know how to deliver savings without sacrificing reliability
Adept at balancing reliability and security rigor with practical developer-experience constraints
Requirements
7+ years of software development experience, with at least 2+ years in a people leadership role
Expert in backend technologies with strong distributed systems experience
Demonstrated proficiency with AI-assisted and agentic development workflows (AI coding agents, automation of engineering and operational tasks)
Experience operating systems at scale with a strong reliability and uptime mindset
Experience running or managing an SRE on-call rotation, including incident response and post-mortem culture
Deep experience with cloud infrastructure (GCP) and container orchestration (Kubernetes)
Experience driving cloud cost optimization and FinOps initiatives — right-sizing, spend visibility, and measurable cost savings
Experience with infrastructure-as-code and configuration management tooling (Terraform)
Understanding of security best practices, especially around internal access control and container security
Demonstrated success in managing a team of software developers, with a focus on team and individual performance
Demonstrated ability to mentor other developers and provide technical leadership
Strong problem-solving, debugging, and system design skills
Excellent communication and collaboration skills
It will be considered a plus (nice-to-have):
Experience in the transit industry or another real-time, safety-critical domain
Experience building internal developer platforms and golden-path tooling
Experience applying AI/LLM tooling to SRE or infrastructure operations (AIOps, automated incident triage)
Experience with CI/CD systems at scale, including test sharding and build performance
Experience operating and maintaining PostgreSQL and Redis in production — performance tuning, indexing, replication, backup and restore
Don't meet every single requirement?
Studies have shown that women and people of colour are less likely to apply to jobs unless they meet every single qualification in the job posting.
At Spare, we are committed to creating a diverse and inclusive environment so we strongly encourage you to apply even if you don't believe you meet every single qualification outlined. We also do our best to respond to all applications we receive.
About the Infrastructure Platform Team
The Infrastructure Platform Team's job is to build and run the foundation that Spare's on-demand transportation platform depends on — where circumstances change in real time and downtime is not an option. Primarily this involves keeping our Kubernetes clusters reliable and secure, hardening our application resilience, driving cloud cost efficiency, and making sure our distributed systems can handle real-time updates efficiently.
Why Join Us?
Work on challenging technical problems with real-world impact in the transportation industry
A fast-paced, high-impact role in a rapidly growing startup
The opportunity to take ownership of core systems and drive innovation
A dynamic, collaborative, and supportive team culture
Competitive salary and equity options
If you are ready to tackle complex routing and optimization challenges and thrive in a high-performance environment, we'd love to hear from you!