Jobs Companies Moonshot Head of AI Safety

Über diese Head of AI Safety Stelle bei Moonshot

Moonshot · Remote · Toronto, Ontario, Canada

Moonshot believes that marginalized people in society, including people of colour, Indigenous people, people from diverse socioeconomic backgrounds, women, Disabled people, and LGBTQIA+ people, must be centred in the work we do. We strongly encourage applications from people with these identities, or from other communities currently underrepresented in our workforce.

All qualified applicants will be afforded equal employment opportunities without discrimination based on race, creed, colour, national origin, sex, age, disability, or marital status. We know a diverse workforce will enable us to understand the drivers behind violent extremism and online harms in greater depth, and to do better work to counter them.

About the role

Moonshot is recruiting a Head of AI Safety to lead the delivery, development, and growth of our AI Safety portfolio. The role combines Moonshot's expertise in violence prevention, behavioural risk, and online harms with the emerging practice of evaluating and improving the safety of AI systems. The portfolio addresses harm categories including pathways to violence, extremism, child sexual exploitation, abuse and grooming (CSEA), mental health and crisis, and risks affecting children and teenagers.

The Head of AI Safety will serve as Moonshot's primary applied AI safety counterpart for frontier AI companies, governments, and regulators. The role will work closely with model, policy, trust and safety, product, research, and engineering teams. This is not an engineering or data-science role, but it is a hands-on position requiring the successful candidate to lead and participate directly in red teaming and adversarial evaluation, working in detail with evaluation methodologies, test scenarios, model responses, safety policies, and intervention frameworks. The role holds responsibility for client and partner relationships, project and staff management, methodological quality, and business development. The Head of AI Safety will build and maintain relationships across the wider AI safety ecosystem, including with governments, foundations, regulators, academics, researchers, and civil society organizations.

This is a remote position. Candidates must be based in Ontario, Canada, due to employment and regulatory requirements.

Your responsibilities will include:

Applied AI Safety, Evaluation, and Advisory

  • Lead and quality-assure Moonshot's applied AI safety work across harm categories including pathways to violence, extremism, CSEA, abuse and grooming, mental health and crisis, and risks affecting children and teens, using methods such as red teaming and adversarial evaluation of AI systems.
  • Advise frontier AI companies on how to improve the safety of their models, products, policies, and intervention systems.
  • Translate insights from psychologists, child-safety specialists, violence-prevention practitioners, safeguarding experts, and other subject-matter experts into clear, actionable guidance for model safety, policy, product, research, and engineering teams.
  • Set the methodological approach for the portfolio, translating violence-prevention, safeguarding, and behavioural-risk expertise into structured and testable evaluation frameworks.
  • Lead and participate directly in red teaming and adversarial evaluation, working in detail with test scenarios, model responses, scoring criteria, safety policies, and evaluation results.
  • Identify patterns, edge cases, and potential safety failures, and develop practical recommendations for improving model behaviour and user protections.
  • Maintain rigour and clear documentation across the team's technical deliverables, suitable for technical, government, and foundation audiences.
  • Ensure work is delivered within a clear ethical framework and in compliance with contractual, legal, data protection, and ethics obligations.
  • Identify, manage, and escalate operational, reputational, delivery, and partnership risks.

Client & Partner Management

  • Serve as Moonshot's primary applied AI safety counterpart for frontier AI company partners, governments, regulators, and the wider ecosystem invested in AI safety.
  • Build trusted relationships with model, policy, trust and safety, product, research, and engineering teams.
  • Build and sustain relationships across the wider AI safety ecosystem, including governments, foundations, regulators, academics, researchers, civil society organizations, and specialist practitioners.
  • Represent Moonshot externally in meetings, briefings, workshops, and sector engagement, including with regulators and policymaker audiences.

Team Leadership & Management

  • Provide direct leadership, coaching, and management to Moonshot's AI safety team.
  • Foster a collaborative, accountable, and mission-driven team culture, with particular attention to wellbeing given the sensitive nature of the work.
  • Support workforce planning, performance management, and professional development across the team.
  • Ensure effective coordination with internal teams supporting the portfolio, including operations, finance, research, and technical teams.

Portfolio Development & Growth

  • Develop Moonshot's AI safety portfolio, identifying strategic opportunities, partnerships, and funding.
  • Lead proposal development, scoping, and renewals with technical credibility, using precise, defensible language suited to technical and government audiences.
  • Develop repeatable methodologies, service offerings, and partnerships that allow the portfolio to grow while maintaining methodological rigour and delivery quality.
  • Support external communications, publications, briefings, and thought leadership that establish Moonshot as a credible voice in applied AI safety.
  • Oversee project planning, staffing, budgeting, forecasting, and delivery timelines across the portfolio.

Requirements

Essential:

  • Experience in trust & safety, online harms, or a closely related field such as violence prevention, safeguarding, or public health, and the ability to adapt that knowledge to AI systems.
  • Curiosity about AI and the ability to build technical fluency quickly, enough to engage credibly with technical counterparts at AI companies. Much of this work is new, so comfort learning as you go matters more than existing AI safety expertise.
  • Experience designing research, evaluation frameworks, or interventions for harm categories such as violent extremism, CSEA, self-harm and crisis, or targeted violence.
  • Demonstrated experience managing projects, teams, budgets, partners, and clients, with strong people management skills.
  • Excellent written communication, with experience producing credible (not promotional) material for government, foundation, or enterprise audiences.
  • Comfort and demonstrated resilience working with highly sensitive or graphic content (CSEA, extremist material, crisis content), with awareness of wellbeing practices for this kind of work.
  • Strong judgment and the ability to navigate ambiguity, competing priorities, and sensitive stakeholder environments, including representing organizations externally.
  • Willingness to travel and work outside regular hours where needed to accommodate clients or respond to incidents.
  • Highly trustworthy, with discretion and diplomacy, and willing to undertake relevant security clearance procedures.
  • Experience supporting business development, grant funding, or procurement.
  • Commitment to Moonshot's mission.
  • Candidates must be eligible to work in Canada, and will be required to undertake and pass a standard background check, along with any relevant security clearance procedures per client needs.

Desirable:

  • Direct experience in model safety, red teaming, or adversarial evaluation of LLMs or other AI systems.
  • Understanding of LLM architecture, safety tooling, or trust & safety policy.
  • Prior experience in child safety evaluation, teen-safety product work, or grooming and CSEA detection.
  • Familiarity with government or regulatory engagement, such as briefing officials or supporting policy submissions.
  • Experience with intervention or diversion programme design that can transfer to AI-mediated interventions.
  • Academic or applied background in radicalization studies, forensic psychology, or violence risk assessment.
  • Familiarity with taxonomy or classifier development, including how testing data feeds a classifier.

Benefits

  • 25 days paid vacation leave, plus Statutory Holiday
  • Flexible public holiday policy with the option to work statutory holidays in exchange for a day off at another time.
  • Group healthcare package, including coverage for partners and children (80% Co-Insurance).
  • HSA is restricted to mental health practitioners only
  • Dental & Vision Insurance (80% Co-Insurance).
  • Life & LTD Disability Insurance.
  • 24/7 access to counselling via our Employee Assistance Program.
  • Generous maternity and paternity leave: 26 weeks paid maternity leave, 8 weeks paid paternity leave.
  • All permanent employees are granted share options upon employment.

Salary: $115,000 - $140,000 CAD (depending on skills and experience).

Bereit, sich bei Moonshot zu bewerben?
Bei Moonshot bewerben

Über Moonshot

About Moonshot:

Moonshot is a social impact business dedicated to disrupting and reducing online harms across the globe. We currently operate in more than 28 countries across different forms of targeted violence and other public safety issues. We use data-proven techniques to ensure our clients respond effectively, and our work ranges from targeted intervention programs, threat monitoring, software development, and digital capacity building.

We do this through:

  • Finding new ways to reach individuals at risk of involvement in the milieu that promotes targeted violence or might be at risk of mobilizing to violence themselves.
  • Collaborating with partners and working for clients including governments, NGOs, and private sector organizations from across the globe.
  • Building a multifaceted team with a diversity of backgrounds, both professional and academic, including international development, law enforcement, communications, psychology, data science and software engineering.
  • Investing in the research and development of new technologies and methodologies to counter violence and other public safety issues.

Working at Moonshot:

We’re growing quickly, have big ambitions, and high expectations of our staff. Our dedication to finding effective responses and leading innovation means that our work environment is fast-paced, dynamic and creative. We match this by offering our staff access to a range of learning and development options, scope to advance personal subject-matter expertise, and opportunities for career progression.

Our staff say they value:

  • Our shared sense of purpose: working as a team to find new solutions to global challenges.
  • Personal development opportunities: a chance to learn new things and get even better at what you already do.
  • Our ideas-driven culture: opportunities to work with creativity and autonomy whatever your position in our organisation.
  • The diversity of thought: working with staff from a wide range of personal and professional backgrounds.
  • Open and collaborative working: being part of a team who support each other to achieve great results.

Alle Jobs bei Moonshot ansehen →

Ähnliche Jobs

Moonshot
OSINT Analyst
Moonshot
⚡ Früh bewerben London, England, United Kingdo... Vor Ort £30,784–£40,000
● Neu 👁 Gesehen ✓ Beworben vor 6 Tg.
Moonshot
Manager (Insights)
Moonshot
⚡ Früh bewerben London, England, United Kingdo... Vor Ort £46,000–£55,000
● Neu 👁 Gesehen ✓ Beworben vor 6 Tg.
Moonshot
Head of AI Safety
Moonshot
⚡ Früh bewerben London, England, United Kingdo... Vor Ort £67,000–£80,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Moonshot
Head of AI Safety
Moonshot
⚡ Früh bewerben Washington, District of Columb... · standortgebunden $110,000–$145,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Moonshot
Intelligence Analyst
Moonshot
⚡ Früh bewerben Toronto, Ontario, Canada · standortgebunden CA$55,000–CA$64,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Moonshot
Intelligence Analyst
Moonshot
⚡ Früh bewerben Washington, District of Columb... · standortgebunden $50,000–$57,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Moonshot
Threat Intelligence Analyst
Moonshot
⚡ Früh bewerben Washington, District of Columb... · standortgebunden $60,000–$75,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Wo.
Moonshot
Linguist Talent Pool (Albanian, Hindi, Kurdish (Sorani), Punjabi and Urdu)
Moonshot
⚡ Früh bewerben London, England, United Kingdo... Vor Ort ⚠ 6 Mon.+ alt
● Neu 👁 Gesehen ✓ Beworben vor 6 Mon.
Velan Studios
Gameplay Animator (Senior+)
Velan Studios
⚡ Früh bewerben Toronto, Ontario, Canada Vor Ort $85,000–$85,000
● Neu 👁 Gesehen ✓ Beworben vor 3 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Moonshot

Alle Jobs bei Moonshot ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos