About this Senior Computer Vision Engineer role at Swordhealth
Since 2020, Sword has expanded across Musculoskeletal, Women’s Health, Cardiometabolic, and Mental Health, and is now moving beyond the session to a fully AI-native, 24/7 care program that brings physical activity, therapeutic exercise, psychotherapy, nutrition, and behavior change into one connected experience. More than 1 million members across three continents have completed over 15 million AI sessions, helping 2,000+ enterprise clients avoid more than $1 billion in unnecessary healthcare costs. Backed by 59 clinical studies, 43 patents, and more than $500 million raised from leading investors including Khosla Ventures, General Catalyst, and Founders Fund, Sword is defining a new standard for healthcare.
Thrive is Sword's program for chronic joint and back pain: personalized, clinician-designed physical therapy delivered at home, with an AI care specialist guiding members between visits. What makes it work is movement understanding. During each session, Thrive reads how a member moves from the camera and gives real-time feedback on their exercises, the way a physical therapist in the room would.
The Computer Vision team, part of the Algorithms org, builds the models behind that. We turn a camera feed into an accurate read of human movement, 2D and 3D pose and body dynamics, in real time, on-device or in-the-cloud, and turn that movement into clinical signals. This role owns that computer vision end to end: the models, the data lifecycle that feeds them, the evaluation, and the systems that ship them to members at scale.
AI fluency is a core expectation at Sword Health. Every candidate is assessed against our three-level framework — be ready to share real examples of how AI is already part of how you work.
-
Explorer (Level 1) — Uses AI daily to boost personal productivity
-
Builder (Level 2) — Creates workflows and tools that elevate the whole team
-
Integrator (Level 3) — Embeds AI into products and processes at scale
Every hire must demonstrate at least Level 1. The expected level will vary depending on the seniority of the role.
What you'll be doing
Own core computer vision models, from 3D human pose to statistical body modeling, taking them from prototype to production;
Ship those models to run real-time in the cloud and on-device on tablets, owning the conversion and optimization in between;
Own the data and code lifecycle behind them: training frameworks, annotation workflows, pipelines, auto-labeling, test sets and taxonomy;
Extend movement understanding into multimodal territory, combining it with language and reasoning; build novel approaches in the movement-intelligence domain;
Unify and mature how we train, track, version and deploy models, so every result is reproducible and testable;
Design systems that run without you, automating the loops so the team's output scales past manual effort;
Help grow the Computer Vision team by defining and promoting best practices, establishing principles that scale your impact.
What you need to have
5+ years solving complex problems with Computer Vision, with models shipped to production;
Strong software engineering foundation across architecture, pipelines, MLOps and the full model lifecycle;
Deep, hands-on command of modern Computer Vision (transformers and convolutional models), with real depth in at least one of detection, segmentation, tracking, pose, or 3D;
A data-centric instinct: you cook your own data, build data flywheels and active-learning loops, and treat the dataset as source code;
Solid grounding in multimodal AI, with the ability to build systems that pair vision with language and reasoning when the problem calls for it;
Strong written, asynchronous communication. You think in specs and documents and leave a clear trail others can build on;
Self-direction and full ownership: you scope ambiguity into a plan and ship without being told;
Fluency with PyTorch or JAX, the Python data stack, and comfort picking up new languages and codebases;
AI-native ways of working: you build AI into the workflow itself, from agents to LLM-in-the-loop tooling.
What we would love to see
Experience with pose estimation, body modeling, movement intelligence, or human-motion work;
Experience working with massive data, building full-lifecycle ML systems at a startup or scaleup, wearing different hats;
Experience automating experimentation or model-improvement loops;
Product mindset, users empathy and desire to ship impact to millions of members.
These compensation bands are just the starting point. Once someone joins and proves they’re outlier talent, we adjust quickly to ensure their compensation aligns with their impact.
Our job titles may span more than one career level. Actual pay is determined by skills, qualifications, experience, location, market demand, and other factors. Compensation details listed in this posting reflect the base salary and any potential variable, bonus or sales incentives, and the Company’s estimation of the value of private company stock options, if applicable. The pay range is subject to change, future value of company stock options is not guaranteed, and compensation may be modified in the future. In addition to our total compensation, Sword offers a number of benefits as listed below.