À propos de ce poste AI Product Lead, AI Experience & Quality chez Flodesk
Flodesk is recognized in the Inc 5000 as one of the world's fastest-growing email marketing companies, built to help entrepreneurs sell online and design emails that people love to get. We're committed to giving small businesses simple and intuitive tools that help them grow, nurture, and monetize their email list.
We’re a remote-first company headquartered in San Francisco with a globally distributed team, including in-person hubs in Da Nang (Vietnam), Barcelona (Spain), and Menlo Park (California). Our team reflects the diversity and creativity of the people we serve. Join our mission to level the playing field for small business owners through good design.
About the role
Flodesk is building toward a future where small business owners do not need to be expert marketers to grow. Today, our AI helps members create and edit emails, but that is only the starting point. We are building a system that can understand a member's business, brand, audience and past performance, then help turn a goal into a complete marketing campaign across email, workflows, forms, sales pages, segmentation, scheduling, analytics and recommendations.
Reporting to the Head of Product, AI Systems and Core Experience, you will shape how that system behaves and how we know it is good. You will own Flodesk's prompts, agent instructions, quality definitions, evaluation criteria and continuous-improvement loop. You will work directly with prompts, outputs, traces, research and product tools, partnering with engineers on code and infrastructure.
This is not traditional product management conducted through specs and roadmaps alone. It is a hands-on product role that connects product, language and AI systems. You will need to understand how context, models, tools, product state and structured outputs come together to create a member experience, then identify the right change when that experience fails. You are not expected to build or deploy production code, but you must be able to operate the prompt and evaluation system independently.
What you'll do
- Manage member-facing AI behavior across prompt-to-create, agentic editing and future jobs such as segmentation, scheduling, analytics and recommendations
- Create, test, version and document prompts and agent instructions, with a clear record of what changed, why it changed and how behavior improved
- Define what context, member data, tools and product state the AI needs to do a job well, then translate those needs into clear requirements and acceptance criteria
- Own Flodesk's AI evaluation practice, including representative datasets, behavioral scenarios, scoring rubrics, regression coverage and human review
- Run the recurring evaluation and quality-monitoring process independently using available tools. Set baselines and release gates for prompt, model and agent changes
- Review traces and complete AI interactions to diagnose whether a failure most likely comes from the prompt, context, model, tool, product structure or implementation. Fix what you own and turn the rest into clear, actionable changes for engineering
- Turn member research, support feedback, analytics, production failures and domain expertise into priorities, experiments and durable regression cases
- Own the day-to-day product direction and learning agenda for Flodesk's AI experience. Partner with product, design, marketing, copy and engineering to encode a coherent Flodesk point of view into the system
What you bring
- You have shaped or meaningfully improved real LLM or agentic products, with evidence that your work improved user outcomes or product quality
- You have practical knowledge of prompts, context, tools, agent behavior and evaluation, and understand how those parts interact in a production experience
- You are hands-on with prompt iteration, evaluation tools and trace review, and are comfortable working with JSON, schemas and structured outputs
- You have strong product judgment and can turn ambiguous member needs into clear behavior, quality definitions and testable hypotheses
- You bring experimental rigor. You can design representative test sets, build useful failure taxonomies and separate signal from anecdote
- You think in systems and reusable patterns. You can make today's email experience better without hard-coding assumptions that make tomorrow's surfaces harder to build
- You communicate clearly across technical and non-technical disciplines and can make sound tradeoffs involving member value, quality, latency, reliability and cost
Bonus points if you've:
- Owned AI product behavior, conversation design, computational linguistics or prompt systems
- Built content-generation, creative or marketing tools
- Used Langfuse, Braintrust or comparable evaluation and observability platforms
- Worked on products for small businesses, creators or marketers
What we bring
- $150,000 - $200,000 base salary, depending on your location and experience
- Tier 1 city residents: $165,000 - $200,000
- All other locations: $150,000 - $185,000
- Fully paid health insurance for individual coverage
- 16 weeks paid parental leave for non-birthing parents; 22 weeks paid maternity leave for birthing parents
- Unlimited flexible time off
- 401(k) match (US employees only)
- $1,000 annual stipend for learning and development