Über diese Staff AI Engineer Stelle bei Factored
Fully remote | Complete engagement job
Founded in Palo Alto by Dr. Andrew Ng and Israel Niezen, Factored helps U.S. companies build and scale world-class AI, ML, and Data teams, powered by the top 1% of LATAM talent, with a defining purpose: To empower brilliant humans, unleash their potential, and amplify their impact in the world.
At Factored, you’ll be part of a community that values learning, ownership, and authenticity, where your growth is personal and your ideas matter. We’re transparent, curious, and collaborative. We strive for excellence, celebrate diversity, encourage curiosity, and build an environment where you can truly thrive.
As a Staff AI Engineer at Factored, you will operate at the intersection of technical architecture, Generative AI, and enterprise strategy. Working directly with enterprise clients, you will act as a key technical contributor and trusted advisor.
This role is for technical leaders who combine system architecture mastery and hands-on ML/GenAI engineering with the executive presence needed to navigate ambiguous environments. You will define business problems, architect production-grade AI applications, align senior stakeholders, and own end-to-end delivery to drive measurable impact.
Functional Responsibilities:
- Partner with client executives to translate ambiguous business problems into enterprise AI solution architectures with clear trade-off analyses (cost, latency, risk).
- Design and build scalable backend systems, data pipelines, and APIs integrating LLMs, agentic workflows,and RAGs.
- Implement multi-agent orchestration frameworks and advanced retrieval mechanisms (vector DBs, hybrid search) for complex workflows.
- Deploy and manage cloud-native AI applications across AWS, GCP, Azure, or Databricks using Docker, Kubernetes, Terraform, and CI/CD pipelines.
- Instrument systems with LLM telemetry, cost-tracking, security guardrails, and systematic evaluation harnesses (LLM-as-a-judge) to ensure safety and performance.
- Fine-tune prompts and optimize inference latency using caching, quantization, and cost-reduction strategies.
- Serve as the embedded technical authority within client environments to align cross-functional teams and manage technical risks.
- Elevate team standards (modular code, testing, CI/CD) and mentor client technical staff to build long-term operational autonomy
Qualifications:
- 8+ years of experience in Software/ML Engineering, with 3+ years specifically focused on production GenAI/LLM applications (RAG, agents, tool use) and 2+ years in customer-facing or forward-deployed roles.
- Deep hands-on experience building production systems with Generative AI frameworks (LangGraph, LangChain, LlamaIndex, OpenAI, vector databases).
- Proven ability to architect and scale complex backend microservices and APIs using Python (FastAPI, Django, Flask) alongside relational and NoSQL databases.
- Hands-on expertise building, deploying, and managing cloud-native applications on AWS, GCP, Azure, or Databricks using Docker, Kubernetes, Terraform, MLflow, and automated CI/CD pipelines.
- Experience implementing LLM telemetry, cost-tracking, security guardrails, and systematic evaluation harnesses (LLM-as-a-judge patterns).
- Exceptional ability to structure ambiguous client problems into clear technical requirements and present trade-off analyses (cost, latency, risk) to non-technical executive stakeholders.
- Fluent English communication (written and spoken) with a track record of driving engagements independently in fast-paced, high-stakes environments.
Our Benefits:
- Ownership through equity participation.
- Annual company retreat.
- Education bonus for continuous learning.
- Company-wide winter break.
- Paid time off.
- Optional in-person events and meetups.
- Tailored career roadmaps.
- High-performance culture.