About this Platform Engineer (AI labs) role at Weekday AI
This role is for one of Weekday’s clients
Min Experience: 3+ years
Location: Bengaluru, Karnataka, India, India
JobType: full-time
Our client is seeking a platform engineer to build a robust platform for agent pipelines, ensuring they are modular, reliable, and fast for release to frontier labs. This role involves turning AI training and evaluation workflows into production systems, owning validated releases, and tackling technically challenging problems as frontier models evolve. You will move quickly from experiments to shipped improvements, contributing to a fast-paced and innovative environment.
Requirements
Key Responsibilities
- Create reusable components and stable interfaces for platform architecture, including agents, tools, pipeline stages, execution backends, and lab integrations.
- Choose and operate execution infrastructure for agent runs and evaluation batches, covering scheduling, concurrency control, retries, checkpointing, recovery, and cleanup.
- Build reproducible, securely isolated environments with versioned images, explicit dependencies, controlled access, and reliable setup and resets.
- Own packaging, validation, versioning, and delivery of releases to labs in their required formats.
- Build CI/CD, local tools, and automated checks to enhance developer velocity from code change to evaluated release.
- Expose run status, logs, resource use, failures, and infrastructure cost; improve throughput and recovery, partnering with AI Engineers for debugging.
- Design modular interfaces for evolving components and automate the path from code change to evaluated release.
- Build recoverable batches to improve reliability, throughput, and cost of long-running work, addressing failures in model APIs and stateful environments.
- Automate validation and delivery to ensure releases are reproducible with correct environments, dependencies, and evaluation behavior.
Qualifications
- Strong backend or platform engineering experience, including owning and operating systems in production.
- Ability to design modular APIs, write maintainable Python code, and navigate unfamiliar codebases.
- Hands-on experience with containers, cloud infrastructure, CI/CD, job orchestration, and observability.
- Sound judgment regarding concurrency, recovery, reproducibility, isolation, and infrastructure tradeoffs.
- A habit of building useful developer tools and removing workflow friction.
- Experience with agent runtimes, distributed evaluation, sandboxing, infrastructure as code, or data and ML platforms.
Must-have skills
Backend Engineering, Production Systems, Modular API Design
Good-to-have skills
Cloud Infrastructure, Containers, Observability