Jobs Companies Weekday AI Founding Infra Engineer

Über diese Founding Infra Engineer Stelle bei Weekday AI

Weekday AI · Vor Ort · Gurugram, Haryana, India

This role is for one of Weekday’s clients


Min Experience: 4+ years
Location: Gurugram, Haryana, India, Gurgaon, Haryana, India
JobType: full-time

We are looking for a highly skilled and entrepreneurial Founding Infrastructure Engineer with 4–10 years of experience to build, scale, and own the infrastructure powering our next generation of products and AI workloads. This is an early engineering role with significant ownership, where you will work closely with the founding team to design infrastructure from the ground up and establish systems that are reliable, scalable, secure, and cost-efficient.

The ideal candidate has strong hands-on expertise in infrastructure, GPU computing, and Kubernetes, with a deep understanding of distributed systems and cloud-native technologies. You should be comfortable operating in an ambiguous, fast-paced environment and taking projects from architecture and design through implementation and production operations.

Requirements

Key Responsibilities

  • Design, build, and operate highly available and scalable infrastructure for production and compute-intensive workloads.
  • Architect and manage GPU infrastructure for machine learning, AI, and other high-performance computing workloads.
  • Design, deploy, and maintain Kubernetes clusters across cloud and/or on-premise environments.
  • Optimize GPU utilization, scheduling, networking, storage, and compute resources to maximize performance and cost efficiency.
  • Build infrastructure automation using Infrastructure as Code and modern DevOps practices.
  • Establish reliable deployment, monitoring, observability, alerting, and incident-response systems.
  • Develop scalable solutions for container orchestration, workload scheduling, resource allocation, and service discovery.
  • Manage infrastructure lifecycle, including provisioning, upgrades, capacity planning, performance tuning, and disaster recovery.
  • Work closely with application and ML engineers to provide reliable infrastructure for model training, inference, experimentation, and production services.
  • Identify infrastructure bottlenecks and proactively improve system reliability, performance, scalability, and security.
  • Define engineering best practices around infrastructure architecture, Kubernetes operations, CI/CD, and production readiness.
  • Participate in technical strategy and help shape the infrastructure roadmap as an early member of the engineering team.

Must-Have Skills

  • 4–10 years of hands-on experience in infrastructure engineering, platform engineering, DevOps, or SRE.
  • Strong expertise in infrastructure architecture and operations across production environments.
  • Deep hands-on experience with Kubernetes, including cluster architecture, deployments, networking, storage, scheduling, and troubleshooting.
  • Strong understanding of GPU infrastructure, GPU provisioning, utilization, scheduling, and performance optimization.
  • Experience working with containerization technologies such as Docker and Kubernetes-based workloads.
  • Strong understanding of Linux systems, networking, compute, storage, and distributed systems.
  • Experience with cloud infrastructure and services, preferably AWS, GCP, or Azure.
  • Proficiency with Infrastructure as Code tools such as Terraform and configuration-management/automation tools.
  • Experience building CI/CD pipelines and automated infrastructure workflows.
  • Strong debugging and problem-solving skills across complex production environments.

Good-to-Have Skills

  • Experience with NVIDIA GPUs, CUDA, GPU operators, or GPU orchestration.
  • Experience managing large-scale GPU clusters or AI/ML infrastructure.
  • Knowledge of Kubernetes operators, Helm, service meshes, and cluster autoscaling.
  • Experience with high-performance networking, distributed storage, and workload schedulers.
  • Exposure to ML platforms, model serving, inference infrastructure, or large-scale training systems.
  • Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, or similar technologies.
  • Experience building infrastructure at an early-stage startup or as an early engineering hire.
Bereit, sich bei Weekday AI zu bewerben?
Bei Weekday AI bewerben

Über Weekday AI

At Weekday (backed by YC; also Product Hunt #1 product of the day), we are building the next frontier in hiring. We have built the largest database of white collar talent in India and have built outreach tools on top of it to generate highest response rates.

Alle Jobs bei Weekday AI ansehen →

Ähnliche Jobs

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Weekday AI

Alle Jobs bei Weekday AI ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos